
"“I am an AI Safety Researcher”" by Ashe Vazquez Nuñez
25/09/2026
0:00
27:44
Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and editing.
This post reflects on the tortured distinction between "safety" and "capabilities" in AI research.
Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught, and is currently arguing that key strategic decision-makers in and around "AI safety" have brought about the AI labs' stampede towards Artificial Superintelligence (ASI). This post instead looks at the following problem: how does one conduct alignment research without contributing to capabilities? It proposes decisions an individual or a small research group can take to do good work in AI.
At the end, I discuss possible objections: namely, that my proposals fail to 'maximise impact'. I lay out why this meme is poisonous and usually backfires, and conclude by rejecting it entirely.
Two examples of failure
My first claim is that 'safety' and 'research' are two concepts that are in routine tension with one another. I illustrate this through examples of work that did too much of one at the expense of the other.
Example: (mechanistic) interpretability
In limiting its scope [...]
---
Outline:
(01:12) Two examples of failure
(01:27) Example: (mechanistic) interpretability
(04:30) Example: MIRI and Recursive Self-Improvement
(09:49) The curse of science
(12:11) A note on the AI labs
(15:53) So what do you do?
(17:00) The information you give away
(20:16) The information you let in
(21:46) But what about impact?
(22:42) The virtue of taking things slow
(26:50) Appendix: caveat for policy work
The original text contained 19 footnotes which were omitted from this narration.
---
First published:
September 23rd, 2026
Source:
https://www.lesswrong.com/posts/HekpnSkrt89tMm3Dc/i-am-an-ai-safety-researcher
---
Narrated by TYPE III AUDIO.
This post reflects on the tortured distinction between "safety" and "capabilities" in AI research.
Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught, and is currently arguing that key strategic decision-makers in and around "AI safety" have brought about the AI labs' stampede towards Artificial Superintelligence (ASI). This post instead looks at the following problem: how does one conduct alignment research without contributing to capabilities? It proposes decisions an individual or a small research group can take to do good work in AI.
At the end, I discuss possible objections: namely, that my proposals fail to 'maximise impact'. I lay out why this meme is poisonous and usually backfires, and conclude by rejecting it entirely.
Two examples of failure
My first claim is that 'safety' and 'research' are two concepts that are in routine tension with one another. I illustrate this through examples of work that did too much of one at the expense of the other.
Example: (mechanistic) interpretability
In limiting its scope [...]
---
Outline:
(01:12) Two examples of failure
(01:27) Example: (mechanistic) interpretability
(04:30) Example: MIRI and Recursive Self-Improvement
(09:49) The curse of science
(12:11) A note on the AI labs
(15:53) So what do you do?
(17:00) The information you give away
(20:16) The information you let in
(21:46) But what about impact?
(22:42) The virtue of taking things slow
(26:50) Appendix: caveat for policy work
The original text contained 19 footnotes which were omitted from this narration.
---
First published:
September 23rd, 2026
Source:
https://www.lesswrong.com/posts/HekpnSkrt89tMm3Dc/i-am-an-ai-safety-researcher
---
Narrated by TYPE III AUDIO.
D'autres épisodes de "LessWrong (Curated & Popular)"



Ne ratez aucun épisode de “LessWrong (Curated & Popular)” et abonnez-vous gratuitement à ce podcast dans l'application GetPodcast.








