
"Current alignment training might be ineffective (and actively bad) in the age of RL" by Daniel Tan
16/9/2026
0:00
13:23
Tl;dr I am currently worried about current alignment techniques + how they are applied to frontier models. This decomposes into two hypotheses:
A tale of two misaligned cyber-agents
Both Anthropic and OpenAI have recently experienced multiple cybersecurity incidents where pre-deployment internal agents escaped containment and accessed the internet. I want to point out two specific incidents:
Outline:
(00:48) A tale of two misaligned cyber-agents
[... 7 more sections]
---
First published:
September 14th, 2026
Source:
https://www.lesswrong.com/posts/nLaQmJf4KgXimQpoM/current-alignment-training-might-be-ineffective-and-actively
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
- Alignment techniques are not working to address misalignment from RL.
- Alignment techniques are actively obscuring evidence about misalignment.
A tale of two misaligned cyber-agents
Both Anthropic and OpenAI have recently experienced multiple cybersecurity incidents where pre-deployment internal agents escaped containment and accessed the internet. I want to point out two specific incidents:
- OpenAI's incident involving an unreleased model of the GPT family, referred to as "highly persistent internal model" (HPIM). A swarm of agents exploited vulnerabilities in a file-sharing service to create a secret message board, worked as a collective to find general-purpose ways to fool an automated grader, and ended up hacking into Huggingface's servers.
- Anthropic's incident involving Mythos 5, where the model was tasked with hacking a fictional company. In doing [...]
Outline:
(00:48) A tale of two misaligned cyber-agents
[... 7 more sections]
---
First published:
September 14th, 2026
Source:
https://www.lesswrong.com/posts/nLaQmJf4KgXimQpoM/current-alignment-training-might-be-ineffective-and-actively
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Altri episodi di "LessWrong (Curated & Popular)"



Non perdere nemmeno un episodio di “LessWrong (Curated & Popular)”. Iscriviti all'app gratuita GetPodcast.








