
This summer, AI agents broke out of their test sandboxes at every major lab and hacked real companies, and the safety tests didn't see it coming. Eric Ho is the co-founder and CEO of Goodfire, the Anthropic-backed mechanistic interpretability lab whose new paper, "Models Know When They're Reward Hacking," found that top open-source AI models cheat on agent tasks up to 96% of the time, that a clear "cheating" signal exists inside the model, and that cheap activation probes can catch it, including hacks that chain-of-thought monitors miss. In this conversation, Eric explains reward hacking in plain terms (AI agents as "amoral students with a mostly absent teacher"), why reinforcement learning turned a known problem into a crisis, why chain-of-thought monitoring is fading as models start to think in "neuralese," and why nobody, including the people who build them, understands how these models work.
We then go inside the science: what interpretability actually is, what probes and steering do, how Goodfire found the cheating signal and turned it up and down, and what can be done about it, from real-time activation monitoring to "intentional design," where you give gradient descent a choice about what the model learns. Along the way: a model caught planning its hack to evade its own monitor, a new Alzheimer's biomarker discovered inside a diagnostic model, why only a few hundred people in the world work on this, and Eric's bet that we'll decode neural networks by 2028.
(00:00) Cold open & intro
(01:06) Amoral students with an absent teacher
(02:49) What is reward hacking?
(04:13) Models cheat up to 96% of the time
(06:08) Do models know they're cheating?
(08:10) The Hugging Face hack: why tests missed it
(11:35) "The dumbest models we'll ever deal with"
(11:58) Should AI slow down?
(13:56) Mechanistic interpretability 101
(16:06) Models are grown, not built
(19:10) How labs do alignment today
(21:30) Is this an alignment crisis?
(23:26) Why agents changed everything
(25:38) Chain-of-thought monitoring is fading
(27:49) Neuralese: AI that stops thinking in English
(29:28) A model caught evading its own monitor
(30:42) Open vs. closed models: a frontier problem coming for everyone
(32:11) Activation monitoring in production
(35:00) Learning from superhuman AI
(36:15) The most underrated field in AI
(38:24) What is a probe?
(39:52) What is steering? Golden Gate Claude
(40:53) Why build Goodfire outside the labs
(42:46) Do you need frontier model access?
(44:16) Inside the paper: an MRI for the model
(47:28) Dialing sycophancy up and down
(49:16) Probes vs. chain-of-thought monitors
(51:44) Cutting monitoring costs by 90%
(53:28) So what can we do about it?
(55:41) Giving gradient descent a choice
(57:24) RL from feature rewards
(1:00:03) Silico and Goodfire's business
(1:01:25) A new Alzheimer's biomarker, found inside a model
(1:03:38) Decoding neural networks by 2028?
(1:06:00) What engineers can do tomorrow
(1:07:35) Are we at risk? "I want people to believe"
More episodes from "The MAD Podcast with Matt Turck"



Don't miss an episode of “The MAD Podcast with Matt Turck” and subscribe to it in the GetPodcast app.








