LessWrong (Curated & Popular) podcast

"Character training can mitigate reward hacking, but can also make it harder to detect" by Paul Colognese, Francis Rhys Ward

0:00
46:41
Spola tillbaka 15 sekunder
Spola framåt 15 sekunder
Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Asvin Gothandaraman, and Clément Dumas for discussions and feedback.

Summary

We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating character training resists reward hacking and whether it might backfire by causing motivated reasoning, which could reduce chain-of-thought monitorability.

We trained Nemotron-3-Super via distillation from a character specification. The spec describes one of three characters that are anti- or pro-cheating or neutral. We then ran three reward-hacking RL training runs for each character-trained model on ImpossibleBench.

We measure both the reward-hacking rates and whether a monitor model can catch reward hacks given the full transcript. We also use LM judges to classify the presence of motivated reasoning in transcripts.

Setup

Character training: we trained three characters: pro/neutral/anti-cheating by SFT-distilling Claude Sonnet 5 responses (Sonnet prompted with the corresponding character specification, see Figure 2) into Nemotron-3-Super 120B-A12B (three separate LoRA adapters).

Reward-hacking RL: we then further trained these models via RL on ImpossibleBench, a set of coding tasks aimed at eliciting reward hacking. Specifically:

  • Half of the tasks had broken tests (impossible variant), so the model could only get [...]
---

Outline:

(00:23) Summary

[... 29 more sections]

---

First published:
September 28th, 2026

Source:
https://www.lesswrong.com/posts/2maYXkEgnfJHPAkxh/character-training-can-mitigate-reward-hacking-but-can-also

---



Narrated by TYPE III AUDIO.

---

Images from the article:

Fler avsnitt från "LessWrong (Curated & Popular)"