Arena intern and UCLA PhD candidate, @hgzhou42, introduces Trace-and-Amplify (TA), a framework for collecting training-time reward-hacking trajectories at scale without hacking instructions.
Monitors trained and evaluated on prompt-elicited hacking trajectories can achieve high detection accuracy, but often fail to transfer to training-time reward-hacking trajectories that emerge during RL without hacking instructions. Trace-and-Amplify enables scalable collection of these training-time trajectories, producing monitors that generalize much better to real and held-out hacking types.
Detection accuracy 59.98% (PE-trained) → 90.16% (TA-trained) compared to 97.1% on prompted hacks → 28.0% on training-time hacks.
0:00 – OpenAI's ExploitGym cyberattack benchmark exploit
1:04 – Goodhart's Law and the CoastRunners boat-racing hack (2016)
2:04 – Gaming the evaluator: the robot-hand grasping example (2017)
3:10 – Reward hacking in code generation: hard-coding, test-rewriting, skipping eval
4:20 – A standard defense: reward-hacking monitors
4:58 – Monitor architectures: zero-shot LLMs, fine-tuned BERT, hidden-state probes
6:11 – Where monitor training data comes from today: prompted hacks
7:03 – The core question: do prompted hacks represent real hacks?
7:35 – Why this matters: RL post-training is the standard recipe for frontier models
8:20 – Why nobody's checked this before (hacking is rare, labeling isn't scalable)
9:23 – Introducing the method: Trace-and-Amplify
9:49 – The Tracer: a contradictory unit test that locates evaluation-gaming
10:50 – Amplify: collecting hacking rollouts at scale during RL training
11:32 – Experiment setup: Qwen2.5-Coder, DeepSeek-Coder, LeetCode/TACO
12:16 – Finding #1: prompt-trained monitors don't transfer to real hacks
14:40 – Can strong zero-shot judges (GPT-4.1, o4-mini) do better?
15:48 – Finding #2: monitors trained on real hacks generalize much better to unseen hacks
17:04 – Ruling out artifacts introduced by the method
17:55 – Why the gap? Real hacking is more hidden than prompted hacking
20:05 – Three takeaways, limitations, and future work
Many monitors are trained & evaluated on prompt-elicited hacking trajectories, where models are explicitly asked to exploit the reward signal. But the real test is whether they catch the training-time hacks that naturally emerge during RL training without hacking instructions.
❗️We find a large mismatch:
📉 Detection accuracy: 97.1% on prompted hacks → BUT 28.0% on training-time hacks.
🔍 To study this, we introduce Trace-and-Amplify (TA), a framework for collecting training-time reward-hacking trajectories at scale without hacking instructions.
✅ Monitors trained on TA-collected trajectories perform much better on real inference-time hacks, and generalize better to held-out hacking types:
Detection accuracy 59.98% (PE-trained) → 90.16% (TA-trained).
🌐 Project: https://t.co/unI5L32fRH
📄 Paper: https://t.co/Q7NXjg1ERU
Joint work with @LilichenLi3146, @hgzhou42, @JoLiang17, @zhoutianyi, and @cho_jui_hsieh at PKU, UCLA, and UMD
@PKU1898@umdcs@UCLACS
Grateful for everyone’s contributions and support!