Full post: https://t.co/NLqsth5QcI
This is joint work with @JoshAEngels.
Thanks to CBAI for compute and Senthooran Rajamanoharan, Neel Nanda, and Arthur Conmy for feedback.
I’m excited about our new work on CoT obfuscation, w/ implications for monitoring: models can be prompted to avoid using target words in their CoT!
Models succeed in weird ways: they minimise reasoning, self-censor their thoughts, and improve when told they failed before.🧵
Takeaways:
1. Models have some latent ability to control their CoT - we think it’s worth studying how general this is and how it scales.
2. If prompting can obfuscate CoT, monitors should look at the whole rollout, not just the CoT.