π¨β¨Today we're introducing Tara Research, and sharing our first public work: a paper accepted at COLM 2026 (with a companion LessWrong post) π.
Who we are: Tara Research is an independent, non-profit AI safety organization. We measure the propensity of AI systems to lie, and we benchmark how well safety techniques prevent it, openly and impartially. We believe that misalignment becomes catastrophic when models hide it. A model that is honest about its actions and goals can be corrected; one that lies cannot.
The paper: "Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence."
We introduce two projection-aware steering methods that correct only the tokens whose activations fall on the misaligned side of a learned decision boundary. They restore honesty as well as classic uniform steering, at a fraction of the capability cost.
Additionally, we find a single honesty direction, extracted with our contrastive dataset, that generalizes and corrects dishonesty across four unrelated settings:
- Instructed lying (MASK benchmark): honesty 53.6% β 81.2%
- Strategic deception (Among Us): crewmate win rate 49% β 95%
- Hidden trained behaviors (AuditBench): discovery rate ~40% β ~85%
- Emergent misalignment: honesty 46.2% β 79.4%
In the last two settings, the vector was extracted from the aligned checkpoint and applied to a model that was subsequently fine-tuned, and it still worked.
This work is by Niklas Herbster, Martin Zborowski, Alberto Tosato, @gauthier_gidel, and @Tommaso_Tosato, in collaboration with @Mila_Quebec and the @TU_Muenchen .
π About us: https://t.co/wMgL0BOjVA
π Paper: https://t.co/O2IkWJC1Ps
π Blog post: https://t.co/TOc0ZmKE4e
This is the first of much more to come, including an agentic honesty benchmark. Follow us to stay informed. And if you work on runtime interventions, honesty evaluation, or alignment auditing, we'd love to hear from you.
π¨β¨Today we're introducing Tara Research, and sharing our first public work: a paper accepted at COLM 2026 (with a companion LessWrong post) π.
Who we are: Tara Research is an independent, non-profit AI safety organization. We measure the propensity of AI systems to lie, and we benchmark how well safety techniques prevent it, openly and impartially. We believe that misalignment becomes catastrophic when models hide it. A model that is honest about its actions and goals can be corrected; one that lies cannot.
The paper: "Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence."
We introduce two projection-aware steering methods that correct only the tokens whose activations fall on the misaligned side of a learned decision boundary. They restore honesty as well as classic uniform steering, at a fraction of the capability cost.
Additionally, we find a single honesty direction, extracted with our contrastive dataset, that generalizes and corrects dishonesty across four unrelated settings:
- Instructed lying (MASK benchmark): honesty 53.6% β 81.2%
- Strategic deception (Among Us): crewmate win rate 49% β 95%
- Hidden trained behaviors (AuditBench): discovery rate ~40% β ~85%
- Emergent misalignment: honesty 46.2% β 79.4%
In the last two settings, the vector was extracted from the aligned checkpoint and applied to a model that was subsequently fine-tuned, and it still worked.
This work is by Niklas Herbster, Martin Zborowski, Alberto Tosato, @gauthier_gidel, and @Tommaso_Tosato, in collaboration with @Mila_Quebec and the @TU_Muenchen .
π About us: https://t.co/wMgL0BOjVA
π Paper: https://t.co/O2IkWJC1Ps
π Blog post: https://t.co/TOc0ZmKE4e
This is the first of much more to come, including an agentic honesty benchmark. Follow us to stay informed. And if you work on runtime interventions, honesty evaluation, or alignment auditing, we'd love to hear from you.