🚨✨Today we're introducing Tara Research, and sharing our first public work: a paper accepted at COLM 2026 (with a companion LessWrong post) 🎉.
Who we are: Tara Research is an independent, non-profit AI safety organization. We measure the propensity of AI systems to lie, and we benchmark how well safety techniques prevent it, openly and impartially. We believe that misalignment becomes catastrophic when models hide it. A model that is honest about its actions and goals can be corrected; one that lies cannot.
The paper: "Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence."
We introduce two projection-aware steering methods that correct only the tokens whose activations fall on the misaligned side of a learned decision boundary. They restore honesty as well as classic uniform steering, at a fraction of the capability cost.
Additionally, we find a single honesty direction, extracted with our contrastive dataset, that generalizes and corrects dishonesty across four unrelated settings:
- Instructed lying (MASK benchmark): honesty 53.6% → 81.2%
- Strategic deception (Among Us): crewmate win rate 49% → 95%
- Hidden trained behaviors (AuditBench): discovery rate ~40% → ~85%
- Emergent misalignment: honesty 46.2% → 79.4%
In the last two settings, the vector was extracted from the aligned checkpoint and applied to a model that was subsequently fine-tuned, and it still worked.
This work is by Niklas Herbster, Martin Zborowski, Alberto Tosato, @gauthier_gidel, and @Tommaso_Tosato, in collaboration with @Mila_Quebec and the @TU_Muenchen .
🌐 About us: https://t.co/wMgL0BOjVA
📄 Paper: https://t.co/O2IkWJC1Ps
📝 Blog post: https://t.co/TOc0ZmKE4e
This is the first of much more to come, including an agentic honesty benchmark. Follow us to stay informed. And if you work on runtime interventions, honesty evaluation, or alignment auditing, we'd love to hear from you.
Excited to announce that I will be presenting Model Merging via Data-Free Covariance Estimation at COLM '26!
Thanks again to my amazing co-authors @dtredsox13@PTikeng Colin Raffel and Guillaume Rabusseau
Excited to be giving an Oral at the @icmlconf Continual Adaptation at Scale workshop today at 1:50 PM.
I'll be talking about our recent work on obtaining 𝗱𝗮𝘁𝗮-𝗳𝗿𝗲𝗲 𝗰𝗼𝘃𝗮𝗿𝗶𝗮𝗻𝗰𝗲 𝗲𝘀𝘁𝗶𝗺𝗮𝘁𝗲𝘀 of activations, for model merging.
Excited to share that our work on Model Merging has been accepted for an Oral Presentation at the CATS Workshop @ ICML 2026!
Looking forward to presenting our work next week and connecting with everyone at the workshop.
#ICML2026
https://t.co/LQmAcrghOJ
Reviewers & ACs for #ICML2026 have been recognized for their service!
- Reviewers: 4439 Gold (free registration), 4437 Silver. 17749 total reviewers were assigned >= 1 paper
- ACs: 1647 receive free registration, out of 1691 who were assigned >= 1 paper
TY for your hard work!
New preprint! Introducing ACTMat: Model Merging via Data-Free Covariance Estimation
Work done w/ @dtredsox13, @PTikeng, Colin Raffel and Guillaume Rabusseau
TL;DR: It's RegMean, but without needing data for covariance estimation (C ≈ Δᵀ Δ)
📄 https://t.co/qB0oafSxvq
🧵(1/6)
@introspection@Mila_Quebec 2) We also show that the commonly used L2 norm is not a reliable proxy for explaining grokking.
3) Finally, we investigate how factors such as training dataset choice and model overparameterization impact grokking delay
The figure below summarizes our contribution fairly well
@introspection@Mila_Quebec 1) We show that grokking time scales proportionally to 1/(α * β) when minimizing composite objectives of the form f = g + βh using gradient descent with learning rate α, where g is the training error and h is any regularizer that enforces an inductive bias toward generalization.
📝New preprint out!🔍 We used dynamical systems theory to predict "Grokking" — the phase transition leading to "Generalization Beyond Overfitting" discovered at @OpenAI last year— Big kudos to our @Mila_Quebec dream team, especially @PTikeng!