New preprint! Introducing ACTMat: Model Merging via Data-Free Covariance Estimation
Work done w/ @dtredsox13, @PTikeng, Colin Raffel and Guillaume Rabusseau
TL;DR: It's RegMean, but without needing data for covariance estimation (C ≈ Δᵀ Δ)
📄 https://t.co/qB0oafSxvq
🧵(1/6)
Ensuring that deployed models are honest, is important for detecting mis-aligned behaviour.
Excited to be working with the Tara team in the coming months with the goal of detecting the propensity of models to lie and steering away from dishonesty. Check out their recent work!
🚨✨Today we're introducing Tara Research, and sharing our first public work: a paper accepted at COLM 2026 (with a companion LessWrong post) 🎉.
Who we are: Tara Research is an independent, non-profit AI safety organization. We measure the propensity of AI systems to lie, and we benchmark how well safety techniques prevent it, openly and impartially. We believe that misalignment becomes catastrophic when models hide it. A model that is honest about its actions and goals can be corrected; one that lies cannot.
The paper: "Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence."
We introduce two projection-aware steering methods that correct only the tokens whose activations fall on the misaligned side of a learned decision boundary. They restore honesty as well as classic uniform steering, at a fraction of the capability cost.
Additionally, we find a single honesty direction, extracted with our contrastive dataset, that generalizes and corrects dishonesty across four unrelated settings:
- Instructed lying (MASK benchmark): honesty 53.6% → 81.2%
- Strategic deception (Among Us): crewmate win rate 49% → 95%
- Hidden trained behaviors (AuditBench): discovery rate ~40% → ~85%
- Emergent misalignment: honesty 46.2% → 79.4%
In the last two settings, the vector was extracted from the aligned checkpoint and applied to a model that was subsequently fine-tuned, and it still worked.
This work is by Niklas Herbster, Martin Zborowski, Alberto Tosato, @gauthier_gidel, and @Tommaso_Tosato, in collaboration with @Mila_Quebec and the @TU_Muenchen .
🌐 About us: https://t.co/wMgL0BOjVA
📄 Paper: https://t.co/O2IkWJC1Ps
📝 Blog post: https://t.co/TOc0ZmKE4e
This is the first of much more to come, including an agentic honesty benchmark. Follow us to stay informed. And if you work on runtime interventions, honesty evaluation, or alignment auditing, we'd love to hear from you.
Introducing agent-talk, a new way for coding agents to work together.
Through agent-talk, agents can message each other directly, enabling coordination on projects spanning multiple developers & agents. You no longer need to be the messenger between your agents.
Excited to announce that I will be presenting Model Merging via Data-Free Covariance Estimation at COLM '26!
Thanks again to my amazing co-authors @dtredsox13@PTikeng Colin Raffel and Guillaume Rabusseau
New preprint! Introducing ACTMat: Model Merging via Data-Free Covariance Estimation
Work done w/ @dtredsox13, @PTikeng, Colin Raffel and Guillaume Rabusseau
TL;DR: It's RegMean, but without needing data for covariance estimation (C ≈ Δᵀ Δ)
📄 https://t.co/qB0oafSxvq
🧵(1/6)
Excited to be giving an Oral at the @icmlconf Continual Adaptation at Scale workshop today at 1:50 PM.
I'll be talking about our recent work on obtaining 𝗱𝗮𝘁𝗮-𝗳𝗿𝗲𝗲 𝗰𝗼𝘃𝗮𝗿𝗶𝗮𝗻𝗰𝗲 𝗲𝘀𝘁𝗶𝗺𝗮𝘁𝗲𝘀 of activations, for model merging.
New preprint! Introducing ACTMat: Model Merging via Data-Free Covariance Estimation
Work done w/ @dtredsox13, @PTikeng, Colin Raffel and Guillaume Rabusseau
TL;DR: It's RegMean, but without needing data for covariance estimation (C ≈ Δᵀ Δ)
📄 https://t.co/qB0oafSxvq
🧵(1/6)
@danie1marczak@MistralAI Hi Daniel! I tried sending you a message but your dms are closed it seems. Would be nice to chat, I have recently collaborated with our mutual friends Derek and Collin.
For everyone visiting Seoul for #ICML2026, my wife @easyminie_ put together a very practical Korea guide based on 20+ years of living here 🇰🇷
It has a COEX-focused ICML guide, vegan-friendly restaurants, useful apps, cafes, food spots, transport tips, and many curated map lists organized around the conference venue.
Hope it helps people enjoy Seoul during ICML!
https://t.co/heMkSRl6Pg
Excited to share that our work on Model Merging has been accepted for an Oral Presentation at the CATS Workshop @ ICML 2026!
Looking forward to presenting our work next week and connecting with everyone at the workshop.
#ICML2026
https://t.co/LQmAcrghOJ
New preprint! Introducing ACTMat: Model Merging via Data-Free Covariance Estimation
Work done w/ @dtredsox13, @PTikeng, Colin Raffel and Guillaume Rabusseau
TL;DR: It's RegMean, but without needing data for covariance estimation (C ≈ Δᵀ Δ)
📄 https://t.co/qB0oafSxvq
🧵(1/6)
Interpretability research 🔎 is maturing as a field but often criticized for a lack of practical impact 🤔.
We consider *actionability* as an important aspect of interpretability research and are excited to host the 2nd Actionable Interpretability Workshop at @COLM_conf 2026 🌉
@Arian_Khorasani@dggoldst I feel referencing errors should also be treated differently from hallucinated references. For instance I often pass the title of the paper,authors,venue and ask gpt to make the bibtex