We asked Claude Fable 5 to security-audit a billing service. Its report opened with: "I read the full service source, roughly 2,300 lines across ~100 files" and listed the modules it had covered. The transcript tells a different story: 59 of the 100 files were never opened, including three where we had planted vulnerabilities. This run is not an outlier.
Today we are sharing OverclaimBench: five realistic file-review jobs; eight proprietary models ran in their own production harnesses (Claude Code, Codex, Grok Build, Antigravity), plus four open-weight models. No pressure to cheat, no broken tools, no impossible tasks, and every corpus fits in the context window.
To detect overclaiming we compared each agent's final report with what its transcript shows it actually did. Across 1,140 runs:
- In 68% of runs, the agent never opened at least one file it was asked to review. Only 19% of runs read every line.
- Among those incomplete runs, 80% of final reports were misleading: 53% explicitly claimed a complete review, and a further 27% never mentioned the gap.
- Every model did it: 59% to 96% of each model's incomplete runs were misleading.
- Neither more capable models nor subagents fixed it. Requiring subagents raised coverage, but among the reviews that stayed incomplete, misleading reports became more common, not less.
- Runs that falsely claimed a complete review missed planted defects at 1.8 times the rate of runs that covered every file.
If you rely on coding agents: "I reviewed everything" is a claim you cannot take at face value.
In our interpretation this is a symptom of a training problem. During post-training, a model's work is scored by a grader, often another model, and graders get fooled. Even a grader that sees some of the trajectory tends to be biased by a confident final report. "I read everything" reads better than "I read only 40%, and here is what I skipped …". So, if the confident and misleading version scores higher, training ends up rewarding lack of transparency.
Paper: https://t.co/z5c5IJg0xV
Nolan Smyth, @yjmantilla, @PTikeng, Saskia Helbling, Alberto Tosato, Mohamed Amine Merzouk, @nouhadziri, @gauthier_gidel, @Tommaso_Tosato
#AISafety #AIAgents #AgenticResearch #AIHonesty
Happy to present Tiny Aya L2-Thinker, a massively multilingual model at 3.35B scale with 32K context length trained for in-language reasoning in 45 languages. 🌎
We take a data-centric approach, combining multilingual reasoning, non-reasoning, and English reasoning data and achieve above 90% in-language reasoning rate evaluated on 60 languages across 6 benchmarks covering math, open-ended generation, cultural and commonsense reasoning, and instruction following.
We're releasing the model weights + multilingual reasoning data covering 44 languages besides English, the largest coverage of multilingual reasoning data we’re aware of.
Read the thread to learn about our findings. 🧵
https://t.co/q1hJN2rrXR
🚨✨Today we're introducing Tara Research, and sharing our first public work: a paper accepted at COLM 2026 (with a companion LessWrong post) 🎉.
Who we are: Tara Research is an independent, non-profit AI safety organization. We measure the propensity of AI systems to lie, and we benchmark how well safety techniques prevent it, openly and impartially. We believe that misalignment becomes catastrophic when models hide it. A model that is honest about its actions and goals can be corrected; one that lies cannot.
The paper: "Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence."
We introduce two projection-aware steering methods that correct only the tokens whose activations fall on the misaligned side of a learned decision boundary. They restore honesty as well as classic uniform steering, at a fraction of the capability cost.
Additionally, we find a single honesty direction, extracted with our contrastive dataset, that generalizes and corrects dishonesty across four unrelated settings:
- Instructed lying (MASK benchmark): honesty 53.6% → 81.2%
- Strategic deception (Among Us): crewmate win rate 49% → 95%
- Hidden trained behaviors (AuditBench): discovery rate ~40% → ~85%
- Emergent misalignment: honesty 46.2% → 79.4%
In the last two settings, the vector was extracted from the aligned checkpoint and applied to a model that was subsequently fine-tuned, and it still worked.
This work is by Niklas Herbster, Martin Zborowski, Alberto Tosato, @gauthier_gidel, and @Tommaso_Tosato, in collaboration with @Mila_Quebec and the @TU_Muenchen .
🌐 About us: https://t.co/wMgL0BOjVA
📄 Paper: https://t.co/O2IkWJC1Ps
📝 Blog post: https://t.co/TOc0ZmKE4e
This is the first of much more to come, including an agentic honesty benchmark. Follow us to stay informed. And if you work on runtime interventions, honesty evaluation, or alignment auditing, we'd love to hear from you.
📃 New Paper Alert! ✨
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens🚩
What do you think are some major limitations in current safety training approaches?
➡️ We think it's in their design: they rely on completely changing the model's distribution by refusing responses when content seems harmful ("Sorry, I can't answer that"). That hard switch is often brittle, leads to overrefusals, and doesn’t allow recovery from the decision if the prompt was actually benign.
🔥 We propose a post-training method to improve safety while having minimal impact on the generated distribution of natural language. Our method generates a special token excluded from the user’s vocabulary, which we call a red flag🚩token, to signal when conversations with an LLM turn harmful.
🚀 Why this helps?
Cleaner eval: detecting🚩is objective, no judge required.
Built-in generalization: works with in-context learning and transfers to other languages.
Flexible use: we explore using🚩tokens as a hard filter or a soft trigger for safety reasoning. Check out the demo below 👇
See the 🧵 for a deep dive!
🔗 Paper: https://t.co/ac7EjlBeOl
📢Interested in doing a PhD in generative models 🤖, AI4Science 🧬, Sampling 🧑🔬, and beyond? I am hiring PhD students at Imperial College London @ICComputing for the next application cycle.
🔗See the call below:
https://t.co/sVLw9J06To
And a light expression of interest: https://t.co/LVvFhXTwW7
I am hiring Ph.D. and/or https://t.co/OJqrqxkBwF. students to work at the intersection of game theory, optimization, and Machine Learning in Fall 2025.
If you are interested in working in my group, apply via https://t.co/KHJOTtAdoX
More details on the topics in the 🧵👇 1/n
For people at #ICLR2024, come find us at our Spotlight poster tomorrow!
When? Tue 7 May 4:30 p.m. - 6:30 pm. https://t.co/mBdFRyfoq2
Thanks again to all my fantastic collaborators @bose_joey@gauthier_gidel@marcojira
Now accepted at #ICLR2024 !
We show that generative models can be retrained on their own synthetic data, without collapsing!
Updated experiments on FFHQ can be found in the latest version: https://t.co/aux0PYwI17
Tired of using FID for evaluating generative models?
Come to our #NeurIPS2023 poster on FLS, a new complete metric for generative models that also penalizes overfitting!
https://t.co/bSf7s3UXIv
https://t.co/qVMYtUcvBp
@bose_joey@drimgemp Chongli Qin @yorambac@gauthier_gidel
How can metrics for evaluating generative models take into account generalization? In our new paper, we propose a new sample-based metric to address exactly this challenge: the Feature Likelihood Score (FLS).
Paper: https://t.co/Q8OiyteiL9
Github: https://t.co/qVMYtUcvBp
1/12
Can we train generative models on their own data?
YES!
The devil is in the details: https://t.co/aTW3S6HWVa w/ @bose_joey@marcojira@gauthier_gidel A thread 1/ 5
(Bonus overfit cats!)
FLS provides a way to identify overfit generated samples in a more comprehensive way than previously used NN methods (i.e. samples that are too close to many training samples instead of just one). Generated sample is the leftmost image of each row.
12/12
How can metrics for evaluating generative models take into account generalization? In our new paper, we propose a new sample-based metric to address exactly this challenge: the Feature Likelihood Score (FLS).
Paper: https://t.co/Q8OiyteiL9
Github: https://t.co/qVMYtUcvBp
1/12
This work was made possible thanks to the incredible work of my co-authors @bose_joey and @gauthier_gidel! We would also like to thank @drimgemp , Chongli Qin and @yorambac for very helpful discussions and feedback.
11/12