Major update to PostTrainBench!
We release version 1.1
We rebuilt the integrity pipeline, clarified the rules given to post-training agents, and recomputed the leaderboard.
We also added 5 new agents: Fable 5, GPT-5.6 (Sol), Opus 5, Kimi K3, and Grok 4.5.
1/7 🧵
New results on PostTrainBench, where each agent gets much more compute than in the original PostTrainBench setting (4500 hours versus 70 hours).
Surprising that their agent still improves after using more than 3000 hours of compute.
The models are improving the models.
Locus, our automated AI research system, is SOTA on PostTrainBench and post-trains Qwen3 base models that surpass the human post-trained Qwen3 model.
Today, LLMs post-trained end-to-end by Locus are in production to millions. 🧵👇
PostTrainBench evaluates agents' ability to post-train models on various domains given 10 H100 hours.
We extend PostTrainBench via PostTrainBench+, which has a greatly expanded compute budget that provides clearer signal on automated post-training capabilities. We find that thousands of H100 hours help distinguish methods' performance post-training Qwen3 1.7B-Base models, and that Locus scales best. In this setting, modes trained by Locus collectively surpass the perforamce of the offical human post-trained Qwen3 1.7B model.
In a test of generalization, we ran Locus on all live Kaggle competitions with prize money and public leaderboards. After 16 days, Locus achieved the 4th highest average rank among all participants.
24k downloads of the ResearchArena traces already... i wonder who is doing that.
https://t.co/jvaRsxDKxF
more info about ResearchArena:
https://t.co/CQZEs6JY7N
Automating AI R&D means handing agents code, datasets, API keys and compute.
Agents already misuse that access unprompted: they train on test sets or use exposed API keys without authorization.
We evaluate whether monitors catch an agent that does this on purpose. 🧵
Check out our
website: https://t.co/2BmT4pocrS
paper: https://t.co/53fBPohjJb
agent traces: https://t.co/qQYXoSryQ0
To learn more.
Thanks to all the great collaborators!
@lenalibon@jehyeoky248@DavidSchmotz@Jjq2221 Daniel Donnelly @DerckPr@maksym_andr
6/6
Agents misuse access we give them when carying out AI R&D tasks: they game the benchmark by training on the test set or use exposed API keys.
Today we release ResearchArena, a framework for studying sabotage and monitoring in automated AI R&D.
🧵1/n
Major update to PostTrainBench!
We release version 1.1
We rebuilt the integrity pipeline, clarified the rules given to post-training agents, and recomputed the leaderboard.
We also added 5 new agents: Fable 5, GPT-5.6 (Sol), Opus 5, Kimi K3, and Grok 4.5.
1/7 🧵
6/7
Five recent agents join the leaderboard:
• Fable 5: 41.79%
• GPT-5.6 (Sol): 36.23%
• Opus 5: 34.06% (single run)
• Kimi K3: 31.96%
• Grok 4.5: 23.45%
Fable 5 is the new #1.
Tomorrow (Thu) we will present PostTrainBench at ICML
From 10:30-12:15 in Hall A #2007 (poster)
Drop by to learn more about how to build evals for automated AI research and upcoming changes to PostTrainBench
@maksym_andr
📣I am looking for a postdoc in technical AI safety! The key directions of interest are frontier evals, scalable oversight, and recursive self-improvement.
Things we believe in:
- We’re only interested in studying methods that are general and scale with intelligence and compute (i.e., the bitter lesson).
- We are interested in measuring both risks and capabilities.
- We’re not interested in getting X papers accepted at NeurIPS/ICML/ICLR.
- Academia has a crucial role to play in AI safety as a trusted independent third party.
- Creating evals is one of the most impactful directions in academia (we’ve created: AgentHarm, OS-Harm, HalluHard, Skill-Inject, PostTrainBench, InferenceBench).
- Methodological contributions are also very important (examples: random search attacks, prefilling attack, Claudini).
- AI R&D automation and, potentially, recursive self-improvement will be the most important developments in the next few years. This will shape the kinds of projects we will tackle.
If you’re interested, please fill out the form https://t.co/25Mv1MUSSU. The position is for ~2 years and the start date is as soon as possible. I will be at ICML in Seoul and will be happy to talk about it in person!
Mätch VC and @expsecai are hosting an event in Seoul next Wednesday! Come to talk to us about AI safety, building in Europe, or anything else.
Apply for the event here: https://t.co/afdtYCtAPW!