Our next-gen AI Research Agents are here. 🚀 Building AIRA₂ gave us incredible insights into solving bottlenecks, and we're excited to share them. Async evolutionary exploration, parallel agents, closed generalization gap, scaling laws, aha moments, and more.
Thread & paper👇
So happy I can finally share this with you all.
We spent months building AIRA₃, entered it in a live Kaggle competition against ~4,000 human teams, and watched it teach an LLM to reason on its own.
Gold 🏅 First ever by an autonomous AI system.
Unhackable eval. Early RSI ♻️
As a test of our progress to advance the frontier of AI research, in June we entered the next generation of our autonomous AI research system, AIRA₃, in a live Kaggle competition run by NVIDIA to fine-tune a 30B Nemotron model. The challenge was to teach the model to reason better — all competitors had access to the same information and were graded externally on a private test set.
AIRA₃ placed 8th out of ~4,000 teams to win Gold, outperforming human competitors who had access to the same frontier tools.
We believe this is a reliable signal that AIRA₃ can improve a targeted capability of an AI model at a level similar to human experts.
AIRA₃ won gold in a live Kaggle competition, placing 8th of ~4,000 teams. To our knowledge, the first gold won by an autonomous AI research system.
NVIDIA's competition provided a great eval setup:
- fine-tuning a 30B Nemotron reasoning model
- uncontaminated and externally graded
- comparison with human experts
AIRA₃ performed at the level of the top human competitors.
Excited to share this! A lot has happened since AIRA₂, but the strategy carries forward: find the bottlenecks, address them, repeat. Here we go after three: throughput, evaluation robustness, and smarter agent design. Incredible team effort !
Excited to share AIRA₂ — our next-generation AI Research Agents for ML that address key bottlenecks to scaling.
AIRA₂ achieves SoTA on real-world ML tasks from MLE-bench-30 (81.5% vs 72.7%), exceeds human SoTA on 6/20 diverse AI research tasks from AIRS-Bench (and hacks another 5), while exhibiting strong, predictable scaling properties.
To push the frontier of AI Research, we need systems that scale well. Developing AIRA₂, we learned a lot about the bottlenecks and what it takes to resolve them — insights already driving our next iteration:
1/