Super excited to see the benchmark out! We are excited to scale our tasks further from here with even more frontier tasks and longer horizons. This is just the beginning!
Announcing Autoresearch Bench
A benchmark for coding agents autonomously tackling research problems
While existing coding benchmarks are saturating quickly, the difference between models in autoresearch loops is stark
Announcing Autoresearch Bench
A benchmark for coding agents autonomously tackling research problems
While existing coding benchmarks are saturating quickly, the difference between models in autoresearch loops is stark