Announcing Autoresearch Bench
A benchmark for coding agents autonomously tackling research problems
While existing coding benchmarks are saturating quickly, the difference between models in autoresearch loops is stark
It was very interesting seeing the different behaviors for long-horizon tasks:
- GPT-5.6 and Grok iterated fast & often
- Opus liked to make larger, more substantial changes
We'll be sharing the 12 full hour runs soon! Certain models even plan according to how much time they're told they have
@DavidSHolz hey! 1. big fan of your work on Midjourney
we didn't include Fable results because guardrails routed us to Opus and we couldn't get clean runs
on certain cases where we didn’t get restricted, we did see Fable 5.1 underperforming Opus though!
Announcing Autoresearch Bench
A benchmark for coding agents autonomously tackling research problems
While existing coding benchmarks are saturating quickly, the difference between models in autoresearch loops is stark
We just evaluated Gemini-3.8-Flash on our
@PhyseraAI - TB bench. Our analysis on 5 random tasks from the set:
1. It is excellent at deriving and implementing a coherent numerical method. It turned an internally consistent mathematical model into working code, especially when it can derive a checkable optimum.
2. It is inconsistent on edge-case semantics. It knows the individual primitives, but assemble them in the wrong order when a specification has interacting edge cases.
3. It tends to overbuild static-analysis solutions while missing the hardest coverage cases.
4. In multiple tasks long trajectories do not imply better outcomes. More exploration became speculative scope expansion rather than targeted verification.
I think it is promising low-cost choice for numerical/scientific coding / transformations with crisp formulas / tasks where it can independently check residuals or invariants but had fallbacks for production shell tooling / static analysis / clinical derivations / compliance-style work.
Multiple companies are betting on automating the experimental loop. But which model is the best?
- Opus 5 is the best overall
- Grok 4.6 is second and efficient
Live now: our Posttraining & Midtraining Track from AI Engineer World's Fair 2026.
9 talks. 11 speakers. One thesis: a model does not do what you want. It does what you rewarded.
https://t.co/bQPPLSPDdv
- @CompleteSkeptic, TypeSafe AI
- @SeanZCai, State of Data
- @uri_rolls + @Thom_Wolf, Arithmetic + Hugging Face
- @raymondmfeng, Applied Compute
- @willcb, Prime Intellect
- @_potatodonkey_, Emulated
- @kenbwork, LatchBio
- @olive_jy_song + @realDanFu, MiniMax + Together AI
- Ali Khial, G2i
At @aiDotEngineer, we spoke about the unique challenges we're tackling at @emulated_ai:
- Simulation of infra at scale
- Novel research
- Sandbox infra for clouds
https://t.co/f9e5Zr8no5