Announcing Autoresearch Bench
A benchmark for coding agents autonomously tackling research problems
While existing coding benchmarks are saturating quickly, the difference between models in autoresearch loops is stark
Today, I am excited to announce Animation Bench. Our first benchmark to evaluate frontier models on web animation reconstruction.
Frontend generation is a common commercial use case for coding agents. In addition to static layout it contains transitions and interactions triggered by dynamic behaviours. We evaluated four frontier models on 48 tasks from real websites across three axes:
1. Visual Similarity
2. Motion Consistency
3. Layout Correctness.
We find that motion consistency is often incomplete or missing.
You can kind of just train specialized computer-use agents.
Here is a small Qwen3-VL-2B fine-tuned to do instance segmentation of cells through computer-use actions in a headless labeling app. It takes in screenshots as input and outputs commands to draw polygons, zoom, navigate, etc...
All navigation and labeling in this video is done by Qwen.
Inspired by Astra's computer use abilities, who can also label data quite well, albeit quite slowly.
Autoresearch Bench is our first attempt at measuring a model's ability to make meaningful progress on research problems over extended time horizons.
Very excited to get this out, and so proud of the hard work from our team!
Super excited to see the benchmark out! We are excited to scale our tasks further from here with even more frontier tasks and longer horizons. This is just the beginning!
It was very interesting seeing the different behaviors for long-horizon tasks:
- GPT-5.6 and Grok iterated fast & often
- Opus liked to make larger, more substantial changes
We'll be sharing the 12 full hour runs soon! Certain models even plan according to how much time they're told they have
@DavidSHolz hey! 1. big fan of your work on Midjourney
we didn't include Fable results because guardrails routed us to Opus and we couldn't get clean runs
on certain cases where we didn’t get restricted, we did see Fable 5.1 underperforming Opus though!
Announcing Autoresearch Bench
A benchmark for coding agents autonomously tackling research problems
While existing coding benchmarks are saturating quickly, the difference between models in autoresearch loops is stark
We just evaluated Gemini-3.8-Flash on our
@PhyseraAI - TB bench. Our analysis on 5 random tasks from the set:
1. It is excellent at deriving and implementing a coherent numerical method. It turned an internally consistent mathematical model into working code, especially when it can derive a checkable optimum.
2. It is inconsistent on edge-case semantics. It knows the individual primitives, but assemble them in the wrong order when a specification has interacting edge cases.
3. It tends to overbuild static-analysis solutions while missing the hardest coverage cases.
4. In multiple tasks long trajectories do not imply better outcomes. More exploration became speculative scope expansion rather than targeted verification.
I think it is promising low-cost choice for numerical/scientific coding / transformations with crisp formulas / tasks where it can independently check residuals or invariants but had fallbacks for production shell tooling / static analysis / clinical derivations / compliance-style work.
Multiple companies are betting on automating the experimental loop. But which model is the best?
- Opus 5 is the best overall
- Grok 4.6 is second and efficient