This is exactly what we keep seeing in our benchmarks at @VettoAi : you can’t benchmark the model independently from the harness.
Same model, same tasks — very different economics.
GPT-6 Astra: $0.80/trial with Terminus-2 vs. $2.30 with Codex.
Claude Fable 5.1: $2.00 with Terminus-2 vs. $2.70 with Claude Code.
The performance/cost frontier isn’t just a model property anymore. It’s a model × harness property.
Full leaderboard:
https://t.co/VOYcn5Kol2
5/ We spent the last few years building a supply chain for training models. We now need one for challenging them.
My bet is that part of the data/evals industry becomes a small set of trusted organizations with frontier access, secure infrastructure, deep domain expertise and constantly changing tests.
If we want independent answers about what frontier models can actually do, someone other than the labs needs to be able to test them.
@elonmusk@samaltman@darioamodei@OpenAI@xai@Anthropic - let us test yours.
1/ AI safety is missing a pretty basic institution: the independent auditor.
Today, the companies building the most powerful models are also largely responsible for telling us what those models can and can’t do.
Dario Amodei has been making the case for third-party evaluation, and I think he’s right. As capabilities grow, “we tested our own model” can’t be the final form of oversight.
4/ I think this is also where the role of data and eval companies starts to change.
The best ones already know how to find rare experts, turn real-world problems into controlled environments, build graders, continuously refresh hard datasets and run models against them.
That looks a lot like the foundation of an independent audit layer.
3/ We ran into this building's internal cyber benchmarks.
We ask agents to find and exploit real, post-cutoff vulnerabilities blind. With some frontier APIs, safeguards can stop the run after the model has already found the vulnerability.
Did the model fail, or was it prevented from continuing?
Those are very different safety conclusions.
In the last few days we have been using Astra and got really impressed
Most models climb our benchmark Terminal Tasks v1.0 by thinking longer and costing more. Astra didn't.
67% pass rate, the highest we've measured, on 73% fewer tokens than the next best model.
Pass rate per token is the number that should be in every model card, and almost never is.
Check the full leaderboard bellow: https://t.co/5yRnBzVK4r
I remember many friends studying hard to figure out about Navier Stokes while in college, it is one of those big topics that shape our understanding about the world.
Now, AI is solving not only this but many other big challenges. It's just mesmerizing!
Congrats to all the people involved, regardless of who was the first to do it, in the broad scheme of things it is an outstanding feat for humans.
How safe should a model that can change the world around you be?
Far safer than anything we tolerate today, I suppose. As models gain real agency, “usually aligned” stops being good enough.
Intelligence itself is increasingly being addressed. If digital work becomes close to unconstrained, the bottleneck moves to reality: how quickly agents can experiment, observe, act and learn from what happens.
Physics already gives us a preview. Ideas can move faster than our ability to build and test them.
I think how we feed real-world experience back into models, and how fast we can close that loop, will shape the next frontier of AI.
And once models are acting on reality rather than just producing outputs, safety becomes a very different problem.
If you're an AI researcher looking to be challenged, we want to talk.
Some of the most important questions in AI still don’t have good answers.
A lot of the interesting work starts before there is a benchmark, a dataset, or even a clear experimental setup.
Why is a model failing this task? Is it the model, the data, the grader, or the harness? What would we need to measure to know? What data would actually teach the missing capability?
We work on questions like these with frontier AI labs and companies, across agents, evaluations, human and synthetic data, post-training, and new benchmarks.
Rings a bell? Let's chat.
You’ll take problems end to end: design the evaluation, build the dataset or environment, run experiments, inspect trajectories and failure modes, and figure out what the results actually mean.
We move quickly and researchers here have a lot of ownership. If you like ambiguous problems, building the experiment instead of just running it, and getting very close to how frontier models actually succeed and fail, I’d love to talk.
We’re hiring globally and remotely, all seniorities.
Send me a DM.
How safe should a model that can change the world around you be?
Far safer than anything we tolerate today, I suppose. As models gain real agency, “usually aligned” stops being good enough.
Intelligence itself is increasingly being addressed. If digital work becomes close to unconstrained, the bottleneck moves to reality: how quickly agents can experiment, observe, act and learn from what happens.
Physics already gives us a preview. Ideas can move faster than our ability to build and test them.
I think how we feed real-world experience back into models, and how fast we can close that loop, will shape the next frontier of AI.
And once models are acting on reality rather than just producing outputs, safety becomes a very different problem.
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
1/Frontier models can code, operate computers, use tools, and handle increasingly complex agentic tasks.
But show them a short video where something unexpected happens, and they can still miss what seems obvious to us.
Today we’re releasing a study from our lab: Plot Twist Bench 🧵
3/ AI progress isn’t uniform across capabilities.
Good benchmarks help us identify where the frontier still has gaps, understand what models are actually struggling with, and ultimately point to where there’s room for progress.
That’s what we’re trying to map with Vetto Labs.
Full Plot Twist Bench results: https://t.co/MlpPQ3jwer
2/ We created 26 short videos with deliberate twists and 48 questions that can’t be answered by predicting what usually happens.
You have to understand what actually happened on screen.
Humans: 85%
Best model (GPT-5.6 Sol): 75% pass@1
And one result stood out: on a 27-second clip, all 12 models failed. All 5 human testers got it right.
So is Fable 5.1 the best coding model?
It has a very good case on some of the things we’re measuring.
I’m just less sure that “best coding model” is a very stable concept anymore.
Change the harness, the horizon, or what you count as success, and the ranking can move quite a lot.
https://t.co/SPAtbzFGQJ
@alexalbert__@bcherny@RLanceMartin@mikeyk@trq212@felixrieseberg@DarioAmodei
There’s also real progress here.
We had a similar Fable 5 run hit our $300 budget and 400+ steps.
5.1 is much better.
We’re seeing big jumps from Gemini 3.8 Flash and Muse Spark too, although both still struggle on some of the harder tasks we’ve been testing.