If you're an AI researcher looking to be challenged, we want to talk.
Some of the most important questions in AI still don’t have good answers.
A lot of the interesting work starts before there is a benchmark, a dataset, or even a clear experimental setup.
Why is a model failing this task? Is it the model, the data, the grader, or the harness? What would we need to measure to know? What data would actually teach the missing capability?
We work on questions like these with frontier AI labs and companies, across agents, evaluations, human and synthetic data, post-training, and new benchmarks.
Rings a bell? Let's chat.
You’ll take problems end to end: design the evaluation, build the dataset or environment, run experiments, inspect trajectories and failure modes, and figure out what the results actually mean.
We move quickly and researchers here have a lot of ownership. If you like ambiguous problems, building the experiment instead of just running it, and getting very close to how frontier models actually succeed and fail, I’d love to talk.
We’re hiring globally and remotely, all seniorities.
Send me a DM.
1/Frontier models can code, operate computers, use tools, and handle increasingly complex agentic tasks.
But show them a short video where something unexpected happens, and they can still miss what seems obvious to us.
Today we’re releasing a study from our lab: Plot Twist Bench 🧵
Is Fable 5.1 the best coding model right now?
Maybe. But after testing it across short and very long coding environments, I think there’s a more interesting question:
How much of “coding performance” is actually the model, and how much comes from the harness and the length of the task?
We got some pretty different answers depending on the setup. 🧵
Really believe in the breakthroughs that AI-powered research will bring.
For me this release reflects a new generation of benchmarks that are becoming popular: to measure building blocks of R&D so we get closer to self improvement.
Exciting direction!
What if an agent gets 80% reward… without actually doing the core task?
That question led us to take a closer look at ProgramBench — and ultimately to double down on the same class of long-horizon reverse-engineering tasks pioneered by the @AIatMeta team.
The goal isn’t simply harder tasks.
It’s more trustworthy reward.
@LucasSmaira breaks down what we found and what we changed ↓
As tasks get longer and useful signals get sparser, long horizon tasks have a reward problem.
In order to address reward density in coding long horizon tasks, we're releasing ProgramBench Vetted: reverse-engineering tasks inspired by Meta's ProgramBench.
The setup is simple: give an agent a program it can execute but cannot read, and ask it to rebuild the program from scratch.
It's one of the most interesting long-horizon coding task designs we've seen.
We tested some of the newest models on it ↓
Contenders keep closing the gap with the frontier, it's impressive.
We've updated our Terminal Tasks v1.0 leaderboard with the recent model releases.
Here's the TLDR:
- Grok grinds its way up and surges into the top 5
- Gemini 3.7 Flash basically doubled its performance since 3.5 Flash.
- Glimmer achieves a third of Hy3's performance despite being 10x smaller.
- Deepseek V4 Pro's new checkpoint basically 3x'd its predecessor
Check it out:
https://t.co/IEKrWKbuF5
Congrats to all the teams for the outstanding work @elon@xai@GoogleDeepMind@AIatMeta@deepseek_ai
Congrats @xAI on Grok 4.6
Frontier intelligence? ✅
Frontier skepticism? still a work in progress 😃
We evaluated Grok 4.6 on GulliBench, and like every frontier model we tested, it still struggles with knowing when to verify instead of trust.
@elonmusk the full benchmark is live—we’d love to hear what you think. https://t.co/iNjKwhTnfj
@logangraham
Fun coincidence: @AnthropicAI published work on epistemic vigilance the same week we released GulliBench.
Different benchmark, same intuition:
As models get smarter, measuring reasoning isn’t enough. We also need to measure when they should verify instead of trust.
What happens when the right answer is hidden behind information the model shouldn't trust?
That's the question behind GulliBench.
Instead of measuring whether a model can reason, GulliBench measures whether it knows when to verify before answering.
The results reveal an interesting gap: stronger reasoning doesn't necessarily mean better judgment.
This is the first benchmark in a broader research direction we're excited to explore as AI systems move from benchmark environments into real-world production.
Read Lucas' thread below for the motivation, methodology, and results.
Frontier models are becoming incredibly good at reasoning.
But anyone deploying them in production has probably seen something like the model confidently trusting information it shouldn't.
It never questions the data it's given although it can perfectly execute every step of the task.
Check it out: https://t.co/oqOQzZB7P5
We believe that pushing the frontier of AI requires scalable and automated evaluation systems.
Today's release is our first public step in that direction.
Computer Anthology is the beginning of our long-term effort to build rigorous, continuously evolving evaluation infrastructure for AI.
Read Lucas' thread below.👇
For the past couple of months, I've been quietly building at Vetto what I believe is essential for the next generation of model evaluation and post training.
If the pace of model development is relentless, so must be the data and evaluation infrastructures that fuel it.
As benchmarks demand greater headroom and fairness, keeping pace through traditional services is increasingly unsustainable.
Computer Anthology is the premiere of our proprietary, scalable engine towards automating an evaluation system that consistently differentiates the frontier.
https://t.co/sp1yBJ4vgN
Selected researchers get expert, human-in-the-loop data labeling from Vetto + mentorship from experienced advisors. Everything needed to go from idea to publication.
Vetto is taking the stage at @BrazilatSV 2026 — one of the most important events connecting Brazilian tech leaders to the global ecosystem. Our CTO, Rodrigo Schmidt, is bringing our vision to the center of that conversation 🚀
#VettoAI#BrazilAtSiliconValley#BSV2026
We’re excited to welcome Rodrigo Schmidt as Co-Founder & CTO.
Former Meta engineering leader, Rodrigo joins us to build the next generation of data intelligence for frontier AI — where learning depends on how experience is structured, not just scale.
Welcome to the team, Rodrigo 👋