I love TLA+ as much as the next person but CellD & Walgit both had TLA+ specs and I exposed trivial distsys bugs in both in about an hour
I don’t even consider myself good at this stuff I’m just a passionate learner
AI is great but for now the operators skills still matters
We spend a lot of time trying to find where frontier models still fail. Making a benchmark hard is no longer the challenge. Making it hard for the right reasons is.
Push a task far enough and eventually a model will fail. But did you find a real capability gap, or did the task become ambiguous, under specified, or impossible?
This doesn't make benchmarks less valuable. It makes good ones more valuable, and much harder to build.
We pretrained a 2.3B MoE (360M active) Hybrid Mamba-2 that lands within a few points of Llama-3.2-3B using <1% of its pretraining FLOPs.
No dedicated cluster. The run hopped between H100s, A100s, V100s (yes, V100s) and TPU v5p/v6e on a single codebase.
Meet Rigel 🧵
There's been a ton of talk about the role of humans in code review, and when and how humans should be signing off on changes.
I believe that, long-term, humans have no role in routinely reviewing code.
In the last few days we have been using Astra and got really impressed
Most models climb our benchmark Terminal Tasks v1.0 by thinking longer and costing more. Astra didn't.
67% pass rate, the highest we've measured, on 73% fewer tokens than the next best model.
Pass rate per token is the number that should be in every model card, and almost never is.
Check the full leaderboard bellow: https://t.co/5yRnBzVK4r
If you're an AI researcher looking to be challenged, we want to talk.
Some of the most important questions in AI still don’t have good answers.
A lot of the interesting work starts before there is a benchmark, a dataset, or even a clear experimental setup.
Why is a model failing this task? Is it the model, the data, the grader, or the harness? What would we need to measure to know? What data would actually teach the missing capability?
We work on questions like these with frontier AI labs and companies, across agents, evaluations, human and synthetic data, post-training, and new benchmarks.
Rings a bell? Let's chat.
You’ll take problems end to end: design the evaluation, build the dataset or environment, run experiments, inspect trajectories and failure modes, and figure out what the results actually mean.
We move quickly and researchers here have a lot of ownership. If you like ambiguous problems, building the experiment instead of just running it, and getting very close to how frontier models actually succeed and fail, I’d love to talk.
We’re hiring globally and remotely, all seniorities.
Send me a DM.
1/Frontier models can code, operate computers, use tools, and handle increasingly complex agentic tasks.
But show them a short video where something unexpected happens, and they can still miss what seems obvious to us.
Today we’re releasing a study from our lab: Plot Twist Bench 🧵
As tasks get longer and useful signals get sparser, long horizon tasks have a reward problem.
In order to address reward density in coding long horizon tasks, we're releasing ProgramBench Vetted: reverse-engineering tasks inspired by Meta's ProgramBench.
The setup is simple: give an agent a program it can execute but cannot read, and ask it to rebuild the program from scratch.
It's one of the most interesting long-horizon coding task designs we've seen.
We tested some of the newest models on it ↓
Frontier models are becoming incredibly good at reasoning.
But anyone deploying them in production has probably seen something like the model confidently trusting information it shouldn't.
It never questions the data it's given although it can perfectly execute every step of the task.
Check it out: https://t.co/oqOQzZB7P5
For the past couple of months, I've been quietly building at Vetto what I believe is essential for the next generation of model evaluation and post training.
If the pace of model development is relentless, so must be the data and evaluation infrastructures that fuel it.
As benchmarks demand greater headroom and fairness, keeping pace through traditional services is increasingly unsustainable.
Computer Anthology is the premiere of our proprietary, scalable engine towards automating an evaluation system that consistently differentiates the frontier.
https://t.co/sp1yBJ4vgN
A huge amount of the Anti-AI code sentiment massively overestimates the quality of human code outside of a very small set of open source and high quality company codebases. Human Slop is everywhere and can trivially be improved on by any opus level model.
“All the good ideas are taken.”
“If I was born 5 years earlier I’d be rich.”
“I missed the window.”
People said this about dot-com, Web 2.0, social, mobile, cloud… and now AI. There’s always another window. The people who aren't complaining are the ones who are seizing it.