Great insights from @sanyamsatia on why most evals fail, with a simple recipe on what makes a good eval
"I've found that good evals have the same shape usually: they are easy to understand, have a simple prompt (few lines at most) and are simple to verify, but require a lot of work"
Good evals take serious time and rigor to build, yet most never get that attention despite the compute benchmarks steer during model and harness development.
I wrote up the most common patterns I've seen that make evals lose their signal. https://t.co/TRZpUeUJHr
FrontierBench is live! Excited to have been a task author and reviewer on this effort, with @boolean_ai as a data partner.
The team was rigorous in keeping the quality bar high at scale, across a wide range of domains. Congrats to @ryan_marten and everyone involved on the launch!
@martin_casado Unreal stuff !!!
Watching that goal live from behind the net, surrounded by a stadium packed with Argentinians who were non stop “alentando” their team, was a surreal experience
Vamos 🇪🇸🇪🇸🇪🇸🇪🇸!!!!
We ran a frontend eval from an in-progress internal benchmark. Kimi K3 is not at the same level as current frontier models like GPT 5.6 Sol or Fable 5. It's closest to Opus 4.7 on this eval so it's 3 months behind frontier. An impressive result nonetheless.
Evaluating frontier models is going to increasingly require very high taste and in-depth domain expertise.
@tszzl@cpaik I wouldnt be that sure. We ran blind experiments with hundreds of engineers using claude code where models were switched randomly by a proxy. Chinese open-source got solid feedback. The data clearly shows that the benchmaxxed but unusable era is ending.
You asked for it and now you have it: Sticky models are here 🚀
Now you can make any model such as GPT-5-Codex, Gemini, Grok, GLM-4.5 default in Claude Code.
Comment + RT for free access this week.
GPT-5-Codex is now available in Claude Code with Bonsai! Add @gpt5 in your prompt to route it to GPT-5-Codex. No markups, no data retention. Free this week.
Software thrives on freedom - choosing tools, mixing workflows, building without barriers. Today’s AI coding assistants are turning into walled gardens, you can only use Claude models inside Claude Code, and OpenAI models inside Codex. We believe that if you pay for a model, you should be able to use it anywhere.
We are tearing these walls down. With Bonsai you can use any major model in Claude Code today. Next, we’ll liberate every coding assistant. We will build a future where AI development is open, interoperable, and in your hands.
For the last couple of weeks I've been working with @shreyshahi and @sanyamsatia figuring out how to make better software.
We created Bonsai, a handy tool that allows us to use any model from Claude Code.
Try it out for free at https://t.co/RIvwqjXPQa
https://t.co/o7TxX9Rjio
We built Bonsai for ourselves and use it daily. Use GPT-5, Gemini, Claude, and Grok together in Claude Code to cross-review work, catch issues, and ship better code faster. Bring your Claude subscription and use it for free.