tb science was just released like two weeks ago, and opus 5.0 (previous sota) scored around 30%. fable 5.1 scores 52.6 that's a really impressive and meaningful jump. also super interesting how smooth performance scales here with higher thinking modes.
also for tb4, fable/mythos 5.1 are the new champions, again with smooth performance increase at higher thinking levels.
i also find it super challenging to build reliable GUI computer-use settings. i wonder how long it will remain possible to keep scaling up the complexity while still making sure there aren't countless paths that introduce unintended behavior
i'm not sure. i think outside of the tech bubble, people have started using claude/gpt, and they often don't feel substantial progress over the past few months. coding abilities have been sufficient since opus 4.5 for a lot of "office" use cases, but models still have no taste, no "common sense", and their text-writing ability is (subjectively) regressing. at least, these are the vibes i've gotten from talking to people.
for example, handing over a presentation and asking the model to add some slides often requires manual post-processing of every single widget because even frontier models miss the style, tone, and context. it's a lot about taste, but i think improving models in that direction will be important
questionable who's driving adoption in companies and paying the bill, though, but i think a lot of it is the group described above
so annoying. i just wanted to play around with some safety evals. i can change the instructions a bit and proceed a few more steps, but then the conversation gets flagged again, and at some point it can't be recovered anymore. what's the next best choice? going with kimi k3 now
We've pushed a version update to the Terminal-Bench dataset and leaderboard.
Terminal-Bench 4.0 calibrates task resources (time, cpu, memory), implements task fixes, and removes saturated tasks.
just based on my memory, not something iβve verified, but whenever a new benchmark comes out without existing hillclimbing ots datasets from the large data vendors, ant/oai seems still far ahead. i guess picture will change with tb science 0.2+
however, great release!
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains.
An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%.
1/n π