A lot of people are saying Anthropic has already nerfed Claude Opus 5.5.
We have every BridgeBench result for Opus 5.5 from day 1 of launch.
We're retesting soon.
If the scores drop, you'll see it here first.
Have you noticed a difference?
The Amish girls are educated up until 13-14. They have a fertility rate of 6.1, church attendance at 85%, marriage age 21 to 22, virginity rate 90-95%, and divorce rate less than 1%.
College educated career woman have a fertility rate of 1.73, church attendance of 68%, marriage age 27 to 28, virginity rate 25-30%, and divorce rate of 32%.
But go ahead and tell me how important it is that your daughters go to college and have a career.
OpenAI models are way more token efficient than Claude.
Opus 5.5 is the most token hungry model on the chart. Over 4x more than GPT 6 Astra for the same task.
More tokens means higher costs, slower outputs, and burning through your usage limits faster.
Hot take: this might be the most accurate AI model tier list rn
I’ve used almost 70% of these models myself, and this is the closest ranking I’ve seen to reality.l
If you’re confused about what’s actually good, mid, or straight-up bad right now
Just go with this this..
Eu, sênior, anos de experiência, vou fazer 100% vibe codec. você, estudante, júnior, DEVE escrever o máximo possível na mão, deixar a IA fazer tarefas mais mundanas, em paralelo, e você DEVE revisar tudo manualmente. Você precisa de dezenas de milhares de horas de prática deliberada, estudo, pra se tornar alguém que, depois, vai saber o que mandar a IA fazer por você, não antes.
Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max) and Claude Opus 5.5 (xhigh) currently top it at 63% and 62%
Terminal-Bench-Science 0.1 was announced in August 2026, built by @StevenDillmann and researchers at Stanford with the @terminalbench team and a global community of open source scientific contributors. It has 70 expert-curated tasks based on real research in five domains: life sciences (19), physical sciences (17), mathematical sciences (17), engineering sciences (9), and earth sciences (8).
As with Terminal-Bench, each task drops an agent into a sandbox environment with the data, tools and instructions for a task, and the agent runs from start to finish. Automated tests grade every task pass/fail, and we report the average pass@1 over 3 attempts.
Key takeaways:
➤ Only two models score above 50%: GPT-6 Astra (max) at 63%, and Claude Opus 5.5 at 62% (xhigh) and 59% (max). Recent releases have made large advancements, but even the top model, GPT-6 Astra, has significant headroom
➤ Life sciences is the lowest-scoring domain for most leading models, though domain scores are noisy with 8 to 19 tasks each. Claude Opus 5.5 (xhigh) passes 71% of mathematical sciences tasks but 46% of life sciences tasks
➤ The best open weights models, GLM-5.3 (max) at 10% and DeepSeek V4.1 Flash (max) at 9%, sit more than 50 points below the leaders
jev + opus 5.5... i simply can't comprehend why everyone isn't building this yet.
in my workflow, this cut costs and time by ~80%. i think it's one of the best ways to use it.
→ pick relevant project notes before loading the context
→ route suitable tasks to a faster worker
→ choose a recovery path when a tool fails
→ run focused checks before the full test suite
opus handles the hard reasoning. jev picks from options the harness prepares and validates.
i explain how to build the decision layer in the article below:
REPARE NOS INGREDIENTES DOS DOIS PÃES LADO A LADO.
Bruno Monteze compara a composição do pão de padaria com a do pão de forma industrializado e lê os componentes de cada um: farinha, água, fermento e sal de um lado; farinha, açúcar, vinagre, óleo vegetal e conservantes do outro.