Anthropic has launched Claude Sonnet 5.5: it scores 56 on the Artificial Analysis Intelligence Index, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we’ve seen
With max effort, Sonnet 5.5 gains 18 points over Sonnet 5 and to #2 on the Intelligence Index behind only Opus 5.5 (max). Anthropic has priced Sonnet 5.5 identically to Sonnet 5 at $0.2/$2/$10 per 1M cache input/input/output tokens, however it outputs a higher number of Output Tokens per Task and costs $7.60 per task (~50% higher than Sonnet 5’s Cost per Task)
Key takeaways:
➤ Meets leading models on agentic terminal use and knowledge work: in Terminal-Bench 4.0, Claude Sonnet 5.5 reaches 64% against 60% for Opus 5.5 and GPT-6 Astra. On AA-Briefcase (1811 vs 1822 Elo), GDPval-AA (1844 vs 1846 Elo), and AutomationBench-AA (71% vs 70% headline score), Sonnet 5.5 reaches parity with Opus 5.5, albeit with significantly higher token usage to achieve it
➤ Heaviest token use we have measured: at max effort, where it reaches performance nearing that of Opus 5.5, Claude Sonnet 5.5 used ~193k Output Tokens per Intelligence Index Task. This is the highest token use we have measured on around 60% higher than Opus 5.5 (max) or Sonnet 5 (max) and ~7x GPT-6 Astra (max)
➤ Pricing remains at $2/$10 per million tokens of input/output, matching GPT-6 Sol. At this pricing Claude Sonnet 5.5 sits off the Intelligence vs. Cost per Task Pareto Frontier. At high effort levels it sits behind Opus 5.5, while lower efforts have GPT-6 Astra or Sol configurations delivering equivalent performance for lower cost. The high effort setting is the most competitive on this basis, sitting very narrowly behind GPT-6 Sol on Intelligence at effectively the same Cost per Task
➤ Behind Opus 5.5 on factual knowledge and scientific reasoning: as a smaller class model, Sonnet 5.5 still lags on factual knowledge in AA-Omniscience compared to Opus 5.5. It scores 54% against 66% for factual accuracy, though with a lower hallucination rate (47% against 59%). It also sits ~6 points lower on Humanity's Last Exam and SciCode compared to Opus
These evaluations were conducted on a pre-release deployment of Claude Sonnet 5.5, which Anthropic found to have a bug that can degrade responses to requests that use structured outputs. This is fixed for the public release and Anthropic expects minimal change or slightly understated performance, but we will be re-running relevant evaluations soon.
Other model details:
➤ Context window: 1 million tokens with image and text input, unchanged from Sonnet 5
➤ Pricing: unchanged from Sonnet 5’s latest $2/$10 per 1M input/output tokens; cache writes at $2.5, cache reads $0.2
➤ Effort settings: five (low, medium, high, xhigh, max). Intelligence Index evaluations were run at all five with Anthropic's default fallback enabled. We see Sonnet 5.5 fall back in ~0.1% of tasks across the Intelligence Index, primarily in TerminalBench 4.0, falling back to Sonnet 5 in all cases.
I don’t watch live sports anymore, huge waste of time. Instead I consult a panel of frontier models, they simulate the game and report back on the outcome along with a curated “suite” of emotions I would’ve felt while watching.
Opus 5.5 is setting unreal expectations for Fable 5.5.
Do we even need more intelligence at this point? Give these models enough context about your work, and they'll do all of it.
Why does Opus 5.5 feel practically unlimited when Fable 5.1 was so heavily limited?
It's a combination of two things: Opus's efficiency, and the 50% limit on Fable.
Claude Code subs are very generous with their usage, but only half is allowed to be used by Fable.
Separately, Opus is way more gentle with costs. Opus 5.5 on High is over 2x cheaper than Fable 5.1 High.
The result is that, roughly, going from Fable 5.1 high to Opus 5.5 high is a 4.3x increase in limits.
Fable 5.1 xhigh to Opus 5.5 high (move I made) is a 6.6x increase 🤯
Simple illustration of how effort works.
Effort = how hard the model tries on your prompt.
I assumed effort was a harness-level setting. It's not. The model reads the effort level as part of its input, and it's trained to decide how much work a task deserves.
Opus 5.5 is an insane release. It feels like Anthropic wrote down everything they needed to beat OpenAI, then did all of it.
It's cheap, it's fast, and most importantly it talks in plain English.
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023.
That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR.
We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use.
Read more: https://t.co/RuEosScSMb
Opus 5.5 communicates more naturally, addressing some of the most common feedback we heard on Opus 5.
It puts the most important information up front and follows the writing rules you give it, which makes long sessions easier to follow.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
EXCITED TO LAUNCH: Akai (https://t.co/mMBMmX4NCr)
Deel added >$140M ARR in 90 days without increasing headcount by automating~600 Full Time Employees' equivalent in work with Akai.
Akai was an internal tool to automate our painfully repetitive operations in Finance, HR, Accounts Payable, and Compliance, etc. We never intended to make this a product.
But we watched revenue per employee grow from $130K to $215K
We built >8k agents that do the work of ~600 employees
It had such a dramatic impact on our business that today we are launching it for everyone.
How it works: Say you're automating payment reconciliation:
1. Record your screen while manually matching a messy transaction and Akai will capture your screen, voice, server requests
2. Akai will see that you pulled unformatted wire transfer info from an archaic bank portal, put it in some excel sheet, checked NetSuite invoices, payment history, and put a ticket on Zendesk
3. Akai reads between the lines and build a workflow + steps + conditional guardrails. It learns tacit edge cases, like resolving malformed invoice references without you writing a single regex
4. Simply connect NetSuite, your ledger, Zendesk, PSPs, and even legacy bank portals with zero API access
5. Run the workflow and tell it what to adjust in plain English: "strip slashes on wire memos and auto-apply partial payments." It adapts instantly
6. Once it works for you, add 100s of colleagues. Your entire payment ops team forks and extends the workflow for new PSPs, secondary ledgers, or regional settlement rules
7. We automated 85% of our payment reconciliation end to end, eliminating 500+ hours of soul-crushing manual grunt work every single week.
Claude Code/Codex can't do this in multiplayer mode. Every person rebuilds the same skill from scratch in their own way.
Deel built Akai to:
1. Understand backend operations edge cases (it had to work for our 7000 person team first)
2. Collaborative across 1000s of employees
3. Self-Learning from millions of runs
4. Optimises cost and gets cheaper every run
We're so confident that we're announcing an Automation Guarantee:
If our engineers can't automate a thousand of hours of work in your first 30 days, you get a full refund.
Book a demo: https://t.co/VESnztdck3 if you're an exec at a company with hundreds of employees