this is quietly huge
Anthropic just made Haiku 5.5 10x cheaper than Haiku 4.5, and told you to run it as a subagent under Opus 5.5 and Sonnet 5.5
most people send every agent task to one big model. you pay top price for file reads, lookups and summaries
the setup splits the stack:
> lead: Opus 5.5 or Sonnet 5.5 plans and makes the hard calls
> subagent: Haiku 5.5 does summaries, compaction, db queries and classification
> hard agent coding still goes to Sonnet 5.5 or Opus 5.5
the numbers:
> $0.10 per 1M input tokens. Haiku 4.5 was $1
> 20x cheaper than Sonnet 5.5 per token, for prompts up to 100K
> Devin Fusion holds 66.2 on FrontierCode with Haiku 5.5 as the sidekick (Cognition)
Opus thinks. Haiku runs.
you stop paying your smartest model to read files. you let the small one run
@Argona0x the 30% vs 8% gap makes sense.
every new message makes the bot re-read the whole chat .so message 50 costs way more than message 1.
any simple rule for when to cut a chat and start fresh?
1 person. A day job. Anthropic's AI hunting bugs for 31 days.
It found 171. 4 out of 5 got thrown out.
The win came from the setup, not a smarter model:
> AI #1 hunts. Its only job is to find suspects
> AI #2 tries to kill each suspect. Only survivors move up
> another model checks again, so one model's blind spots don't pass
> a human makes the final call. In the paper, AI found 0 bugs fully alone
Result: about 135 suspects dead. 4 became official security bugs (CVEs).
Same idea, another study (Monash University):
just asking AI "is this bug real?" caught 36.4% of fake alarms.
an AI agent that checks the code caught 95.5%.
I animated my version: Jev flags, Opus 5.5 checks, my spider pulls only what survives.
Try it today: when AI finds a bug, open a new chat and ask it to prove the bug is fake.
Finding bugs is cheap. Proving them is the job.
Your AI found bugs today?
Until something tries to kill them, they are guesses.
5 AI agents. 1 task. 60 seconds.
only 1 finished.
agents from OpenAI, Anthropic and two made-up ones.
(a skit, not a benchmark. you'll argue anyway.)
JEV was done in 4 seconds. the website was one button: YES. disqualified.
DOTS never left the start line. still "watching".
ASTRA took the lead. at 82% it clicked "delete project".
my money was on SONNET. it hit 97%. still adding dark mode.
OPUS sat at 0% for 50 seconds. just thinking.
then it built the whole site in the last 10.
the winner spent 50 of the 60 seconds thinking.
who do you bet on: the agent that answers in 4 seconds, or the one that thinks for 50?
@noisyb0y1 breakouts are where slippage is worst: everyone hits the same level at once. with a -1% stop, even 0.2% slippage each way eats 40% of your risk
@Argona0x The actual trick from Anthropic's Opus 5.5 guide is missing here: a time budget.
Add a line like "elapsed 340s / 1200s" to every message. The lead agent then keeps more helpers running in parallel.
And the guide says quality stayed "comparable", not better.
@marfinxx read the paper. no Opus 5.5 or GPT-6 Astra in it. they tested 1.5B-8B open models like Qwen2.5 and LLaMA-3.1. The 58% on MATH500 is a 1.5B model