Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
Anthropic cut 80%+ of Claude Code's system prompt for the Claude 5 gen models. Coding evals seems to be same.
Most of what we wrote into CLAUDE.md was scaffolding for older models. It's dead weight now.
Some things to change : https://t.co/vXMIZjTNJm
Every time I open LinkedIn or X, someone is telling me that reviewing code is a productivity killer.
For a side project a dozen people will run? Sure. For code that touches prod? I beg to differ, and I've been trying to articulate why for months.
Then I read 𝙒𝙝𝙮 𝙎𝙤𝙛𝙩𝙬𝙖𝙧𝙚 𝙁𝙖𝙘𝙩𝙤𝙧𝙞𝙚𝙨 𝙁𝙖𝙞𝙡 by Dex at HumanLayer, and it's the most honest thing I've read on AI coding all year.
The pitch for the "lights-off" software factory is seductive: agents build, agents review, agents test, nobody reads the diff. Ship more. You're the bottleneck.
Here's why it breaks, and it isn't a skill issue.
Coding models are trained on a reward that's basically binary: did the test you were asked to fix pass, and did you avoid breaking the others. That's it. Nothing in that loop punishes wrapping everything in try/catch, lazy type casts, or spreading one concept across eleven files.
Tests give feedback in seconds. 𝘽𝙖𝙙 𝙖𝙧𝙘𝙝𝙞𝙩𝙚𝙘𝙩𝙪𝙧𝙚 𝙗𝙞𝙡𝙡𝙨 𝙮𝙤𝙪 𝙞𝙣 𝙬𝙚𝙚𝙠𝙨 𝙖𝙣𝙙 𝙢𝙤𝙣𝙩𝙝𝙨. 𝙔𝙤𝙪 𝙘𝙖𝙣'𝙩 𝙗𝙖𝙘𝙠𝙥𝙧𝙤𝙥 𝙖𝙣 𝙞𝙣𝙘𝙞𝙙𝙚𝙣𝙩 𝙞𝙣 𝙈𝙖𝙧𝙘𝙝 𝙩𝙤 𝙩𝙝𝙚 𝙙𝙚𝙨𝙞𝙜𝙣 𝙙𝙚𝙘𝙞𝙨𝙞𝙤𝙣 𝙩𝙝𝙖𝙩 𝙘𝙖𝙪𝙨𝙚𝙙 𝙞𝙩 𝙞𝙣 𝙅𝙖𝙣𝙪𝙖𝙧𝙮.
So the models got dramatically better at solving one-off problems, and barely moved on keeping a codebase healthy. And there is no benchmark that even measures the second thing.
His team ran the lights-off experiment for real. It ended with a cofounder spending two weeks in VS Code replumbing patterns by hand.
What actually works is the boring stuff we knew before AI, just moved earlier:
→ Agree on the product problem and what success will read like
→ Sketch the architecture: services, contracts, schemas
→ Do program design, the underrated one. Types, method signatures, call-stack trees, file-tree diffs, before a line gets written
→ Build in vertical slices you can poke at, instead of letting an agent hand you 2,000 lines and a mystery
𝙏𝙝𝙞𝙧𝙩𝙮 𝙢𝙞𝙣𝙪𝙩𝙚𝙨 𝙤𝙛 𝙥𝙡𝙖𝙣𝙣𝙞𝙣𝙜 𝙨𝙖𝙫𝙚𝙨 𝙝𝙤𝙪𝙧𝙨 𝙤𝙛 𝙧𝙚𝙫𝙞𝙚𝙬. Every one of those is a decision you'd otherwise make during code review, at the most expensive possible moment to change your mind.
My favourite line from the piece: 𝘆𝗼𝘂 𝗱𝗼𝗻'𝘁 𝗵𝗮𝘃𝗲 𝘁𝗼𝗼 𝗺𝗮𝗻𝘆 𝗣𝗥𝘀, 𝘆𝗼𝘂 𝗵𝗮𝘃𝗲 𝘁𝗼𝗼 𝗺𝗮𝗻𝘆 𝗯𝗮𝗱 𝗣𝗥𝘀.
I'd rather be 2-3x faster and still able to sleep than chase a 10x number I'll pay back with interest during an outage.
Read the code.
https://t.co/9Chw5lCbHj
Do you know what you don't know? https://t.co/j1GvUeMoBA was my secret weapon when preparing for my Deel interview.
When leveling up, the hardest part isn't learning the material: it's figuring out what to actually learn.
Check out https://t.co/JQmyHGS3qr
I Asked Claude Opus 5 to Make Me Perlin noise flow field
Claude Opus 5 built: 20,000 particlesriding a Perlin noise flow field.
Every color, every curve of motion, every parameter was the model's choice.
The whole thing is one HTML file.
Real-time simulation, no rendering, no After Effects. Drag your finger across it and the flow bends around you.
Built with: Claude Opus 5 (Anthropic) Technique: Perlin noise flow field #claudeopus5 #claudeai #generativeart #oddlysatisfying
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.