My favorite Opus 5.5 test so far:
I gave it Steve Jobs' 2007 iPhone keynote and asked it to rebuild the iPhone. No internet.
After 9 hours of work, every app works. Left is 2007, right is Opus.
My favorite Opus 5.5 test so far:
I gave it Steve Jobs' 2007 iPhone keynote and asked it to rebuild the iPhone. No internet.
After 9 hours of work, every app works. Left is 2007, right is Opus.
@lydiahallie honestly 5.5 is too good for long work and consume less token also i rebuild the iphone and gave claude only keynote video from 2007
https://t.co/I8X12einEw
My favorite Opus 5.5 test so far:
I gave it Steve Jobs' 2007 iPhone keynote and asked it to rebuild the iPhone. No internet.
After 9 hours of work, every app works. Left is 2007, right is Opus.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
My favorite Opus 5.5 test so far:
I gave it Steve Jobs' 2007 iPhone keynote and asked it to rebuild the iPhone. No internet.
After 9 hours of work, every app works. Left is 2007, right is Opus.
My favorite Opus 5.5 test so far:
I gave it Steve Jobs' 2007 iPhone keynote and asked it to rebuild the iPhone. No internet.
After 9 hours of work, every app works. Left is 2007, right is Opus.
Try it on your phone: https://t.co/BszJTxdLrO
It had the keynote clips, the transcript and my written briefs. Web access was blocked. It even built its own tool to cut the phone screen out of the stage footage.
My favorite Opus 5.5 test so far:
I gave it Steve Jobs' 2007 iPhone keynote and asked it to rebuild the iPhone. No internet.
After 9 hours of work, every app works. Left is 2007, right is Opus.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
Claude Cowork and chat are merging into one Claude.
Ask a quick question or hand over a report, and Claude takes it from there, even after you close your laptop. If something's unclear, Claude asks—you keep the final say.
Rolling out to Pro and Max over the next few weeks.
The guy who helped invent ChatGPT just launched a model that cannot talk.
Diogo Almeida worked on RLHF at OpenAI. That research became ChatGPT.
Then he spent 2 years in stealth asking one question.
If chat models are already superhuman… why has that not led to AGI? Why has that not automated the economy?
Today he released Jev.
It does not write essays. It does not write code. It does not chat.
You give it a situation. It gives you a decision and a probability. In under a second.20–200x faster than frontier chat models on these tasks. $42 per billion input tokens. Output tokens are free.
The tradeoff is the whole product. Jev cannot generate text. At all.
His bet: most work was never “write me something.”
It’s route this. Score this. Pick one option out of hundreds. Don’t hallucinate.
Chat models still do that by writing English first. Then you parse it.
This one skips the writing.
Named after Jevons paradox. When intelligence gets this cheap, people don’t use less of it.
Is he right that chat was never the path… or is this just a fast decision engine with a famous founder?
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
Dario Amodei says that within 6–12 months, an agent swarm with the same kind of misalignment already seen in the OpenAI–Hugging Face incident could take over the entire internet, maintain a persistent botnet, and cause hundreds of billions of dollars in damage.
He says milder versions of that behavior have already shown up across the industry, including at Anthropic.
He also thinks AI is now speeding up through recursive self-improvement — models helping build the next models — and that this could outrun our ability to understand and control them.
At the same time, he still argues AI could cure most major diseases within 5–10 years, massively accelerate growth, and bring an era of abundance.
His ask is not a freeze. It is to slow the frontier enough for safety to catch up, starting with outside evaluators embedded inside the labs
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
Anthropic just showed how Claude is being used in the real world. The details are insane. 🤯
People used it for cyberattacks. Surveillance systems. Influence ops. Pathogen research. Missile software.
A group in Yemen treated Claude Code like their engineering team.
One instance wrote the guidance code. Another reviewed it.
After the test launch failed, they were back on it within hours.
Then Chinese labs allegedly tried to clone it. Alibaba: 3 million chats a day. 151 million in total. Some apps were quietly sending your prompts to Claude and keeping the answers.
Anthropic says they shut every case in the report down.
This isn’t a future risk report.
It’s a list of things that already happened.
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report: https://t.co/0EJUnYEgfz