We asked ten Claude Opus 5.5 agents to devise a faster shortest-path algorithm and prove it in Lean. Within 15 hours, they produced C-HD: a formally verified improvement over the published bounds.
We want the OpenAI API to feature the best model at every price point and to be the best at every modality (text, code, image, video, etc).
And then we want you all to come up with great ideas and build them and to get to be happy users. The best ideas will come from you all.
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe.
GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale.
We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
We rebuilt customer support around Grok Bot to scale our operations without adding headcount.
It works autonomously to respond to customers, resolve tickets, and manage the queue.
https://t.co/Vyw224m8Dh
As part of our efforts to pace the frontier, we’re committed to supporting independent assessments with deep levels of access across training, evaluation, and deployment.
That access should enable third party assessors to challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards.
We’re outlining four priority areas for deeper assessment, alongside principles for rigorous, secure, and independent work:
https://t.co/z2b15zsqzO
On RSI Index, Opus 5.5 is the top model on four of the five research tasks. On LM Training, where the model trains its own small language model within a 24-hour budget, its result beat the published human baseline. This shifts our timeline for full RSI from August 2027 to July 2027
On Terminal Bench 4 it scores 61.6% vs. 45.5% for Opus 5. On Proof Bench it ties for #1 at 100%, and is the cheapest run in the top 10 at $0.92/problem.
Claude Opus 5.5 takes #1 on RSI Index and is the first model to beat the published reference on LM Training under our protocol, marking a major step forward for long-horizon agentic work.
Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that.
Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads.
Opus 5.5 is our first model since we called for pacing the frontier. As with previous models, it was tested by external evaluators before release, including METR and Frontier Design.
On our most comprehensive alignment test, it achieves the strongest score to date.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.