Across our benchmarks, the model sets a new standard.
It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.
Introducing Gemini Robotics 2: the next era of truly adaptable robots from @GoogleDeepMind 🤖
Powered by three new models, it brings whole-body control, fine dexterity, and teamwork to complex tasks — making robots more helpful in the real world. https://t.co/aey7mLztYc
Claude Opus 5 is narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at 26% lower Cost per Task
We supported @AnthropicAI to evaluate Claude Opus 5 ahead of release: it sets the highest GDPval-AA v2 and AA-Briefcase scores so far. Opus 5 (max) scores 61 on the Artificial Analysis Intelligence Index, effectively tied with Claude Fable 5 (max, 60), and ahead of GPT-5.6 Sol (max, 59), Kimi K3 (57), and Claude Opus 4.8 (max, 56)
Key takeaways:
➤ New leader in agentic knowledge work: Claude Opus 5 (max) scores 1861 Elo on GDPval-AA v2, >100 points ahead of Claude Fable 5 and GPT-5.6 Sol (max). On AA-Briefcase, our proprietary agentic knowledge work benchmark, it scores 1720 Elo, +146 ahead of Fable 5. These benchmarks test the ability of models to produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup
➤ Joint first place on the Coding Agent Index: Claude Opus 5 (xhigh) with Claude Code leads the Artificial Analysis Coding Index, including the highest score on SWE-Atlas-QnA
➤ Frontier intelligence with reduced cost: Claude Opus 5 (max) costs $2.03 on average per Intelligence Index task, below Claude Fable 5 (with fallback) at $2.75, but still above Claude Opus 4.8 (max) at $1.80 and Claude Sonnet 5 (max) at $1.53. However, at high and xhigh reasoning efforts Opus 5 can outperform both Opus 4.8 and Claude Sonnet 5 at a lower cost per task
➤ Frontier agentic terminal use: 89% on Terminal-Bench v2.1 at max effort, roughly in line with the leader, GPT-5.6 Sol (xhigh)
➤ Outperformance on scientific reasoning: Along with leading agentic performance, Claude Opus 5 scores 53% on Humanity’s Last Exam in line with Fable 5; on CritPt, a frontier physics evaluation developed by Argonne and UIUC researchers, it also matches Fable 5 but sits behind GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra
➤ Factual knowledge still lags Fable 5: As expected from the models’ size classes, Opus 5 still has lower factual knowledge on AA-Omniscience than Fable 5. It improves +7 points on AA-Omniscience Accuracy over Opus 4.8, but answers more often when uncertain - its hallucination rate rises +14 points to 50%
➤ Improving efficiency, but only on the Intelligence vs. Cost per Task Pareto frontier at high Intelligence levels: Opus 5 outperforms Fable 5 at lower cost, but at lower effort levels it sits just behind the GPT-5.6 family on the Intelligence vs. Cost per Task frontier
Other model details:
➤ Context window: 1 million tokens (equivalent to Opus 4.8)
➤ Pricing: As with recent Opus launches, tokens cost $5/$25 per million tokens of input/output; cache pricing remains at a 25% premium for cache writes ($6.25 per million tokens) with 5-minute time to live, and 90% discount for cache hits ($0.50 per million tokens)
➤ Five effort settings (low, medium, high, xhigh, max), and support for server-side fallback as with Fable 5. Intelligence Index evaluations were run with Opus 4.8 fallback enabled
Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open weights model
Key results:
➤ Strong agentic task performance: @Kimi_Moonshot's Kimi K3 reaches an Elo rating of 1668 on GDPval v2. This is a marked improvement over K2.6’s 1190, surpassing GLM-5.2 (1514), GPT-5.5 (1494), and Claude Opus 4.8 (1600). However, it still lags behind Claude Fable 5 (1760). Kimi K3 also scores an impressive 53% and takes the #1 position on AutomationBench-AA, our implementation of Zapier’s Agentic SaaS workflow evaluation.
➤ Second-highest performance on AA-Briefcase (agentic knowledge work): On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5. It is well-rounded: its rubric scoring and analytical quality almost reach Claude Fable 5’s scores, while GPT-5.6 Sol continues to outperform other leading models on presentation quality.
➤ Set to lead open weights models once weights are released: Moonshot AI has not yet released the weights but expressed plans to do so. Once available, Kimi K3 would clearly lead other open weights models including GLM-5.2 (51) and DeepSeek v4 Pro (44). However, at 2.8T parameters, it is significantly larger than its open weights peers (eg. GLM-5.2 at 753B params and DeepSeek V4 Pro at 1.6T), as well as the Kimi K2 to K2.6 models (1T params).
➤ Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers: Moonshot AI’s pricing for K3 is significantly higher than their K2 pricing (K3’s output token price is $15/1M tokens while K2.6 was $4). This positions the model as cheaper on a cost per task basis than Opus 4.8, similar to GPT-5.6 Sol ($1.04) and more expensive than open weights peers, GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04)
➤ Improved token efficiency alongside higher intelligence: Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6. The new model used approximately 132M output tokens to complete all nine evaluations, compared to approximately 166M for K2.6, while achieving higher scores.
➤ Native multimodal capabilities: Kimi K3, like K2.6, is released with native image and text multimodal input. If weights are released, this will position Kimi K3 as one of the leading open weights models with multimodal input capabilities
Other model details:
Context window: 1M
Size: 2.8T total parameters
Pricing: The first-party API is priced at $3.00/$15.00 per 1M input/output tokens, with cached input discounted 90% to $0.30 per 1M tokens.
Modality: Native multimodal input supports text and images, and the model remains text-only for output.
Accessibility: Accessible at launch through Moonshot’s first party API. Model weights are not yet released but Moonshot AI has expressed plans to do so.
Introducing ChatGPT Work, a new agent in ChatGPT powered by Codex and GPT-5.6.
It can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work.
It’s a whole new way to get work done.
GPT-Live is now fully rolled out to all ChatGPT users on Go, Plus, and Pro plans. Free user rollout is in progress.
Update to the latest version of the ChatGPT app on iOS or Android to try it out.
@OpenAI Interesting how many still want back the 4o. 4o was as the latest model powerful, but in comparison with the latest, it is a legacy one. You can modify your gpt in settings for fine tuning if required.
4o is not away. It goes on as a part of the latest upgraded models.
Announcing Grok 4.5, our first model trained specifically for coding and agents. It was trained with Cursor and offers frontier intelligence at leading speeds and cost efficiency.
https://t.co/i8HpU7w64k
GPT-5.2 just overtook Claude Opus 4.5 to achieve the highest score in GDPval-AA, a benchmark that focuses on performance in real-world economically valuable tasks
However, GPT-5.2 is also the most expensive model to run GDPval-AA: GPT-5.2 cost $620, compared to Claude Opus 4.5’s $608 and GPT-5.1’s $88.
This was driven by @OpenAI's GPT-5.2 using >6x more tokens than GPT-5.1 (250M compared to 40M), and OpenAI raising prices by 40% ($14/$1.75 per million input/output tokens compared to $1.25/$10).
GDPval-AA uses our agentic harness, Stirrup, to run models on OpenAI's GDPval dataset, and measures their performance using an AI-based grading pipeline.
Full set of Artificial Analysis Intelligence Index benchmarks are in progress and we will be sharing a full update when complete.