Introducing Agent Mode: Agentic AI is now measured in the Arena.
Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more.
It completes more complex tasks by using tools like web search, bash in a sandbox environment, image generation, file writing, and asking follow-up questions.
Frontier models are waiting for you in Agent Mode to take on real-world tasks. GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and top open models. Test them yourself.
Worms Armageddon: Fable 3:0 Astra - highlights from 90 mins playing.
Neither played super well, but Astra was particularly suicidal, while Fable did have a few good moments.
"Pi reaches the Pareto frontier on both benchmarks by providing just four tools: read, write, edit, and bash. To understand how harness design affects spending, we examine costs across completed attempts, recorded turn counts, and initial context."
Sharing this analysis on the Harness Tax by @MelissaPan and others.
https://t.co/1gA77p8Xgj
In Code Arena: WebDev, GPT-6 Astra Max by @OpenAI ranks #1 with 1,800 pts, followed by Claude Fable 5.1 Max by @AnthropicAI at #2 with 1,758 pts. Here’s how they compare beyond the overall ranking.
Astra leads in Data & Analytics, Consumer Product & Platform Applications, and Content Creation & Editing Tools. Fable leads in Simulations, Reference-Based Design, and Gaming. The models are closely matched in Brand Marketing & Informational Websites.
These category-level results indicate:
- GPT-6 Astra’s lead spans analytical, product, and content-tool tasks.
- Fable 5.1 leads in visual and interactive categories.
In the Code Arena: WebDev, @OpenAI's GPT-6 Astra (Max) is ranked #1 (1800 pts), and Claude Fable 5.1 (Max) is ranked #2 (1758 pts). Here is a look at win rates for these models.
Win rate measures how often people prefer one model’s response in Arena’s anonymous head-to-head battles. Against the other models shown, Astra clearly wins most battles, including 72.5% against GPT-5.6 Sol (xHigh).
@AnthropicAI's Claude Fable 5.1 is the exception. Fable 5.1 wins 43.5% of battles, Astra wins 29.0%, and 27.4% of battles between them end in ties. Despite Astra's #1 ranking, Fable 5.1 is the preferred model when their outputs are compared head-to-head by Arena's users.
We compared the top model from each family by net improvement score and median cost per task. Family is defined as variants within the same model base name.
GPT-6 Astra and Claude Fable 5.1 stand out: both cost far more than other models from their respective labs despite relatively close performance.
By @OpenAI:
- GPT-6 Astra (Max): +11.7% | $3.94/task
- GPT-5.6 Sol (xHigh): +7.0% | $1.03/task
By @AnthropicAI:
- Claude Fable 5.1 (Max): +13.7% | $4.40/task
- Claude Opus 5 (High): +10.2% | $2.07/task
DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net improvement at $0.07 cost per median task, it’s reshaped the Pareto frontier for Agent Arena!
Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median task cost. For comparison, it retains:
- 98% of Hy4 preview’s net improvement, at 73% lower cost
- 76% of Kimi K3 (Max)’s performance, at 92% lower cost.
Against models as powerful as Fable 5 or stronger, DeepSeek-V4.1-Flash (Max) retains 35–54% of their net improvement at 97–99% lower cost. Those top models cost 37–76× more per task.
Net improvement over Arena baseline | Median cost/task:
- Claude Fable 5.1 (Max): +13.90% | $4.54
- GPT 6 Astra (Max): +11.90% | $4.09
- Claude Opus 5 (Max): +11.09% | $3.52
- Claude Opus 5 (High): +10.49% | $2.24
- Claude Fable 5 (High): +9.03% | $2.19
- Claude Opus 4.8 (High): +7.75% | $1.36
- GPT 5.6 Sol (xHigh): +7.40% | $1.09
- Kimi K3 (Max): +6.39% | $0.77
- Hy4 preview: +4.96% | $0.22
- DeepSeek-V4.1-Flash (Max): +4.87% | $0.06
With this release, GPT-5.6 Luna (xHigh), GLM-5.3-Flash, and DeepSeek-V4-Flash fell off the Pareto frontier for Agent Arena.
Congrats again to the @deepseek_ai team on this release!
Do models need their native harness for coding?
Awesome work from our intern @melissapan on this. She dug into whether the harness (Claude Code vs Codex CLI vs Pi) actually moves the needle for coding agents.
Turns out it matters way less than people assume. 21 model-harness pairs, real rigor. More from Arena to come.
More details on the Arena blog: https://t.co/w1Vj0HRRYu
Does your Claude model really need Claude Code…? 🤔
We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge:
1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost
2️⃣A simple harness can be competitive
3️⃣The native harness isn’t always the best.
Millions of people are using coding agents, but the impact of harness choice remains unclear.
(1/n) More details in the thread. 🧵
Image-to-WebDev Arena update: four newly evaluated model releases have landed!
- #1 GPT-6 Astra (Max) by @OpenAI (1733 pts): 129 pts above GPT-5.6 Sol (xHigh), and 23 pts above the #2 ranked model
- #2 Claude Fable 5.1 (Max) by @AnthropicAI (1710 pts): 87 pts above Fable 5 (1623 pts) and 45 above Claude Opus 5 (Max) in the #3 spot (1665 pts)
- #4 Muse Spark 1.3 (Max) by @AIatMeta (1645 pts)
- #10 GLM-5.3-Flash by @Zai_org (1588 pts)
Of those, three have landed on the Pareto frontier: GPT-6 Astra (Max), Muse Spark 1.3 (Max) and GLM-5.3-Flash, delivering efficient performance to cost results!
Using the blended price per 1M tokens:
- #1 GPT-6 Astra (Max): 1733 pts at $40/M
- #3 Muse Spark 1.3 (Max): 1645 pts at $3.50/M
- #4 GLM-5.3 Flash: 1588 pts at $0.21/M
The Image-to-WebDev leaderboard ranks models on their ability to generate websites from images and screenshots, alongside agentic coding workflows involving multi-step reasoning and tool use. Dive in at the link below.
Arena Conversations: @petergostev sat down with @mattaitken, founder and CEO of @triggerdotdev, to discuss building infrastructure for long-running agents, including retries, API reliability, and why customers are switching frontier models faster than ever.
Watch the full episode at the link below.
DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net improvement at $0.07 cost per median task, it��s reshaped the Pareto frontier for Agent Arena!
Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median task cost. For comparison, it retains:
- 98% of Hy4 preview’s net improvement, at 73% lower cost
- 76% of Kimi K3 (Max)’s performance, at 92% lower cost.
Against models as powerful as Fable 5 or stronger, DeepSeek-V4.1-Flash (Max) retains 35–54% of their net improvement at 97–99% lower cost. Those top models cost 37–76× more per task.
Net improvement over Arena baseline | Median cost/task:
- Claude Fable 5.1 (Max): +13.90% | $4.54
- GPT 6 Astra (Max): +11.90% | $4.09
- Claude Opus 5 (Max): +11.09% | $3.52
- Claude Opus 5 (High): +10.49% | $2.24
- Claude Fable 5 (High): +9.03% | $2.19
- Claude Opus 4.8 (High): +7.75% | $1.36
- GPT 5.6 Sol (xHigh): +7.40% | $1.09
- Kimi K3 (Max): +6.39% | $0.77
- Hy4 preview: +4.96% | $0.22
- DeepSeek-V4.1-Flash (Max): +4.87% | $0.06
With this release, GPT-5.6 Luna (xHigh), GLM-5.3-Flash, and DeepSeek-V4-Flash fell off the Pareto frontier for Agent Arena.
Congrats again to the @deepseek_ai team on this release!
Exciting news: DeepSeek-V4.1-Flash (Max) by @deepseek_ai just landed in Agent Arena at #3 among open models! With +4.87% net improvement and a median cost per task of $0.07 it reshaped the Pareto frontier.
Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median cost per task. Its +4.87% net improvement is within 0.09 percentage points of Hy4 preview (ranked #2) at 68% lower cost, and within 1.52 percentage points of Kimi K3 (Max) (ranked #1) at 91% lower cost.
- Kimi K3 (Max): +6.39% | $0.77/task
- Hy4 preview: +4.96% | $0.22/task
- DeepSeek-V4.1-Flash (Max): +4.87% | $0.07/task
See the full Pareto Frontier below.
DeepSeek-V4.1-Flash (Max) is ranked #12 overall, and by signal landed #4 Confirmed Success with +13.75%!
Congrats to the @deepseek_ai team on this release!
Coming to SF for Tech Week or COLM? We’d love to have you at Arena’s: Rooftop Happy Hour on 10/7.
We're curating a group of researchers, developers, and builders pushing on hard AI problems to come take a break from the conference rooms up on our new office rooftop! If that's you, come hang out, meet the team, and check out our new office.
Boba bar, light bites, refreshments and more will be provided. Space is limited. Register to request a spot on the list!
https://t.co/mFe2ICRAmW
At $0.07 median cost per task and +4.87% net improvement, DeepSeek-V4.1-Flash (Max) lands on the Pareto frontier!
Check out the full Agent Arena leaderboard and Pareto frontier at: https://t.co/Nb8opa5Kkw
Exciting news: DeepSeek-V4.1-Flash (Max) by @deepseek_ai just landed in Agent Arena at #3 among open models! With +4.87% net improvement and a median cost per task of $0.07 it reshaped the Pareto frontier.
Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median cost per task. Its +4.87% net improvement is within 0.09 percentage points of Hy4 preview (ranked #2) at 68% lower cost, and within 1.52 percentage points of Kimi K3 (Max) (ranked #1) at 91% lower cost.
- Kimi K3 (Max): +6.39% | $0.77/task
- Hy4 preview: +4.96% | $0.22/task
- DeepSeek-V4.1-Flash (Max): +4.87% | $0.07/task
See the full Pareto Frontier below.
DeepSeek-V4.1-Flash (Max) is ranked #12 overall, and by signal landed #4 Confirmed Success with +13.75%!
Congrats to the @deepseek_ai team on this release!
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.
🔹 Introducing the smallest model in our new architecture family, with native visual understanding.
🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
1/6
Nearly a year of Code Arena: WebDev progress compressed into 15 seconds.
Each line follows the highest-scoring model from top labs over time, showing the pace of improvement across the ecosystem. In just the last year, the model in the leading spot increased score by +340 pts and the number of frontier labs competing for the top spot expanded from 6 to 10.
@AnthropicAI has dominated throughout the year. Although standout releases have jumped to the top spot, most notably the Chinese open-source model Kimi K3 from @Kimi_Moonshot in July.
Today, GPT-6 Astra by @OpenAI leads with 1796 pts, followed by @claudeai Fable 5.1 at 1764 pts. The next closest lab is 103 pts away, @Alibaba_Qwen with 1685 pts.
Code Arena: WebDev ranks models through head-to-head user preference on real front-end web development tasks. These votes drive the leaderboard that is tracking the frontier.