Introducing Agent Mode: Agentic AI is now measured in the Arena.
Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more.
It completes more complex tasks by using tools like web search, bash in a sandbox environment, image generation, file writing, and asking follow-up questions.
Frontier models are waiting for you in Agent Mode to take on real-world tasks. GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and top open models. Test them yourself.
Agent Arena tracks the number of tokens a model takes to complete real-world tasks.
We see the performance of Opus-series models has improved significantly (Opus 4.7 to 4.8 to 5), and at the same time token usage has also increased substantially (~8.5k for Opus 4.7 up to ~21k for Opus 5), when counting both reasoning and output tokens.
The opposite trend appears for top GPT models. Between GPT-5.5 and GPT-5.6-Sol, we see the token usage declined from ~10k to ~8k, despite a notable performance gain of ~1.5pp of Net Improvement.
Exciting news: Muse Spark 1.2 (xHigh) by @AIatMeta is #4 in the Text Arena (1498 pts), and has reshaped the Pareto frontier!
It is priced at $1.25/$4.25 per MToken.
Congrats again to the @AIatMeta team on this release!
Introducing Muse Code (beta), a terminal coding agent built for long-horizon software engineering, powered by our new Muse Spark 1.2 model.
Muse Code plans, implements, and validates complex, multi-file changes across large repositories with persistent sub-agents that solve difficult problems faster, more accurately, and with less intervention.
🧵👇
Muse Spark 1.2 (xHigh) by @AIatMeta is #14 in the Code Arena: WebDev, with 1,545 pts!
This is an improvement from Muse Spark 1.1 at #18. See its biggest gains by category in the post below.
Congrats to the @AIatMeta team on this release!
Introducing Muse Code (beta), a terminal coding agent built for long-horizon software engineering, powered by our new Muse Spark 1.2 model.
Muse Code plans, implements, and validates complex, multi-file changes across large repositories with persistent sub-agents that solve difficult problems faster, more accurately, and with less intervention.
🧵👇
Muse Spark 1.2’s biggest gains by category compared to Muse Spark 1.1 are:
- HTML #24 to #8 (16 spots)
- Gaming #23 to #13 (10 spots)
- Frontend #19 to #13 (6 spots)
- Data & Analytics #13 to #8 (5 spots)
Today, we are announcing Arena's factuality leaderboards, which evaluate the objective performance of AI models in reality and guard against hallucinations.
Arena's chat leaderboard has historically focused on human preference. Although our style control methodology adjusts for the effect of emojis, response length, and other stylistic factors, we wanted a more direct way of measuring objective signals of chatbot utility and anti-hallucination.
Hence, we developed the factuality leaderboards, which speak to the fraction of correct factual claims that a model makes. We basically extract all factual claims and cross-reference them against the internet to ensure they are based in evidence. Models that are supporting their claims with stronger evidence get a better score, and this gives a strong basis for objective performance beyond preference alone.
The big winners, as you can see on this plot, are OpenAI and Grok. Notably, Claude models, as well as Ernie and Muse-Spark-1.1, decrease in ranking.
We're so excited to share this with you! Hope you enjoy!
Muse Spark 1.2 by @AIatMeta is in the Agent Arena!
Bring your toughest prompts to power the leaderboards.
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
In addition to Agent Arena, Muse Spark 1.2 is available in Text, Vision and Code Arena: Frontend.
Introducing Muse Code (beta), a terminal coding agent built for long-horizon software engineering, powered by our new Muse Spark 1.2 model.
Muse Code plans, implements, and validates complex, multi-file changes across large repositories with persistent sub-agents that solve difficult problems faster, more accurately, and with less intervention.
🧵👇
AI Capability Lead at Arena @petergostev told us that his way to measure being in the singularity is how much research is actually being recursive:
"My impression now, and I've been trying to run some math problems myself, my sense is that they're still bad researchers".
"They don't really understand like what problems to focus on. They go down to the way too low level of detail. They don't step back".
"They don't learn from their mistakes and so on. So it's kind of makes them a bad researcher for now".
DeepSeek-V4-Flash (High) by @deepseek_ai has reshaped the cost-performance Pareto frontier in Agent Arena, with a $0.024 median cost per task!
It lands to the right of GPT-5.6 Luna (xHigh) which has a $0.026 median cost per task, and to the left of DeepSeek-V4-Pro (Thinking) at $0.020. It's the model that delivers a positive net improvement at the lowest price point on the chart.
Price per task is based on real-world Agent Mode usage. We calculate each task’s cost from tokens consumed and the model’s pricing, including cache hits and misses.
Dive into the Fullstack Leaderboard for more details at https://t.co/irGHxLSOOZ and learn more about fullstack capabilities at: https://t.co/9UrqAQgqiU
Code Arena just leveled up with fullstack capabilities 🚀
Introducing the new Fullstack Code Arena. We’re moving beyond frontend prototypes to fullstack development complete with databases, API keys, and fast deployments. Build, iterate, and ship real-world software — all in one place.
Models now act as agents in the Code Arena, using structured tool calls to plan, execute, and refine in real time with real world tasks.
Read more about it in the thread 🧵
Big news: Opus 5 (Max) is now #1 in the Fullstack Code Arena with 1,699 points!
The Fullstack Leaderboard shows overall rankings across AI models on full-stack web development tasks: multi-step reasoning, tool use, and end-to-end app generation.
Congrats again to the @AnthropicAI team on Opus 5 (Max)!
Exciting news: Claude Opus 5 with Max reasoning is #1 in the Frontend Code Arena and Text Arena with factuality on!
Claude Opus 5 with default reasoning high is also very strong landing #3 in Frontend Code Arena, right behind Kimi K3 - and #2 in Text Arena (factuality on).
This is real world data that @AnthropicAI's newest model holds up on real world tasks: agentic web coding, document reasoning, and general chat capability.
Claude Opus 5 Max’s score is still preliminary. We’ll continue to see how scores converge and share updates.
Congrats to @AnthropicAI on the SOTA release!