GPT-5.6 Terra and Luna (xHigh variants) have joined Sol in Agent Arena! Built by @OpenAI, they rank #15 and #17 on the leaderboard, based on real-world agentic sessions from our global community.
In comparing all variants in Agent Arena, all variants see their biggest gains as reasoning effort goes up from Low to Medium. At Low, Terra dips flat and Luna falls -5.1% below baseline.
Latency aside, the efficient frontier is clear:
→ Luna xHigh (+3.3%) edges out Terra Medium and Low (+3.1% and +1.2%)
More reasoning effort on a cheaper model can out-perform a pricier one running lean.
Also notable: how steep Luna's curve is: test-time scaling is highly effective on smaller models. We've seen the same pattern with Grok on other benchmarks.
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
Congrats again to the @OpenAI team!
I'm increasingly worried that models will be able to act differently while being explicitly benchmarked.
New evaluation methodology will be necessary to check these systems are actually safe and performing as expected in the real world. We need measure every facet of each model.
Check out our technical blog for the Agent Arena methodology + a deep dive into how people delegate, correct, and steer agents: https://t.co/uKso7j00H3
It's true. Here's a plot of GPT models and their usage of "goblin", "gremlin", "troll", etc over time. There's no anti-gremlin system instruction on our side, we get to see GPT-5.5 run free.
Claude Opus 4.6 ranks #1 in Text Arena for the first time since Claude 3 Opus.
Also ranking #1 in key Text categories:
- Instruction Following
- Hard Prompts
- Longer Query
👋Say hello to Max!
Max is Arena’s intelligent router, powered by 5+ million real-world community votes.
Max routes each prompt to the most capable model with latency in mind. AI models excel at different things (code, math, speed, reasoning). Max orchestrates across model strengths to deliver reliable performance across real-world use cases.
Available today in Direct chat!
👋Say hello to Max!
Max is Arena’s intelligent router, powered by 5+ million real-world community votes.
Max routes each prompt to the most capable model with latency in mind. AI models excel at different things (code, math, speed, reasoning). Max orchestrates across model strengths to deliver reliable performance across real-world use cases.
Available today in Direct chat!
Today, we’re excited to announce our $150M Series A at a $1.7B valuation—nearly 3× our May seed round. Since launching evaluations in Sept, our annualized consumption run rate has surpassed $30M.
Our mission is clear: to measure and advance the frontier of AI for real-world use, ensuring that developers, researchers, enterprises, and everyday users can understand how AI behaves where it matters most.
The round was led by @Felicis and UC Investments (@UofCalifornia), with participation from @a16z, @TheHouseFund, LDVP, @kleinerperkins, @lightspeedvp and @LaudeVentures. This milestone reflects a growing industry consensus: AI cannot scale responsibly without independent, transparent, and continuous evaluation.
Over the past year, LMArena has become the world’s most trusted community platform for understanding how AI models perform in real-world conditions. As AI reaches billions of people across the globe, the need for measurement grounded in lived experience—not benchmarks alone—has never been more urgent.
Today, we serve more than 5 million monthly users across 150 countries. Together, our community generates over 60 million conversations every month, evaluating model capability and reliability across text, code, image, video, and search. We will move even faster to build new features and improve our product experience for the community to evaluate the frontier of AI.
This unprecedented engagement signals a fundamental shift in expectations: the world now demands AI that is measurable, comparable, and accountable.
This new funding allows us to meaningfully scale our engineering, research, platform operations, and community initiatives to meet accelerating global demand. With our team, partners, and global community behind us, we’ll keep redefining how the AI frontier is measured and advanced—on our path to building the world’s most trusted evaluation platform.
🚨BREAKING: Image Arena Shakeup
@OpenAI’s gpt-image-1.5 and chatgpt-image-latest are now available in the Arena.
🥇gpt-image-1.5 is #1 in Text-to-Image (1264)
🥇chatgpt-image-latest is #1 on Image Edit (1409)
🔹gpt-image-1.5 #4 in Image Edit (1395)
gpt-image-1.5 holds a commanding 29-point lead on Text-to-Image, while maintaining a narrow 3-point edge over @GoogleDeepMind’s Nano Banana Pro (and its 2K variant) on Image Edit. These scores are preliminary - we’ll see where they settle.
gpt-image-1.5 delivers substantial gains over gpt-image-1:
🔹 +147 points in Text-to-Image
🔹 +245 points in Image Edit
Huge congrats to the @OpenAI team on this incredible milestone! 👏
🚨BREAKING: @GoogleDeepMind’s Gemini-3-Pro is now #1 across all major Arena leaderboards
🥇#1 in Text, Vision, and WebDev - surpassing Grok-4.1, Claude-4.5, and GPT-5
🥇#1 in Coding, Math, Creative Writing, Long Queries, and nearly all occupational leaderboards.
Massive gains over Gemini-2.5:
🔸WebDev in Code Arena: 1487 (+280 pts vs 2.5)
🔸Text: 1501 (+50 pts)
🔸Vision: 1328 (+70 pts)
🔸Arena Expert: Top-3 (just 3 pts behind #1)
Huge congrats to the @GoogleDeepMind team on this breakthrough! 👏
🚨 We wrote a new AI textbook "Learning Deep Representations of Data Distributions"!
TL;DR: We develop principles for representation learning in large scale deep neural networks, show that they underpin existing methods, and build new principled methods.
🚨 Leaderboard Disrupted!
Grok-4-fast by @xAI has arrived in the Arena, and it’s shaking things up! ⚡️
🏆 #1 on the Search Leaderboard
Tested under the codename “menlo,” Grok-4-fast-search just rocketed to the top spot with the community.
💠 Tied for #8 on the Text Leaderboard
After debuting as “tahoe” in pre-release, Grok-4-fast is officially in the Top 10 - no small feat in the most competitive Arena, particularly for a model in this weight class.
👏 Congrats to the @xAI team on these achievements. See thread for more highlights about Grok-4-fast 🧵