For the first time, China has taken the lead over the US in Frontend Code Arena with the launch of Kimi-K3 by @Kimi_Moonshot.
The last time a Chinese model came close was in early 2025, with DeepSeek-R1.
Fastest growing company I ever worked at, @uber included. It's been a blast evaluating frontier models from the world's leading labs. If you're an engineer or designer looking for a new role, let's chat.
Arena has crossed $100M in annualized revenue run rate, eight months after launching our evaluation product.
With our recent release of Agent Mode, millions of users on Arena are doing real work with agents, from coding to document analysis, in long-running, multi-turn sessions with hundreds of tool calls. Arena now evaluates objective criteria like task completion rates, hallucination rates, and more, far beyond our original human preference voting model. This expansion has taken us from a student project at Berkeley to one of the fastest growing companies in history. Go Bears! 🐻
Our core thesis is simple: to align AI with human values, we must directly measure its impact on people in the real world. Today's milestone is proof that Arena’s platform is the de-facto standard for post-deployment evaluation of AI.
Today we launched Agent Mode and landed in the The New York Times 🤘 Super proud of the team for all it took to get here.
Try it out on @arena – and check out the agent leaderboard to see which models came out on top.
https://t.co/5dqwhxWRXo
Introducing Agent Mode: Agentic AI is now measured in the Arena.
Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more.
It completes more complex tasks by using tools like web search, bash in a sandbox environment, image generation, file writing, and asking follow-up questions.
Frontier models are waiting for you in Agent Mode to take on real-world tasks. GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and top open models. Test them yourself.
Max, Arena's model router powered by 5M+ community votes, is now multimodal.
Starting today, Max is the default in Direct chat across every modality: search, vision, image generation, image editing, and front-end coding with the same latency-controlled performance as the original router for text.
Learn more about Multimodal Max in thread.
Exciting news - GPT-Image-2 by @OpenAI has claimed the #1 spot across all Image Arena leaderboards!
A clean sweep with a record-breaking +242 point lead in Text-to-Image - the largest gap we’ve seen to date.
- #1 Text-to-Image (1512), +242 over #2 (Nano-banana-2 with web-search aka gemini-3.1-flash-image)
- #1 Single-Image Edit (1513), +125 over #2 (Nano-banana-pro aka gemini-3-pro-image)
- #1 Multi-Image Edit (1464), +90 over #2 (Nano-banana-2)
No model has dominated Image Arena with margins this wide.
Huge congratulations to @OpenAI on this major breakthrough in image generation! More performance breakdowns by category in the thread below.
Meta is back in the Arena!
Muse Spark debuts as a top frontier model across both Text and Vision:
- Text Arena: #3 tied with Gemini-3.1-Pro and Claude-Opus-4.6
- Vision Arena: #2 tied with Claude-Opus-4.6
This marks Meta’s first major release since early 2025.
Highlights:
- #4 Hard Prompts, #6 Coding, #9 Creative Writing, #10 Instruction Following, #27 Expert
- #3 tied for Business, Management, & Financial Ops, #7 Legal & Government, #12 Writing & Literature
Meta is back at the frontier. Huge congrats to @AIatMeta on this incredible milestone!
The key to building a leaderboard that can’t be gamed is starting with the right structural foundations: neutrality and rigorous methodology.
Arena scores are powered by a constant stream of real-world prompts from millions of users across the globe comparing responses from the latest frontier AI models for both expert workflows and everyday tasks.
Those interactions continuously shape the rankings, and unlike static benchmarks, the leaderboards reflect how models actually perform in the wild.
Our co-founder and CEO @ml_angelopoulos recently discussed this approach and the broader vision behind Arena on Equity podcast.
Thanks to @TechCrunch, @RebeccaBellan and the @EquityPod team for the conversation.
See the full interview on YouTube. Link in thread.
👋Say hello to Max!
Max is Arena’s intelligent router, powered by 5+ million real-world community votes.
Max routes each prompt to the most capable model with latency in mind. AI models excel at different things (code, math, speed, reasoning). Max orchestrates across model strengths to deliver reliable performance across real-world use cases.
Available today in Direct chat!
LMArena is now Arena.
A name that takes us back to our roots with a powerful mission: to measure and advance the frontier of AI for real-world use.
We have grown from a small PhD research project to a platform powered by a global community of millions. This rebrand has been shaped by the people who use it.
👇 Take a look inside the rebrand.
🚨BIG NEWS: 🎬 Video Arena is now live on the web!
Test out Veo 3.1, Sora 2, Seedance v1.5 Pro, Kling 2.6 Pro, Wan 2.5 & more.
What started last summer as a small Discord bot experiment has grown into a rigorous way to measure and understand how frontier video models perform with real-world use. Thank you to our wonderful community for all the feedback!
Today, we’re opening up access by making it available on the web.
🎥 Generate videos with 15 different frontier AI models and compare them head-to-head.
📊 Vote for the best output to power the leaderboards.
LMArena has raised $150M+ at a valuation of $1.7B+ 💪🏼
In the past 7 months, @arena has:
Grown our userbase 25x. 35M+ unique users.
Grown our revenue from 0 to >>$30M+ ARR in 4 months. Our products help labs and enterprises measure the real utility of AI and understand their strengths and weaknesses for real users.
Grown our team to 40+ world-class experts in machine learning, product engineering, design, marketing, BD, and more.
We are looking for world-class ML scientists, engineers, marketers, and more. Evaluation is one of the hardest and most important problems in AI, and we need brilliant methodologists and builders to help. If our mission resonates with you, apply to join us. My DMs are open and I will personally look at every message I receive. Come work side-by-side with me, my cofounders @istoica05@infwinston, and our amazing team solving important new challenges in research and product every day. Shape the future with us; all types are welcome, from academics to founders.
We are building one of the fastest growing businesses in the world. Our mission is to measure and advance the capabilities of AI for real users. New product experiences are coming on lmarena for our community. New analyses and evaluations are coming for labs and enterprises to improve their AI systems based on real feedback. We are so grateful to all our users and customers and excited to serve you!
The fundraise was led by @felicis and @UofCalifornia investments, with participation from @a16z, @thehousefund, LDVP, @kleinerperkins, @lightspeedvp, and @laudeventures. Our whole cap table is digging in after seeing our growth. We are also excited to invite @pxd, GP at Felicis and former VP of Consumer Product at OpenAI, and Jagdeep Baccher, CIO of UC investments, to our board. Along with the amazing @AnjneyMidha who incubated us from the early days and cofounders @infwinston and @istoica05, we are building the strongest team in the world to solve AI evaluations for reliable deployment.
Onwards!
🚀Introducing Code Arena: the next generation of live coding evals for frontier AI models. Built to test how models plan, scaffold, debug, and build real web apps step-by-step.
Try Claude, GPT-5, GLM-4.6 and Gemini in Code Arena today!
🚀 Introducing Arena Expert: a new LMArena evaluation framework to identify the toughest, most expert-level prompts from real users, powering a new Expert leaderboard.
We also introduce Occupational Categories that underlie eight new leaderboards:
💻 Software & IT Services
✍️ Writing, Literature, & Language
🔬 Life, Physical, & Social Science
🎭 Entertainment, Sports, & Media
📈 Business, Management, & Financial Ops
🧮 Mathematical
⚖️ Legal & Government
🩺 Medicine & Healthcare
Explore how models perform across fields in thread 🧵 👇
@ericbahn Want to give Pastel a try? We auto-draft all your emails for you based on 1) your past emails and 2) any knowledge you give us. And 3) if you have an admin, it's super easy to share the load with them: https://t.co/LFgh7dq2hU
@zapiet app is currently down and blocking checkout. Second time in under a week. Please resolve and share the post-mortem and plans for better code testing.
@DoorDash_Help Phone support for restaurants isn't working right now! I called and it said "we cannot take your call right now". I need to cancel a dash but can't do it online.