Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model.
This release marks a major leap in agentic performance over Kimi K2.7 Code (#23 to #4). Based on 8K+ live agentic sessions, Kimi K3 leads on confirmed task success rate (#1). It also posts a strong +20.6% on praise vs. complaint (#3). It currently lags the field in steerability (#14) and bash recovery (#17).
Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents.
We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model.
Here's a primer on the 5 signals:
User-satisfaction proxies
- Confirmed Success: an explicit "yes that worked" feedback from the user
- Praise vs. Complaint: implicit sentiment in users reactions
- Steerability: can the model course-correct when you push back?
Tool-use proxies
- Bash Recovery: how it recovers from CLI errors (primary signal for tool use)
- Tool Hallucination: does it call tools that don't exist
Below we break down how Kimi K3 scored across the 5 signals, drawn from tasks submitted by a global community of users.
Congrats @Kimi_Moonshot on another big milestone!
@AnthropicAI Not on a monday bro come on:(
Claude down, again.
API Error: 500 {"type":"error","error":{"type":"api_error","message":"Internal server error"},"request_id":"req_011Ca22EE5oiz6czHW1t3E1S"}