Exciting news: Claude Opus 5 with Max reasoning is #1 in the Frontend Code Arena and Text Arena with factuality on!
Claude Opus 5 with default reasoning high is also very strong landing #3 in Frontend Code Arena, right behind Kimi K3 - and #2 in Text Arena (factuality on).
This is real world data that @AnthropicAI's newest model holds up on real world tasks: agentic web coding, document reasoning, and general chat capability.
Claude Opus 5 Max’s score is still preliminary. We’ll continue to see how scores converge and share updates.
Congrats to @AnthropicAI on the SOTA release!
Kimi K3 weights have been released! Kimi K3 is now the leading open weights model at 57 in the Artificial Analysis Intelligence Index
Moonshot has released the weights of their 2.6T parameter model under their 'Kimi K3 License' which we have labelled ‘Commercial Use Restricted’.
Restrictions compared with more permissive licenses such as MIT or Apache 2.0 include requiring model-as-a-service businesses with more than $20 million in revenue to enter into a separate agreement. Commercial products with more than 100 million monthly active users or $20 million in monthly revenue also need to display “Kimi K3” in the user interface.
@Kimi_Moonshot has also released their technical report with insights into model’s architecture and training approach.
Links below to the weights on Hugging Face, the Technical Report and further benchmarks 🔽
Gemini 3.6 Flash just dropped and it is the EXACT same model as 3.5 Flash.
50 on the Intelligence Index. 3.5 Flash: also 50. Identical intelligence, slightly faster.
Another flop from Google. They shipped a speed patch and called it a version.
Six months of silence for THIS?
Nobody wants a faster Flash.
Everyone wants Gemini 3.5 Pro.
Ship the real one, Google.
KIMI K3 is here, and it blew my mind 🤯
yes, it is expensive and very slow
Took over 30 minutes, and blew an ENTIRE $20 kimi sub before it even finished the one prompt for this game.
BUT OMG...
the output geniuinly has fable 5 level magic built into it, this model falls between gpt 5.6 and fable 5 imo.
you can play the game at kimisurfers (dot) com right now, and see what I mean.
This model is too expensive to replace fable 5, BUTT it shows that the chinese open source labs are CATCHING anthropic and open ai FASTTT.