It's disappointing to see the Financial Times portray the leader of an AI company dedicated to pursuing AGI while freely contributing its model weights and research to the world, with such an ugly and exaggerated cartoon illustration.
🚨 Head-to-Head opus 5 vs kimi k3
tested my chilli brand website prompt on opus 5 too
> opus absolutely nailed the 3D generation. It clearly beats kimi k3 in that area
> but kimi k3 is still better at the other parts of frontend. It has a much better understanding of @greensock and produces more consistent motion across all sections
@AnthropicAI still needs to improve its frontend generation
I wasn't expecting Kimi K3 to outperform Opus 5 twice in a row
Tried making Jelly Jungle using the exact same prompt with Claude Opus 5.0 (Claude Code) and Kimi K3 (Kimi Web).
Kimi K3 nailed the gameplay (once again). Physics, movement, sprinting, jumping, and obstacle logic all worked surprisingly well
Opus 5.0 looked much better visually with a cleaner UI, but the gameplay fell apart, controls were inverted, sprint barely worked, movement felt off, and the game was not playable.
Runtime and cost:
- Kimi K3: ~19 min, ~$6.1
- Opus 5.0: ~16 min, ~$11
that's roughly 2× the time and 3× the cost, yet Kimi produced the more playable game with the exact same promp
🚨 Kimi K3 just beat Claude Opus 5 at recreating Fall Guys. 🎮
Same prompt.
Same goal.
Completely different outcome.
We asked both models to recreate Fall Guys from scratch:
🟢 Kimi K3 (Kimi CLI)
✅ Fun gameplay right out of the box
✅ Smooth movement
✅ Solid physics
✅ Surprisingly polished overall feel
🔵 Claude Opus 5 (Claude Code)
✅ Better-looking UI
❌ Broken movement
❌ Kept adding unnecessary visual effects
❌ Gameplay lagged
❌ Spent most of its time iterating instead of converging on a playable result
Runtime & Cost
🟢 Kimi K3 • ~9 minutes • ~$4.40
🔵 Claude Opus 5 • ~17 minutes (still iterating) • ~$13+
That's roughly:
⚡ 2× faster
💰 3× cheaper
🎮 And the more playable game with the exact same prompt.
For rapid game prototyping and browser-based game generation, Kimi K3 is becoming a serious contender. 🚀
Kimi K3 just cooked GPT-5.6, Claude Opus 5, and Claude Fable 5
Opus 5 just delivered a very strong result on my WebGL creature trading card benchmark.
Both Opus 5 and Kimi K3 handled the core technical challenge well, rendering 400 floating cards with the correct motion and overall interaction.
The real difference showed up in the artwork.
Opus 5:
• Generated 17 distinct creature body types instead of relying on variations of the same base model.
• Produced 18 fully painted scenes with lighting that matched the creatures and environments.
Kimi K3:
• Leaned on a single simplified parametric creature model for most variations.
• Generated 21 basic three-stop gradient backgrounds rather than fully illustrated scenes.
Technically, both models passed the benchmark.
Artistically, Opus 5 took a clear lead. The diversity of creature designs, stronger environmental art, and more cohesive lighting made the overall presentation feel significantly richer.
Qwen 3.8 Max doesn't cease to impress me.
It generated a simulation of what a SUPERNOVA explosion looks like, rendered entirely in my browser.
It's become a permanent part of my workflow. Qwen 3.8 is a very reliable workhorse, especially when orchestrated by Fable 5 or GPT-5.6 Sol.
I am officially a Qwen fanboy.
KIMI K3 EXPLAINED IN 13 MINUTES. THIS IS THE BEST VIDEO YOU’LL WATCH ON IT
So much for dense models taking over the world.
Moonshot just proved the doubters wrong with Kimi K3, a 2.8 TRILLION parameter open beast that’s benchmarking right up there with GPT 5.6 and Fable 5.
It serves as proof that you can scale sparse models efficiently without bottlenecking your throughput 👊
Watch this, then read this great guide about Kimi K3 by @kirillk_web3 ↓
Tried recreating Fall Guys using the exact same prompt with Claude Opus 5.0 (Claude Code) and Kimi K3 (Kimi CLI).
Kimi K3 nailed the gameplay out of the box. Physics, movement, and the overall feel were surprisingly solid
Opus 5.0 had a shinier UI, but movement was broken, it kept adding unnecessary visual effects, lagged during gameplay, and spent most of its time reiterating instead of converging
Runtime and cost :
- Kimi K3: ~9 min, ~$4.4
- Opus 5.0: ~17 min (still iterating), ~$13+
That's roughly 2× the time and 3× the cost, yet Kimi produced the more playable result with the exact same prompt.
Kimi K3 cooked GPT-5.6, Claude Fable 5, and Opus 5 in a game demo.
Kimi K3 is making a strong case as one of the most capable models for AI-assisted game development right now.
Across recent game demo comparisons, it's consistently delivering impressive results against models like Claude Fable 5, GPT-5.6, and Grok 4.5.
What stands out isn't just the visuals. K3 has been showing strength in gameplay systems, 3D environments, and end-to-end prototype generation.
There's also growing speculation around the anonymous Arena .ai model "Kivine." Many in the community believe it's Kimi K3, though that hasn't been officially confirmed.
Regardless of the mystery model's identity, one thing is clear: competition among frontier AI models for game development is accelerating, and the quality of generated interactive experiences is improving at an incredible pace.
Ooo that's finally an impressive move from Ant Ling!
This is, more or less, a Kimi K3-lite (not from Moonshot). It's more or less on par with V4-Flash, but more than 2x smaller and faster. KDA+MLA, 1/64 sparsity, very interesting
Kimi K3 will be truly open-weight and 3x faster on Monday
Get ready to move all your standard workloads immediately
Anything that is on Sonnet or GPT 5.5 can move — it's faster, cheaper, and better 🚀🚀
🚨 Qwen3.8-Max-Preview keeps cooking
I asked it to build a Minecraft-style voxel game from a single prompt
This is the result:
> Procedural voxel world
> Terrain generation
> Flying mechanics
> Strong first-pass code
What else should I Test ?👇
Kimi K3 passed Fable 5 on Code Arena just 6 weeks after Fable's release, landing #1 in 6 of 7 frontend domains.
We sat down with Arena CEO and co-founder @ml_angelopoulos for a rapid fire conversation on what the data from real-world use tells us 👇
0:00 Biggest release of the year, or overreaction?
0:33 Kimi K3 beats Fable on Code Arena: verifying the votes
1:35 1679 points: a 17-place jump from K2.6
2:37 #1 in 6 of 7 frontend domains — but still #2 in gaming
3:15 Permanent shift or snapshot? Labs distilling Chinese models
4:23 Regulating Chinese models: the case for and against
6:30 $3 in, $15 out: open source is no longer 10% of the price
7:30 If you're Anthropic or OpenAI, what's your move?
8:09 How long before the leaderboard flips again?
🚀 Macaron-V1-Venti from @Macaron0fficial is now live on Novita.
🎁 Free through August 7.
Built for advanced agent workloads:
🔹 748B-parameter flagship model
🔹 744B base + four 1B LoRA specialists
🔹 Mixture-of-LoRA architecture
🔹 The first model post-trained on GLM-5.2
prompt engineering → context engineering → harness engineering → loop engineering → graph engineering?
You don't need to learn any of these.
you need to run /better-harness once and fix what it finds
Better Harness shipped in Qoder Desktop v1.18.0
Upgrade to the latest and have a try.