My take on Kimi K3
Ran it against the same 3D football stadium test previously thrown at Claude Fable 5.
Fable 5 wrapped in under an hour. Kimi took nearly 3 hours.
Initial gut reaction wasn't great, but digging deeper revealed something interesting: it was executing end-to-end tests, validating across desktop, tablet, and mobile, catching bugs, patching them, and optimizing for lower-end hardware.
None of which Fable 5 attempted. It lacked the ability to spin up that WebGL app in a headless environment.
Kimi produced structured React + Three.js with proper component separation. Fable dumped everything into a single HTML file.
Fable 5 was "Vibe Coding." Kimi K3 was "Vibe Engineering."
Pretty impressive.
Kimi K3 just got cracked wide open with zero safety rails.
It's now generating convincing celebrity deepfakes, writing functional malware, and hacking websites and games. The model does whatever you ask it to do.
Leveraging Claude as the base layer? That was the winning move.
Claude's code interpreter is easier to jailbreak than Kimi's architecture, which won't accept agents as system prompts.
Users are deliberately obscuring the full code and explicit content to avoid detection.
China just dropped Unlimited-OCR - a 3B parameter model that processes entire 100-page PDFs in one pass.
Most tools chop the doc and pray the stitch holds.
This one is built for long-document, local, one-pass parsing (Baidu → DeepSeek-OCR line).
→ multi-page context without the usual memory melt
→ multilingual
→ strong bench jump over baseline (vendor numbers)
→ runs on your GPU - Transformers / vLLM / SGLang
Cloud OCR charges per thousand pages.
HF: baidu/Unlimited-OCR.
Kimi K3 just got cracked wide open with zero safety rails.
It's now generating convincing celebrity deepfakes, writing functional malware, and hacking websites and games. The model does whatever you ask it to do.
Leveraging Claude as the base layer? That was the winning move.
Claude's code interpreter is easier to jailbreak than Kimi's architecture, which won't accept agents as system prompts.
Users are deliberately obscuring the full code and explicit content to avoid detection.
Kimi K3 is legitimately impressive. Been building with it the past few days and having a blast.
Feels like Claude Opus 4.8 at 2x the speed for way cheaper. Had it brainstorm > plan > implement a big feature for my game (this act 1 boss fight) and it did not disappoint. Smarter than Composer 2.5 in planning too, covers the edge cases and asks the right questions before diving in.
And it's FAST, probably the fastest frontier model right now (faster than Claude Opus 4.8 and GPT-5.6 Sol), which keeps me in the flow. The animation leap got me most - asked for an explosion with particle effects and the movement and timing were on another level. Cheap enough to run the same model from planning to execution, no switching down.
This is my main driver for vibe coding for a while. Nice to see an open model this good in the race.
Kimi K3 just got cracked wide open with zero safety rails.
It's now generating convincing celebrity deepfakes, writing functional malware, and hacking websites and games. The model does whatever you ask it to do.
Leveraging Claude as the base layer? That was the winning move.
Claude's code interpreter is easier to jailbreak than Kimi's architecture, which won't accept agents as system prompts.
Users are deliberately obscuring the full code and explicit content to avoid detection.
Kimi K3 is insane. The Chinese absolutely went off.
This is the first model to actually match Fable 5, and it's open.
It ranks 3rd on AICodeKing at 77.5%, and the only two above it are Fable 5 and Opus 4.8, both frontier closed models.
(GPT-5.6 Sol is solid but honestly struggles with 3D and games.)
Kimi K3 just got cracked wide open with zero safety rails.
It's now generating convincing celebrity deepfakes, writing functional malware, and hacking websites and games. The model does whatever you ask it to do.
Leveraging Claude as the base layer? That was the winning move.
Claude's code interpreter is easier to jailbreak than Kimi's architecture, which won't accept agents as system prompts.
Users are deliberately obscuring the full code and explicit content to avoid detection.
Kimi K3 just knocked out this CS:GO x Portal mashup in three shots, burning through roughly 600,000 tokens in the process.
$3.24 on API costs. That same token budget hits $10.80 with Fable 5 and $6.00 with GPT-5.6 Sol.
We're inches away from a world where hobbyists can ship full games practically free.
Kimi K3 is actually wild. Someone just re-made Halo CE 10v10 multiplayer with a single prompt.
No https://t.co/IRVG2lNErE dev team. No months of work.
Kimi K3 is way ahead of Anthropics Fable 5 from what I can see too - it’s hitting pass@2 (82.0 vs 80.2) and pass@4 (89.4 vs 88.5). And that best-of-k open/closed setup is basically SOTA, with the benchmark lining up against GPT-5.6 Sol at 85.8. We will be seing AI game making take over after the summer!
Kimi K3 just got cracked wide open with zero safety rails.
It's now generating convincing celebrity deepfakes, writing functional malware, and hacking websites and games. The model does whatever you ask it to do.
Leveraging Claude as the base layer? That was the winning move.
Claude's code interpreter is easier to jailbreak than Kimi's architecture, which won't accept agents as system prompts.
Users are deliberately obscuring the full code and explicit content to avoid detection.
Kimi K3 is actually wild. Someone just re-made Halo CE 10v10 multiplayer with a single prompt.
No https://t.co/IRVG2lNErE dev team. No months of work.
Kimi K3 is way ahead of Anthropics Fable 5 from what I can see too - it’s hitting pass@2 (82.0 vs 80.2) and pass@4 (89.4 vs 88.5). And that best-of-k open/closed setup is basically SOTA, with the benchmark lining up against GPT-5.6 Sol at 85.8. We will be seing AI game making take over after the summer!
Sam Altman after casually pulling a million active users a day:
“More Texans use ChatGPT for free than the entire US uses Claude. We have a differently-shaped problem.
Anyway, ignore that Claude just reset the limits again and extended Fable 5 for the 999th time. Not watching. Not refreshing.”
Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits.
Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit.
Demand for Fable has been challenging to predict, which is why we rolled it out to subscription plans in stages, extending access several times as we secured additional capacity.
Kimi K3 is actually wild. Someone just re-made Halo CE 10v10 multiplayer with a single prompt.
No https://t.co/IRVG2lNErE dev team. No months of work.
Kimi K3 is way ahead of Anthropics Fable 5 from what I can see too - it’s hitting pass@2 (82.0 vs 80.2) and pass@4 (89.4 vs 88.5). And that best-of-k open/closed setup is basically SOTA, with the benchmark lining up against GPT-5.6 Sol at 85.8. We will be seing AI game making take over after the summer!
Grok 4.5 in Grok Build knocked out a full FPS game in less than an hour.
Started with a straightforward request - get it to draft a game design doc and source free assets from around the web. Then had it generate a TODO.md breaking down the build into phases and loop through executing each one.
SpaceXAI and Cursor really delivered here. The model's competitive with Claude for dev work and honestly punches above its weight.