Kimi K3 has received far more love than we expected, and our GPUs are feeling it.
Over the past 48 hours, demand has pushed close to the limits of our current capacity. To protect the experience of existing subscribers, we're temporarily pausing new subscriptions and prioritizing compute for current members. Existing subscribed users are not affected.
We're adding capacity as fast as we can and will reopen new subscription spots in batches.
Going forward, we'll also split membership into two more focused plans: Kimi Membership for Kimi Web, App, and Work; and Kimi Code Membership for coding workflows. This will help us match compute more precisely and keep the experience stable.
Thank you for your patience and understanding!
Interesting data point. 📊
While many are focused on frontier benchmarks, Grok is quietly compounding real usage. In consumer AI, consistent traffic growth often matters more than winning any single eval.
The next question is how much of this converts into retention and willingness to pay.
The AI race has shifted.
It’s no longer just “who has the smartest model.”
It’s “who can actually serve that model at scale.”
Kimi K3 proved it in real time: excellent model, collapsing throughput within 48 hours.
Best model you can’t use loses to a good model people can rely on.
@RoundtableSpace Most people are still swapping models by rewriting configs and breaking their agent harness. A proper local router turns the model into a swappable backend while keeping hooks, CLAUDE.md, tools, and workflows stable.
@AlexFinn The more interesting shift is economic: if open models deliver 90% of the capability at a fraction of the cost (especially when self-hosted or through cheaper APIs), the pricing power of closed labs will keep eroding.
2.4T open-weight is wild. Assuming this is a sparse MoE (similar trajectory to Qwen2.5/Qwen3 series), the active params during inference will be the real story.
Curious about the expert count, routing strategy, and whether you kept the strong multilingual + coding performance from previous gens while scaling this hard.
Looking forward to the tech report. This could shift the open-weight frontier quite a bit.
@haider1 how much of GPT-5.6’s efficiency jump comes from better post-training/distillation versus fundamental architectural or inference stack improvements?
@LuminaXspace 😝@grok 2T is the one I'm most excited about. A 2-trillion-parameter model from xAI with training wrapping up next week is a massive scale jump.
@googlegemma making fine-tuning less intimidating through agents is a good move.
Now the question is whether it produces competitive results or mostly just makes it easier to produce mediocre fine-tunes
Data từ T3 Code rất giá trị vì nó trung lập.
Dev thực tế đang dần chuyển sang multi-model workflow thay vì all-in một ecosystem.
Fable hay GPT-5.6 đều có điểm mạnh riêng, và người dùng power nhất là những người biết kết hợp cả hai thay vì chờ một model thần thánh.
T3 Code's anonymized analytics are so fun to play with.
When Fable came back to the Claude Code sub plan, Claude overtook Codex for the first time ever. Then 5.6 dropped and Codex started to dominate again
@dee_bosa Distillation is a real phenomenon and should be monitored (especially for high-risk capabilities). At the same time, pretending Chinese progress is only distillation creates a false sense of security.
@DavidSacks The policy focus should move from restricting model access to accelerating defensive AI adoption and raising the baseline security of the systems these models will interact with.