K3 is genuinely a very strong model, and it's incredibly token efficient.
one-shot generated this 3D WebGL landing page with scroll-trigger animations.
every leaf is individually modeled.
Total cost: $2.37.
So much demand for Kimi-K3 that the company is forced to halt new subscriptions for the moment. That's because everybody is ditching Anthropic and OpenAI and switching to Kimi-K3.
It's just as good in nearly every way, yet a fraction of the cost.
I wonder how long it will take U.S. stock market investors to figure out we've just witnessed the second "DeepSeek moment." Yet in this case, it's Kimi (from Moonshot), not DeepSeek.
Alibaba is a major investor of Moonshot. Kimi, in turn, uses a big chunk of AliCloud GPUs (oftentimes competing away Qwen team's resources)
10 days ago, Alibaba previewed quarterly earnings saying AliCloud is growing 40+%. This was ofc before K3
Could be a monster quarter for APAC's biggest cloud
Kimi K3 has received far more love than we expected, and our GPUs are feeling it.
Over the past 48 hours, demand has pushed close to the limits of our current capacity. To protect the experience of existing subscribers, we're temporarily pausing new subscriptions and prioritizing compute for current members. Existing subscribed users are not affected.
We're adding capacity as fast as we can and will reopen new subscription spots in batches.
Going forward, we'll also split membership into two more focused plans: Kimi Membership for Kimi Web, App, and Work; and Kimi Code Membership for coding workflows. This will help us match compute more precisely and keep the experience stable.
Thank you for your patience and understanding!
Qwen 3.8 Max is actually a very good model. Outperforms all Opus models in my benchmarks. There are some issues in tool calling with some harnesses. But, the raw intelligence is just crazy.
Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:https://t.co/YRvcGdB9Bv
China:https://t.co/PKMUNwUuRp
today's fun ai coding tricking - when doing the high-level part of your plan, do an amazon-style "working backwards" approach - write the customer-facing blog post about the feature BEFORE you start building
this can surface so many edge cases and help you dial in on prioritizing the important parts, and just is much nicer to read (plus you can post it after you ship 🙂)
We tested Kimi K3 to build liquid metal text that wraps with your cursor in Three[.]js
- Used the /design command
- 3 prompts to get this effect
- Full session cost: $0.77
Open models deliver their best when used with the right harness.
Dynamically Loaded Tools let you inject tools on demand during a conversation: start with only a few core tools, and insert additional tools into messages when the conversation actually needs them — reducing token usage and improving tool-selection accuracy at the same time. For the reasoning behind this design (lazy loading, tool registry) and combined practices.
👉 https://t.co/EcPhEljrBP
When your application needs a large number of tools, declaring all of them up front in the top-level tools field of every request leads to Tool Definition Bloat: every request carries the descriptions and parameter schemas of all tools, driving up token usage, and the more candidate tools there are, the more likely the model picks the wrong tool or constructs invalid call arguments.
what’s the best tooling to let agents use a browser rn? agents always seem to burn a ton of tokens using things like agent-browser esp using helium (@uwukko maybe you have some good guidance here?)