k3 report reaction thread. pre-registering my hopes:
- data recipes/techniques
- QB ablations at scale + refinements to the technique
- KDA ablations at scale + refinements
- PPO? i'm unsure if i really want to see PPO here tbh
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
Introducing the world's fastest tokenizer implementation, Gigatoken!
Gigatoken is ~500-1000x faster than HuggingFace, and ~100x faster than OpenAI's tiktoken for most tokenizer definitions on most machines.
These baselines are already multithreaded Rust implementations! 🧵
Similar to DeepSeek in January 2025, Panicans may think that the AI networking switch TAM will massively shrink because Kimi K3 uses KDA Attention, which reduces KV-transfer networking bandwidth by up to 10x. But the opposite is true, as we explain below. 👇️ 1/8🧵
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
As AI models continue to grow in scale and capability, shaping a model matters just as much as its size.
We're introducing a new series on AI Model Co-Design exploring the synergy between models and hardware. The first post focuses on how model dimensions influence GPU performance, and how the right design choices improve both system throughput and per-user responsiveness.
You can read it here: https://t.co/tvK9kRjgQx
Starbucks spends $400 million a year on software. Yesterday they announced they're moving off IBM and Microsoft to build their own custom systems in-house.
IBM dropped 3% and Salesforce dropped 4% on the news.
And honestly this is, unequivocally, the biggest signal I've seen since OpenAI and Anthropic launched their consulting arms back in Q1. The largest companies in the world are done paying for software that half fits how they work.
We saw this coming about a year ago. Moved everything we build off Airtable and low-code tools and went fully custom. Already paying off, and it's only going to compound from here.
This is the opportunity right now.
You get all of a company's data into one system. You build out a single operating system for the entire business. You cut out bad, redundant processes. Then you layer AI on top of it, under the correct processes.
That's the core of AI consulting. Helping companies actually operate better.
There are a lot of fly-by-night offerings circulating right now when it comes to Ai Services.
For example, 'second brains'.
Throwing scattered data into a second brain while the processes underneath stay broken does nothing. The companies who will absolutely destroy their competition over the next 5 years are rebuilding how they work from the ground up.
Starbucks is showing you what other companies will be doing over the next several years.
Your job is to position yourself to facilitate that process for as many companies as you can.
The derivatives of the Position vector with respect to time have interesting names:
Velocity (v) = change in Position
Acceleration (a) = change in Velocity
Jerk (j) = change in Acceleration
Snap (s) = change in Jerk
Crackle (c) = change in Snap
Pop (p) = change in Crackle
🧐
A Visual Introduction to Information Theory
(bookmark it)
Information Theory is such an beautiful and powerful subject.
In the era of AI, it's worth spending time learning about it.
Here is a highly-recommended read for anyone who wants real intuition for entropy and mutual information.
It's a visual, intuition-first guide to information theory.
It assumes only familiarity with basic probability, so it stays accessible while still reaching the fundamental limits of compression and transmission.
Paper: https://t.co/ydLqsF9ag8
Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX
new paper from our work at Meta!
**GPT-style language models memorize 3.6 bits per param**
we compute capacity by measuring total bits memorized, using some theory from Shannon (1953)
shockingly, the memorization-datasize curves look like this:
___________
/
/
(🧵)
an interesting property of traditional capitalism is that businesses often have to pick between scale and culture
but with AI, small teams are scaling revenue exponentially without losing their culture
a new pareto frontier is emerging in this regard
Introducing NEO’s 25 Degrees of Freedom, tendon-driven hands — nearing or surpassing human-level dexterity, strength, speed, and reliability.
For seventy years, robotics worked around the hand problem. The humanoid bet is the reverse: it lives or dies at the fingertips.
Zuck: “The pricing from some of the other labs is very extreme and has very high margins. We think that there’s a real ability to be able to offer frontier or very high-level intelligence at a much more affordable cost.”
Epic pricing war breaking out among agentic models.
Successfully training models on TPUs has been demonstrated by Anthropic through the past five-plus successful Claude releases. This is positive for the ML community, as Google TPUs continue to gain market share outside of internal Google workloads, giving frontier AI labs a viable alternative for training. 1/4🧵