New podcast with @xeophon on all things open models. More on Kimi K3, Qwen 3.8, GLM-5.2, Xi's WAIC speech, distillation, the open-closed Gap, and what's next.
Chapters:
00:00 Welcome & context
04:38 Living with / using Kimi K3
08:53 GLM 5.2’s continued role
12:47 How are the Chinese models this good?
17:41 Data, environments, and a tour of the Chinese labs
19:47 Roundup of Chinese providers: Qwen, DeepSeek, MiniMax…
24:08 The US open-model ecosystem
30:25 Frontier vs. near-frontier, and the cybersecurity case against bans
34:58 Distillation and the Ben Thompson debate
44:12 Predictions and a frontier tier list
48:36 Wrap-up
Hoping to keep doing a few more of these on @interconnectsai. Crucial times in AI, we're working hard to share our expertise.
Terence Tao, UCLA professor and the most decorated mathematician alive:
"Hedge funds pay $500K a year for one skill: telling a real pattern from noise. I've spent my life on it, and most people get it exactly backwards."
this free lecture holds the entire "find the signal" problem those firms pay for, given at UCLA by the most decorated mathematician alive and posted for nothing.
at the board it's simple. Tao's lifelong theme is that almost nothing is purely one thing. There is perfect structure, like a clock, and perfect randomness, like a coin, and essentially everything real is a mixture of the two. The work is separating them. The primes are the perfect test case: they look scattered and lawless, and yet Tao and Ben Green proved they contain arbitrarily long evenly spaced runs. Order was hiding inside apparent chaos the whole time. That's the whole "signal detection", minus the marketing.
He gave this lecture at UCLA and it has been free ever since. Same point as my article above: the mathematics those firms pay half a million dollars for is public, documented and free to anyone tonight.
the lecture is free and anyone can watch it. what nobody can sell you is the judgment to know when a pattern is real and when your own eyes invented it. That judgment is the entire job, and it takes years to build.
New lecture! This one is a recap of a bunch of history of preferences, the nature of rewards, how RLHF is formulated, which were once seen as central problems in the field. How much as changed.
Still... super interesting to understand our optimization tools today. Books coming soon :D
00:00 Intro & context
07:34 A short history of preferences (from Aristotle to the VNM Utility Theorem)
20:17 A brief overview of preference data (from the last two years of my practice)
31:11 Open questions in RLHF data
Lecture 8, covering Chapters 10 & 11 of my book.
Richard Feynman stood at a Cornell blackboard in 1964 and explained the problem every AI lab is fighting in 2026. The BBC filmed it. Almost nobody watches it.
The lecture is about why nature only answers in mathematics. Every team forcing language models to reason is hitting the wall he mapped 62 years ago.
He was 46. The Nobel Prize came 11 months later. The footage survived on film reels and now sits free on YouTube with fewer views than a keyboard unboxing.
Watch the blackboard section near the middle. He takes 1 of Kepler's laws and rebuilds it from nothing, with notation a 12-year-old can follow.
No slides. No jargon. 1 piece of chalk.
An ML engineer I know paused it 4 times and made his whole team watch it before standup.
You're 62 years late. The lecture is still free.
My friend applied to 300 tech jobs in two years. No MIT. No Stanford.
Last month Anthropic offered him $750,000.
I asked him how he broke in from zero.
He sent me the exact video that got him in. Andrej Karpathy's 3-hour course on building a full LLM from scratch.
Anthropic's own researcher shows you how LLMs like ChatGPT & Claude are actually built.
I watched it last night.
Halfway through, I realized it's embarrassingly simple to break into an AI lab.
Bookmark this and read the article below.
• 00:00 - intro to LLMs
• 12:41 - LLM training pipeline
• 31:58 - LLM tools & plugins
• 41:14 - LLM transformer architecture
• 2:19:02 - scaling the LLM
David MacKay, Cambridge professor:
"Citadel and Jane Street pay $900K for people who get one idea I taught free at Cambridge: information is measurable, and whoever measures it best wins. the casino knew it first. the machines know it now."
this free lecture holds the entire "secret alpha" the funds and the AI gurus sell, and the man teaching it put his whole course and his whole book online for free and never locked a single page.
at the board it's simple. information is just how much a signal cuts your uncertainty, and you can measure it in exact bits. the more an edge shrinks what you don't know, the more it is worth, and Shannon proved there is a hard ceiling on how fast you can compound it into money. that is the whole "edge", and the whole "AI magic", minus the marketing.
MacKay taught it at Cambridge and gave the book away free in 2003. it has been free ever since. same point as the casino piece above: the house wins with a small, measurable edge, repeated and sized right, not with a secret.
the theory is free and every fund and every model already runs on it. what they cannot sell you is the data to feed the machine and the discipline to size the bet. that is the part that actually compounds, and no course can put it in a formula.
Similar to the panic over DeepSeek R1, some uneducated people think Kimi K3’s use of linear attention (KDA) is bad for NVIDIA, HBM, DRAM, and networking because it has relatively lower KV-cache requirements. The opposite is true, and we explain why below. 👇️ 1/8🧵
el fundador de una empresa china de IA valorada en más de $20,000,000,000 acaba de dar una clase de 40 minutos sobre enjambres de agentes
la explicación más clara que he visto sobre sistemas de IA a gran escala
cámbiala por tus 2 horas de Netflix de esta noche
Kimi K3 achieved huge success today.
This was no accident.
Three months ago, Kimi CEO Zhilin Yang had already revealed the secret in this 30-minute GTC keynote.
He said the next scaling frontier had three dimensions:
- Token efficiency
- Context length
- Number of agents
The important point was that Kimi was not planning to win simply by training a larger model.
They were investing in the underlying machinery: better optimizers, training stability, memory efficiency, long-context architecture, and agent swarms.
Kimi K3 is now the product version of that thesis:
- 2.8T total parameters
- 1M-token context
- Sparse MoE architecture
- Strong long-horizon coding and research
- Multi-agent workflows
The most predictive line from the talk was probably this:
“Open models cannot be just open. They also have to be great.”
Kimi understood that open source only becomes strategically powerful when the model is good enough that developers actually want to use it.