Congratulations to Zhilin Yang, founder and CEO of @Kimi_Moonshot, on the latest Kimi release. What a huge win for the open-source community!
It feels like just yesterday Zhilin was graduating from my lab at CMU, jointly co-advised with William Cohen. Not only did he complete his Ph.D. in just four years, but he also made truly fundamental contributions to ML during his time at CMU.
What a spectacular career! Congrats again Zhilin, and thank you and the entire Kimi team for everything you're doing for the open-source community.
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
What does it really take to build a general purpose AI agent that works reliably in the real world?
Join us at the next AI Tuesdays for a fireside with @itsumeshk (@runable_hq) and @SveeVJ (@ramainai) as we unpack general purpose agents. ����
If you’re building AI agents, this room is for you.
Limited seats. Apply: https://t.co/lbqzyfzE87
#AITuesdays #AIAgents #AIBuilders #Founders
@varadmaniyar @NexusVP
I had the honor of giving a keynote at the International Conference on Machine Learning in Seoul last week titled “What will be left for us to work on?” I addressed the widespread anxiety about how we should adapt as AI capabilities increase. I was thrilled by the talk’s reception, so I have made my slides available, annotated with a lightly edited transcript: https://t.co/vNgRJCL57B
I made three arguments. First, the "AI as Normal Technology" framework is a correct and useful as a way to think about AI’s impacts, unless and until there is some future discontinuity such as through recursive self-improvement. Second, even though we should take recursive self-improvement seriously, there is no milestone that companies might achieve in the lab that will suddenly put us all out of work. Third and finally, jobs of the future will be radically different, and a lot of adaptation will be needed. I shared my thinking about what this might look like and ended with a vision of human/AI “co-superintelligence”.
🇮🇳🥇🥇🥇🥇🥇 ALL FIVE Indian students win GOLD at IPhO 2026 — India ranks World #1 (joint #1)
Having trained at IMOTC myself, I know exactly how brutal that grind is.
Massive respect to Kanishk, Riddhesh, Rishit, Shresth, Svarit and @HBCSE_TIFR for building this pipeline year after year.
🇮🇳GOLDEN SWEEP FOR INDIA 🏆
All 5 Indian students win GOLD at the 56 International Physics Olympiad (IPhO) 2026 in Bucaramanga, Colombia -placing India at Rank #1 in the world (jointly with China, Kazakhstan, Russia, South Korea & Taiwan) among 381 students from 87 countries 1/5
Imagine a fable 5 quality model that’s 3-4x less expensive in less than 6 months. And an Opus 4.8 grade model that can run on a local device in less than 12 months. Greater than 50% chance that these events will happen. Worth keeping in mind when you make predictions about the future.
1/ muse spark 1.1 is an industry-competitive agentic and coding model. across many agentic evals it rivals gpt-5.5 and opus-4.8.
available now through the new meta model api and in meta ai. 🧵
Math proofs are cool, but the real revolution?
Verifying real-world claims in accounting, tax law, compliance, medicine & more; reliably, efficiently, and at scale.
Proof trees look different across domains. One-size-fits-all won't cut it. AI prover architectures must be purpose-built to navigate the specific reasoning trees, resource constraints, and rule stability of the environment they are verifying.
Read why specialized provers are the future → https://t.co/FbNiJ3Knhy
Authored by: @ArnavAMehta
Watch this space for more to come from our Technical blog-post series.
The most valuable AI workflows don’t show up on benchmarks. They live in terminals, browser tabs, and internal automations.
We’re bringing those builders together for an AI Builders Show & Tell.
No panels. No pitches. Just live demos.
📅 July 7, 6:30 PM (link below)
@NexusVP
We've kept hearing how GLM-5.2 beats Opus 4.8, and are skeptical of benchmarks - so we tested them on a real bug from the Cline repo. While both models fixed the issue, GLM was the winner in terms of cost and code quality:
- GLM used twice as many tokens (GLM 1.1m vs Opus 660K) but cost half as much (GLM $0.41 vs Opus $0.81)
- Opus finished quicker - 1.6 min and 12 tool calls vs GLM 4.7 min and 28 tool calls
- GLM cleaned up dead code and verified the build compiled before completing. Opus didn't - it left type errors that passed tests but broke the production build.
Both runs used the same Cline harness prompting and tools, so it seems GLM is RL trained to spend more tokens verifying its work before completing. Impressive work by the @Zai_org team!
There’s a big misconception about how GLM 5.2 was trained. Yes, they distilled Claude and GPT 5.5 — but distillation is not how they matched Opus quality. Distillation only fixed the cold start problem in RL.
RLing an agentic coding model isn’t rocket science. In simplified terms:
1. RL needs trajectories — rollouts where the model actually completed a task in some env
2. No successful trajectory on a task = zero gradient = you can’t RL it. This is the cold start problem
3. Distillation solves it. You seed your model with knowledge from a smarter one (Claude, GPT) on tasks it can’t do yet
4. Now it produces positive trajectories on those tasks
5. RL on those trajectories and hill climb agentic coding
6. At that point you no longer need to distill and can solely hill climb RL to better models
This is an interesting curve. I’d argue it’s harder to get to Opus 4.8 from scratch than to go from Opus 4.8 → Fable/Mythos tier.
GLM 5.2 is already producing positive trajectories, so they have plenty to RL on — they’ll keep climbing to Mythos quality without distilling any further. They no longer need American models.