Honored to be named to @TIME's TIME100 List of the World's Most Influential People in AI, even more so to share it with my co-founder @annadgoldie!
Anna and I started working on AI for Chip Design almost a decade ago. Last year, we started @RicursiveAI to transform end-to-end chip design from years to days! Watching that vision become reality piece by piece has been the most thrilling / fulfilling experience ever!
https://t.co/fsSMaFhuRg
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰
As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.
For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench.
Try it out today: https://t.co/UhudjKDaYI
More on verification scaling in my previous post.
(1/6) If a prompt can generate every layer of the stack, what are programming abstractions worth?
GPU programming is at a weird inflection point. For decades we stacked layers of abstraction to ease our cognitive load. This year, we find ourselves constantly deleting them.
I had so much fun creating the Self-Improving AI Agents course with @achowdhery, and teaching it twice in one year at Stanford!
We also collaborated with Stanford Online to make the course available online:
YouTube: https://t.co/t1uZlWz5eq
The field is moving incredibly fast, but we tried to focus on the core concepts that help us build better AI systems. I hope you enjoy it!
(1/9) I'm thrilled to share the open-source release of Mixture-of-Kittens (MoK), our MoE megakernel for NVL72s! MoK fuses all mixture-of-experts communication and computation into a single, fully deterministic kernel, and powers Composer training across tens of thousands of GPUs.
Joint work with @nash_c_brown, @hmwildermuth, @tmwilliamlin168, and @ellev3n11
What if we could directly bake knowledge into an LLM brain? Our most recent work shows how to build facts into a Transformer block. No gradient descent required!
(link to blog and paper in comments!)
We can instantly build knowledge into a Transformer block, no gradient descent required!
New work w/ amazing team @jerrywliu, @ronnygjunkins, @EyubogluSabri, Atri Rudra and @HazyResearch!
To learn more, checkout our:
📝Blogpost: https://t.co/fLQpWqknvw
📄Paper: https://t.co/pjMqCnVniW
MLPs store facts in language models. Can we write them into Transformers without training?
New work w/ amazing team @garctrob@ronnygjunkins@EyubogluSabri, Atri Rudra & @HazyResearch gives a ✨closed-form✨ recipe for fact-storing, Transformer-ready MLPs. Accepted at COLM 2026!
Introducing the world's fastest tokenizer implementation, Gigatoken!
Gigatoken is ~500-1000x faster than HuggingFace, and ~100x faster than OpenAI's tiktoken for most tokenizer definitions on most machines.
These baselines are already multithreaded Rust implementations! 🧵
How can we extract richer signals from AI Feedback?
Introducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀
The key idea:
- Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale)
- Take the expectation over the full logprob distribution of score tokens
- Scale repeated evaluation and criteria decomposition
You can use these fine-grained signals for more effective test-time scaling, RL, and agent monitoring! It achieves SOTA across Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench 👑
Advised by @Azaliamirh@istoica05@drmapavone@chelseabfinn
🧵👇
So excited to share that I've recently defended my PhD and joined @ColumbiaLaw as an Associate Professor! I'm absolutely ecstatic to be joining such an incredible community, and looking forward to much future work at the intersection of law and AI.
Doing a JD/PhD in computer science at @StanfordLaw and with @HazyResearch was the experience of a lifetime, and I'm so grateful for the opportunity. If you're a student interested in this intersection–please reach out!!
Suppose you're handed a fine-tuned LLM that secretly favors a certain entity. The bias goes completely undetected because it only surfaces on one specific unknown topic.
So how do you catch a bias you can't search for? You amplify it.
Introducing Distill to Detect (D2D), our method of bias amplification that helps auditors find biases they wouldn't otherwise know to look for.
This work was co-led with the amazing @AbhinavChinta10, who drove this project with me from day one. Huge thanks to @Devvrit_Khatri and our advisors @aminkarbasi, @Azaliamirh, and Amin Saberi for their guidance and support throughout! 🙏
📄 Paper: https://t.co/nU9eAMtFIU
📝 Blog: https://t.co/XhuuQYTxmf
💻 Code: https://t.co/CKWC6yVXcK
For more information, please see the thread below. 🧵
Happy Fourth of July!
We’re a proudly American AI lab, and we’re excited and optimistic about what our country will do together for the next 250 years—no doom and gloom here!
Thank you to the patriots and people who built America and the people still building her more perfect union. We love this country and believe service to God, Country, and family are prerequisites for a meaningful life and to keep our country great.
God Bless America!
Is there really anything more American than a Stanford CS lab filled with some of the smartest AI folks I've ever had the pleasure to work with being located... in a strip mall? 🇺🇸🤖🎇🎆 Happy Fourth!
🔎 In high-stakes domains, a missed search result is not just a search error - it can mean a missed opportunity, delayed decision, or denied service.
Yet most retrieval systems still rank by keywords and embeddings. They can surface candidates that look relevant while missing those that actually satisfy the constraints.
What if an agent could turn those constraints into a retrieval tool that searches for viable options directly?
🌐 Project: https://t.co/sdRZD6QCHt
📄 Paper: https://t.co/GfAg16TI2e
With @Yufei_1001, @kelakexyl, @YuChiangWang1, @ChiehJuChao1, and @MonicaSLam. Grateful to our collaborators across Stanford and Mayo Clinic, and to the broader @stanfordnlp and @StanfordHAI communities that helped shape the environment for this work.
#InformationRetrieval #AIAgents
<🧵1/n | 𝗦𝗮𝘁𝗜𝗥>
Together with my co-founders Michael @MichaelPoli6, Stefano @Massastrello and Armin @athmsx, I am excited to announce @RadicalNumerics is emerging from stealth with a $50M seed round to build general biological intelligence.
We’re also sharing an early preview of our new model Omnii, the most powerful genome language model to date.
Omnii preview link:
https://t.co/ouikMtRVwf
At Radical Numerics, our mission is to master the code of life, and to drive the frontier of biological AI for both design and defense.
This is our dual mandate, which comes from something our own team helped make possible.
Our founding team trained Evo and Evo 2, the largest biological AI models (40B params) trained on DNA sequences. Trillions of tokens across all of life, from microbes to mammals. It’s fully open source, and created the field now known as generative genomics.
Last year, scientists used Evo to generate the world’s first complete genome from scratch using AI. Turns out it was a bacteriophage—a type of virus. It functioned in the real world, and in this case it was harmless. But for us, it was a clear turning point.
It showed that AI is no longer just analyzing biology. It is on the cusp of generating functional lifeforms. Eventually, AI will have the power to design and control life itself.
That should make all of us incredibly excited, and incredibly uneasy. (Anyone can design DNA with a new function, and have it synthesized and delivered, like something from Amazon Prime).
The same technology that will help us cure cancer is the very technology that might create the next global pandemic, or worse, allow the creation of bioweapons that can wipe out populations.
We believe these forces are inseparable. If you work on the frontier of biology, you have to build technology to safeguard it from its misuse. Existing biosecurity tools are sorely losing the arms race, relying on outdated “have I seen this exact thing before?” style algorithms.
We founded Radical Numerics to turn the tide.
And we can’t do that by training on textbooks and natural language. We must understand the language of biology from the raw physical data itself, to reason across every molecule and modality, from DNA to proteins.
The next frontier for AI goes far beyond chatbots or video generators to models that can understand and engineer life.
Today, we’re previewing Omnii, which is already far surpassing Evo 2, and will continue improving as we scale and add new modalities (training now).
1. For human health, Omnii can read and write whole genomes (more on writing later). It’s state of the art (SOTA) on detecting causal variants for disease, and can rank Alzheimer's mutations zero-shot. We’re partnering with a diagnostics company to use Omnii for early cancer detection (pancreatic and multi-cancer).
2. For defense, Omnii is SOTA at detecting AI-generated pathogens. We benchmarked existing detection tools, and they simply can’t detect the AI-generated ones (“deepfake viruses”). We’re partnering with a US national lab to pilot Omnii for detecting the next pandemic, both natural and AI-generated.
We have a data center full of Blackwells in construction now to build the most powerful biological AI models ever. This mission takes a new kind of AI lab that can actually scale on physical, biological data: new alignment research (mid/post training), scaling long context, building out mech interp teams to dissect what these models learn, new architectures and systems designs, all from the ground up.
Our team is made up of AI researchers and scientists from top labs and institutions (e.g. Stanford, MIT, Google DeepMind), but more importantly, we all share the belief that this is the most important challenge of our lifetime. If you feel similarly, we are hiring. We aim to bring the brightest minds in AI and science together to save lives.
Thanks to our partners on this journey, led by Emergence Capital @emergencecap, with Obvious Ventures @obviousvc, Triatomic @TriatomicCap
, and Patrick Collison @patrickc. Our advisors include Eric Horvitz @erichorvitz, CSO of Microsoft, Chris Re @HazyResearch of Stanford, George Church @geochurch of Harvard, and Andrew Weber @AndyWeberNCB, former Assistant Secretary of Defense for Nuclear, Chemical and Biological Defense Programs.
Fortune article: https://t.co/L3f3f1329T
Jobs: https://t.co/EzsHSMcGJ1
What happens when multi-agent systems stop relying on a central “controller” agent? Can agents coordinate by sharing results directly with each other?
Introducing Decentralized Language Models (DeLM): we let agents coordinate asynchronously through a shared context. Agents claim tasks from a queue and write back compact, verified results as they finish, making progress visible to all workers without requiring a main agent to merge, filter, and rebroadcast it.
New paper with @azaliamirh!