Ten years ago, my AKLT and KT papers were suddenly in the spotlight after Haldane’s Nobel Prize. Today, my lesser-known 2018 and 2025 papers seem to be getting some attention --- ironically, because a new OpenAI model solved a problem I had long wanted to tackle.
For years, my long-term life plan was to keep trying to prove the Haldane gap for as long as my brain kept working. Now AI has forced me to revise that plan.
Still, perhaps I should have seen this coming: I chose HAL as my first name when I wrote my very first paper.
Another one: The Kakeya problem in 4-D!
Just this summer, Hong Wang won the Fields medal for the Kakeya proof in 3-D (Wang-Zahl, 2026)
https://t.co/NygzEfU4CZ
OpenAI is downplaying this for PR reasons. If you’re a mathematician, you must feel like a nuclear bomb hit, and you’re at ground zero.
Reasoning models are two years-old. In that time they went from incapable of basic arithmetic to solving problems humans couldn’t solve for decades.
Math is only the beginning. AI will revolutionize the entirety of the human scientific endeavor. Most of us don’t appreciate what that means.
What a time to be alive!
Quasi-RH?!?!???! Are you kidding me? If a human did this, it would be an instant Fields Medal, no questions asked.
RH says zeta has no zeros in Re(s)>1/2. The best we had until a second ago was a region that got thinner and thinner the higher up the imaginary axis you go. I thought maybe they’d fatten that up a bit, that’d be a massive breakthrough. But no. They got a zero free strip!!!! Insane
Backprop has been the only credit assignment algorithm capable of training large neural nets.
Introducing Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation (bitter lesson!)
- Dust is on the order of 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers.
- We chose pretraining transformers because it's arguably the hardest task possible for zeroth-order optimization.
- Search in high dimensions is really misunderstood: zeroth-order usually gets better with overparameterization, not worse, i.e. at a fixed population size, larger models (up to 120x) reach a lower loss than smaller ones.
- Backprop-like gradients emerge from a large population of activation perturbations, and the alignment holds up at every scale we tested, up to 1B tokens, which is critical for scaling.
The core idea is a *virtual population*, which helps scale to large populations in transformers by perturbing activations in parallel.
w/ @bishmdl76, @cs_serdar, @akshayvegesna
This morning it finally came home to me. I will probably never write another program in C again.
Which is a big deal. I've been writing C since 1983. 43 years is a long time and a lot of skill investment.
It's not like I didn't see this coming. I've been writing about the decline of C for years, predicting that it would be displaced by newer systems languages with better safety guarantees and more coherent designs.
It wasn't better language design alone that did it, though. It's how good LLMs now are at generating code from specs and prompts.
This means I can speak a program into existence in Rust or Go without any more effort than would be required to conjure it up in C. What makes this particularly dramatic is that I never learned to hand-code in Rust. And now - I don't need to.
What a long, strange trip it's been.