It's here: We just hit superhuman performance on AI kernel optimization!
Real customer models & production settings. Not toy problems (what I typically see).
This is the year that Claude writes its own kernels, Codex its own kernels, for every new GPU that it wants to run on -- something that takes months to port between GPU generations today.
This has a massive impact to scaling intelligence. More compute means getting the next frontier model sooner.
So excited to welcome @theworldlabs and @drfeifei to the @AMD family! I’ve always been a huge fan of Fei-Fei and her pioneering research in AI. Together, we’ll combine World Labs’ deep expertise in AI and world models with AMD’s compute leadership to power the future of AI and strengthen the open AI ecosystem. Can’t wait for all we’ll accomplish!
1/ Biological threats can bring a country to a halt without a single shot being fired. Yet much of our biodefense infrastructure has barely changed in decades, even as America’s institutions produce extraordinary scientific breakthroughs.
@adayaratna7t7@AMD Realizing that RLing the SW optimization capabilities into the models is super effective so customers switching to us has become trivial for many (already visible) - so for devs, it’ll come down to the HW specs and being reasonable to work with
@anandavati Haha yes except so much has been optimized for AR so my prediction is that there will be further AR optimizations that make this tradeoff than switching (fully) over to diffusion
there is about to be a large wave of algorithms, architecture changes, and performance optimizations that trade more flops for less memory because the relative cost of flops per token to hbm per token is dropping. excited to see what we invent!
Surprisingly not
- if you’re changing your model arch or experimenting with new ones
- if new techniques like flash attention are invented by the community or your agent - could include smaller things like better kernel fusion/unfusing or like it finds a great megakernel for your model
- changing how you’re serving like maybe you decide to disagg something
- if the distribution/shape of user requests change over time eg when agents started becoming a bigger thing, things grew to be more prefill heavy
- adding new types of gpus or compute to better optimize your stack
NVIDIA CUDA and AMD ROCm are software. And software is free. Anthropic just ran Claude for a weekend and it got a version of itself running on GPUs that it's never seen. Every compute customer (frontier lab, neolabs, etc.) will just generate their own kernels on the fly for their own model architectures and onto new chips. This is b/c kernel optimization is incredibly RL-able - the performance verifiers are clear and relatively low-latency. Compute RSI is happening fast.
Not saying that. The capabilities will be native to the major LLMs for anyone to program their GPUs to their custom model arch, kernel shapes, data types/quant, disagg setup, etc. And bc it’s highly verifiable, reliability can be made p high, arguably higher than highly inefficient manual flows today
Anthropic is an AMD customer and will be making ~2GW of AMD GPUs go brrr! (That’s hundreds of thousands of GPUs) This is partly possible bc Claude can write Claude’s own kernels, and will only get better down the stack as many tasks to optimize chips are verifiable
We are excited to announce that AMD and @AnthropicAI are expanding our strategic partnership to accelerate the development and deployment of next-gen AI infrastructure. Tune in at 9:30am PT tomorrow to hear from Dr. @LisaSu from the #AdvancingAI keynote stage!
✅ Up to 2 GW of AMD Instinct MI450 Series GPUs in AMD Helios
✅ AMD has committed to make a strategic equity investment of up to $5B in Anthropic
✅ Deep engineering collaboration across Claude, ROCm and AMD Instinct
More on the news: https://t.co/zA0vQHTT1c
We need better GPU performance benchmarks for RLing capabilities into frontier models for RSI, and I have a few ideas.
The goals would be to
1) help models improve their own compute efficiency (basically “living” RL envs with kernels that matter for new model generations AND new compute generations)
2) push performance with significantly less compute (useful in a compute crunch, also enables higher num iterations within a fixed compute budget; also more doable within existing RL setups)
3) invite more people to contribute to this space (by lowering the compute barrier to entry, including type of compute like not just GPUs/TPUs but even CPUs; but also different people like AI folk in addition to MLSys folk).
The idea is that we need a production-grade performance benchmark at the kernel level, that’s vendor-agnostic. It should measure the most useful kernels for top open weight LLMs, like GEMMs, attention, MoE, but also distributed ML stuff like collectives, so all-gather, dispatch/combine, etc. Make it interesting by including kernels that have strong differences between hardware vendors.
We have increasingly great benchmarks at the end-to-end model level (InferenceX @SemiAnalysis_ and Endpoints @MLPerf), but that takes significant compute to run useful things on, like online RL, or even to reproduce all configurations (precision, types of parallelism, etc) very frequently. Another benefit of kernel level perf, not full model perf, is we can also simulate a kernel on CPU, fairly quickly - so you could scale to different, cheaper compute surfaces.
A kernel level leaderboard can be updated in real time, much faster than the model performance stuff.
Here's a start with AMD AgentKernelArena. I find that other AI researchers gravitate to this general setup. https://t.co/jec492JMnQ A vendor-agnostic version would be fire and an incredibly useful asset to the community. It’d also really lower the compute barrier-to-entry to… chip in. Overall, I think this could pretty dramatically accelerate the field.
Today, I’m excited to formally announce @mirendil with my amazing co-founders Harsh Mehta, Shayan Salehian, and Tara Rezaei!
We’re fortunate to work with @a16z and @kleinerperkins, who led our seed round of $200M, followed by a major investment from NVIDIA, among others.
Mirendil exists to accelerate science and technology, and through them, to help solve humanity's most pressing problems.
Self-accelerating AI R&D is the most direct path to delivering on AI's broader promise, which is why we believe the most important application of AI is AI itself. Get this loop right, and it compounds. It fundamentally changes the rate of progress itself across all domains.
We believe this capability should be democratized. It should be used to power all scientific efforts trying to innovate at the frontier. There are far more important problems—and broader ones—than any single lab can take on, so more groups should be able to pursue them.
This pulls concentration of power away from a few labs: businesses and science labs can own their AI and infrastructure, keep their margins, and control their own destiny instead of ceding it all to a single AI lab.
We’re a small team with a singular focus. Our founding team consists of 20 researchers and engineers from frontier institutions including Anthropic, xAI, Google DeepMind, and OpenAI, united by a passion for science and a drive to build the technologies that move it faster. If you want to build the system that builds systems, join us!
@HarshMeh1a, @shayan_, @tararezaeikh
Such an honor to celebrate with the MIT Class of 2026 today. Congratulations to all the graduates and their families.
MIT has had a lasting impact on my life. Very special to be back on campus today.