Neil Movva (@neilmovva) started his career at Nvidia, working on GPUs and kernels, and has an unusually deep understanding of inference, from software to chips to power.
We spend a lot of time on each of those layers, how they connect, and where the important tradeoffs are.
What makes this conversation special is how detailed it is (like a 401-level class), yet Neil makes it remarkably clear and easy to follow.
Today he runs Sail Research, a company building infrastructure for agents to make tokens as cheap as possible.
We discuss:
- Latency versus throughput
- Why there are no bad chips, only bad pricing
- The end of kernel engineering
- Buying chips and power no one else wants
- New chip architectures
- Nvidia lore + his contrarian view of the company
- Open source and the frontier labs
I learned a ton. Enjoy!
TIMESTAMPS
0:00 Intro
0:38 Building a “Token Factory”
4:21 The Future of Background Agents
13:09 Nvidia and the GPU Stack
23:27 Chips, Memory, and Transformers
36:14 The Future of AI Training Data
44:32 Chip Scarcity and Compute Arbitrage
52:44 Reinventing the AI Data Center
59:01 Power and the “Scavenger Strategy”
1:10:10 Open vs. Closed AI
Robotaxis are very cool but there's probably going to be a weird equilibrium where human drivers have some competitive advantage because they're less strict about following speed limits and traffic laws.
Does it count as work when you're waiting idk +/- 5 minutes for an LLM to respond to a complicated query, multiplied over many prompts? It's kind of hard to use that time productively, especially since sometimes "5 minutes" is 20 minutes and sometimes it's 20 seconds.
Someone asked what I'd do if I were 17. I'd learn how to build LLMs from scratch, and then train ones as powerful as I could with whatever hardware I could get access to.
Frontier Labs are incented to be fearful as it provides a moral justification for staying closed.
Chinese Labs take the consistent position of doing what’s good for all and being open.
So the little guys find themselves in the uncomfortable position of rooting for the Chinese…
As we continue to push the frontier of capabilities while improving efficiency, we're dropping API and credit pricing of GPT-5.6 Sol by over 20% for the next 3 months.
when people spend too much time in a corporation they develop an entire profane LinkedIn idiolect. you need to read a book once in a while to prevent total mode collapse
Rare unironic life advice. It can be great for many people but if you're on the fence about whether to go to some sort of graduate school, i.e. you don't have something fairly specific you're trying to accomplish, probably don't. A much worse value proposition than 20 years ago.
Underrated life advice: Embrace looking ordinary to the outside world. Spend less than you make. Love the same person forever. Drive the simple car. Wear what you like. Show up for friends. Avoid drama. Real success doesn't shout. It's a quiet confidence that speaks the loudest.
Why do Americans, despite being rich, live shorter lives than the people of peer countries?
The answer has little to do with healthcare. It has to do with being fat, violent, and risk-seeking!
Americans shoot more guns, crash more cars, and eat more garbage. Simple as that!
The more open-minded you are, the less likely you are to deceive yourself-- and the more likely it is that others will give you honest feedback. If they are "believable" people (and it's very important to know who is "believable"), you will learn a lot from them. Being radically transparent and radically open-minded accelerates this learning process. It can also be difficult because being radically transparent rather than more guarded exposes one to criticism. It's natural to fear that. Yet if you don't put yourself out there with your radical transparency, you won't learn. #principleoftheday