Did a very different format with @reinerpope – a blackboard lecture where he walks through how frontier LLMs are trained and served.
It's shocking how much you can deduce about what the labs are doing from a handful of equations, public API prices, and some chalk.
It’s a bit technical, but I encourage you to hang in there - it’s really worth it.
There are less than a handful of people who understand the full stack of AI, from chip design to model architecture, as well as Reiner. It was a real delight to learn from him.
Recommend watching this one on YouTube so you can see the chalkboard.
0:00:00 – How batch size affects token cost and speed
0:31:59 – How MoE models are laid out across GPU racks
0:47:02 – How pipeline parallelism spreads model layers across racks
1:03:27 – Why Ilya said, “As we now know, pipelining is not wise.”
1:18:49 – Because of RL, models may be 100x over-trained beyond Chinchilla-optimal
1:32:52 – Deducing long context memory costs from API pricing
2:03:52 – Convergent evolution between neural nets and cryptography
@webdevMason Choosing a destination as a means of setting shorter term priorities is a bad idea. It makes it hard to discover your mistake until you're dead.
Hash out what you want NOW—much rarer & harder than it sounds—and think about the long term as one of many ways of criticising that.
Imagine a future where you can REMEMBER EVERYTHING.
Every email, every person, every conversation.
Introducing Engramme.
Our vision is to endow humans with perfect and infinite memory.
All your memories come to you. No more searching or prompting.
https://t.co/LX22KMG2lH 🧠