Did a very different format with @reinerpope – a blackboard lecture where he walks through how frontier LLMs are trained and served.
It's shocking how much you can deduce about what the labs are doing from a handful of equations, public API prices, and some chalk.
It’s a bit technical, but I encourage you to hang in there - it’s really worth it.
There are less than a handful of people who understand the full stack of AI, from chip design to model architecture, as well as Reiner. It was a real delight to learn from him.
Recommend watching this one on YouTube so you can see the chalkboard.
0:00:00 – How batch size affects token cost and speed
0:31:59 – How MoE models are laid out across GPU racks
0:47:02 – How pipeline parallelism spreads model layers across racks
1:03:27 – Why Ilya said, “As we now know, pipelining is not wise.”
1:18:49 – Because of RL, models may be 100x over-trained beyond Chinchilla-optimal
1:32:52 – Deducing long context memory costs from API pricing
2:03:52 – Convergent evolution between neural nets and cryptography
He was Satyendra Nath Bose, an Indian physicist whose quiet brilliance in the 1920s forever altered our understanding of the quantum world.
In 1924, Bose, then a 30-year-old professor in British India, sent a groundbreaking manuscript directly to Albert Einstein. The paper offered a novel, more elegant derivation of Planck's law for blackbody radiation by treating light quanta (photons) as indistinguishable particles—a radical departure from classical statistical methods. Impressed by its insight, Einstein personally translated the work into German and facilitated its publication in the prestigious Zeitschrift für Physik.
This exchange sparked a brief but profound collaboration. Einstein extended Bose's statistical approach to material atoms, predicting a bizarre new state of matter at ultra-low temperatures: what we now call a Bose-Einstein condensate (BEC), where particles behave as a single quantum wave. Bose's original framework became known as Bose-Einstein statistics, and the class of particles that obey it—those with integer spin, including photons, gluons, W and Z bosons, and the Higgs boson—was later named bosons in his honor by Paul Dirac.
Unlike fermions (matter particles like electrons), which obey the Pauli exclusion principle and cannot occupy the same quantum state, bosons can pile into identical states en masse. This "social" behavior underpins extraordinary macroscopic phenomena: the coherent light of lasers, the zero-resistance flow in superconductors, and the collective quantum coherence in BECs.
Despite the monumental impact—his statistics describe half of all fundamental particles and enabled key advances in quantum field theory, condensed matter physics, and particle physics—Bose remained remarkably unassuming. He continued teaching at universities in Dhaka and Calcutta (now Kolkata), mentored students, pursued ideas in X-ray crystallography, unified field theory, and other areas, and never sought the spotlight. Nominated several times for the Nobel Prize (notably for Bose-Einstein statistics and his later work), he was never awarded it, and his name rarely appears in popular accounts of 20th-century physics.
There's a poignant humility in his story: a man whose legacy literally names one of the two fundamental families of particles in the universe, yet whose personal fame never matched the scale of his contribution. Bose reminds us that true influence often arrives without fanfare. Some breakthroughs echo through textbooks and technologies, while their creators work in the background, content to let the universe carry their ideas forward—even if history's spotlight rarely finds them.
@parmita@AllanatrixQ If you think about his goal to be on Mars, you will realize to build things initially on Mars, he won't rely on humans and hence cameras only logic works.
I just came off of a flight from San Francisco to New York.
Midway through the flight one of the engines started to fail due some unknown technical issues.
The pilot was preparing us for an emergency landing in Gary, Indiana.
I pulled out my maxed out MacBook Pro M4, fired up Claude Code, connected to Opus 4.5, and instructed it to sniff out the ports on the internal airplane communication network.
In less than a minute it had full access to all the critical systems.
I instructed it to diagnose the problem with the engine and find a fix.
It reasoned for a few minutes, coded a simple firmware patch update for the engine, and uploaded it to the airplane mainframe.
Before too long the captain was announcing that all the engine issues have been resolved and we will continue on our journey.
We are definitely in the hard takeoff. Like literally.
This is a good question. Distillation is primarily at the pretraining stage. The gains we see with Flash 3.0, sometimes surpassing Pro 3.0, is very likely due to new post-training innovations. For reasoning benchmarks, such as ARC, post-training matters.
Deep Research is the first agent released on the new Interactions API – offering a single endpoint for agentic workflows.
Start building today ↓
https://t.co/FlgtzbDYj7