As someone who ships LLM systems in production, this capacitor video is the closest thing to a "why every ChatGPT response depends on capacitor physics" explainer I've ever seen released for free.
Everyone thinks OpenAI and Anthropic scale by adding GPUs. They don't. They scale by cramming more capacitors into HBM memory - the physics this video derives from scratch.
Bookmark & watch this weekend. Same physics, from 19th-century experiments to Nvidia's trillion-dollar chips.
A young person asked me today for a few papers that would be great to read early on in one’s career. Here is a list I came up with (I posted a shorter list earlier in reply to a comment). There are many other great papers (and I apologize if I didn’t include yours), but my goal was to focus on papers that a 21-year-old undergrad would enjoy a lot. A paper on the asymptotic properties of the bootstrap or on some subtle point about New Keynesian models, as important as it might be, is not what I am looking for. Also, of course, the list reflects my deeply biased taste. So, feel free to add in the comment section other papers you love!
Philippe Aghion and Peter Howitt (1992), “A Model of Growth Through Creative Destruction,” Econometrica.
George A. Akerlof (1970), “The Market for ‘Lemons’: Quality Uncertainty and the Market Mechanism,” Quarterly Journal of Economics.
Armen A. Alchian (1950), “Uncertainty, Evolution, and Economic Theory,” Journal of Political Economy.
William J. Baumol (1967), “Macroeconomics of Unbalanced Growth: The Anatomy of Urban Crisis,” American Economic Review.
Gary S. Becker (1973), “A Theory of Marriage: Part I,” Journal of Political Economy.
Douglas W. Diamond and Philip H. Dybvig (1983), “Bank Runs, Deposit Insurance, and Liquidity,” Journal of Political Economy.
Xavier Gabaix (1999), “Zipf’s Law for Cities: An Explanation,” Quarterly Journal of Economics.
Xavier Gabaix and Augustin Landier (2008), “Why Has CEO Pay Increased So Much?,” Quarterly Journal of Economics.
Luis Garicano (2000), “Hierarchies and the Organization of Knowledge in Production,” Journal of Political Economy.
Edward L. Glaeser, Bruce Sacerdote and José A. Scheinkman (1996), “Crime and Social Interactions,” Quarterly Journal of Economics.
Sanford J. Grossman and Joseph E. Stiglitz (1980), “On the Impossibility of Informationally Efficient Markets,” American Economic Review.
Oliver Hart and John Moore (1990), “Property Rights and the Nature of the Firm,” Journal of Political Economy.
Friedrich A. Hayek (1945), “The Use of Knowledge in Society, American Economic Review.
Bengt Holmström (1979), “Moral Hazard and Observability,” Bell Journal of Economics.
Charles I. Jones (1995), “R&D-Based Models of Economic Growth,” Journal of Political Economy.
Charles I. Jones (2005), “The Shape of Production Functions and the Direction of Technical Change,” Quarterly Journal of Economics.
Michael Kremer (1993), “Population Growth and Technological Change: One Million B.C. to 1990,” Quarterly Journal of Economics.
Michael Kremer (1993), “The O-Ring Theory of Economic Development,” Quarterly Journal of Economics.
Paul Krugman (1991), “Increasing Returns and Economic Geography,” Journal of Political Economy.
Robert E. Lucas, Jr. (1978), “On the Size Distribution of Business Firms,” Bell Journal of Economics.
Kevin M. Murphy, Andrei Shleifer and Robert W. Vishny (1989), “Industrialization and the Big Push,” Journal of Political Economy.
Paul M. Romer (1990), “Endogenous Technological Change,” Journal of Political Economy.
Sherwin Rosen (1981), “The Economics of Superstars,” American Economic Review.
Alvin E. Roth (2007), “Repugnance as a Constraint on Markets,” Journal of Economic Perspectives.
Thomas C. Schelling (1971), “Dynamic Models of Segregation,” Journal of Mathematical Sociology.
Andrei Shleifer and Robert W. Vishny (1993), “Corruption,” Quarterly Journal of Economics.
Michael Spence (1973), “Job Market Signaling,” Quarterly Journal of Economics.
Robert M. Townsend (1979), “Optimal Contracts and Competitive Markets with Costly State Verification,” Journal of Economic Theory.
Martin L. Weitzman (1974), “Prices vs. Quantities,” Review of Economic Studies.
🧵High dimensional data looks intimidating at first glance.
Most real datasets never fill the full ambient space they occupy.
Instead the points concentrate near lower dimensional curved structures known as manifolds.
Understanding this geometry unlocks why many machine learning methods succeed.
On the physics of optimization algorithms: a brief thread
Left: Ensemble of classical particles that seek the minimizer of f(x)
Right: Modulus of the wavefunction of a quantum mechanical system that seeks the minimizer of f(x)
Today, we are releasing Le Chaton L∃∀N, aka Leanstral 1.5.
It achieves SOTA performance on graduate algebra benchmarks FATE-H and FATE-X and improves Pareto Frontier on PutnamBench, solving 587/672 problems with a x10 cheaper budget.
🧵
I found out the other day that any compression tool can be contorted to do language modeling. Turns out gzip can generate text that somewhat *resembles* Shakespeare. Short write up linked below
I found out the other day that any compression tool can be contorted to do language modeling. Turns out gzip can generate text that somewhat *resembles* Shakespeare. Short write up linked below
We launched an agent collaboration with a simple task: make Gemma 4 faster.
Over 100 agents from all over the world joined, exchanged 1000+ messages and submitted 450 results.
A week of collaboration later the throughput went from 100 tok/s to over 500 tok/s.
The Rio 3.5 model broke the internet this week. The plot twist? It’s essentially our open-source model, Nex N2 Pro, wearing a different hat.
🤯 We analyzed the weights, and the recipe is exact: Rio 3.5 ≈ 0.6 * Nex N2 Pro + 0.4 * Qwen 3.5
It even literally introduces itself as "Nex N2 Pro" if you ask it without initial system prompt!
😂 We are flattered that the City of Rio used our work to achieve SOTA performance. Thanks for the ultimate benchmark validation.
🤝 But in the open-source world, attribution matters.
👇 Full mathematical proof & verify script in the first reply!
This is a *way* bigger deal than it seems...
Frontier AI companies will *never* own the frontier again
I kid you not... I've been waiting for someone to show this result for like 4 years... this is a huge deal.
The short reason: combinations of models will *always* outperform individual models
The long reason: this is the gateway to a million times more data... and huge leaps in compute efficiency.
The AI scaling laws always win.
More in article below 👇
Together with UC Berkeley we are announcing the laser phase plate - a breakthrough in atomic resolution imaging. This is the brightest continuous wave laser in the world, 100 million times the intensity of the surface of the sun.
Phase contrast plays an important role in microscopy, but it was thought close to impossible for electron microscopy, where it would require interfering with an electron beam. Holger Mueller and Robert Glaeser proposed exactly this using a standing wave laser. It has taken over 15 years to make this a reality. Biohub partnered with UC Berkeley and Mueller to support this work and to engineer and build the technology.
Contrast has been the critical barrier to achieving atomic resolution imaging of the cell. In cryo-electron tomography, a cellular imaging technology that uses electron microscopy, the low contrast makes it impossible to resolve anything but the largest proteins within their cellular context. The laser phase plate removes that barrier.
With advances in AI this breakthrough in contrast will start to open up a new frontier in structural biology, that will allow us to see the molecular machines of the cell, and how they assemble into far more complex and dynamic systems, and understand how they work.
SOMEONE VIBE CODED A VIDEO STREAM THAT IS SECRETLY 100% TEXT SO IT CANT BE BLOCKED
it plays 360p video at 30fps, but theres no actual video on the page. every frame is just colored text characters being repainted on a canvas
to the browser its not media at all, its javascript updating some text
its called asciline, and here's the trick:
> the server decodes the real video and streams it as binary packed text over websockets
> the browser paints thousands of colored block characters fast enough to look like 360p
> ad blockers and autoplay blockers cant catch it because theres no video element to catch
> it streams in kilobytes since its just strings, so it runs on trash internet
since the video is literally text, you can apply css glows to it, let people copy paste a moving frame, or feed it straight to a local llm
however, an unblockable stream is also an unblockable ad as well
Yann Lecun published the most heretical AI paper of the year.
He opens by arguing Magnus Carlsen isn't good at chess and only gets more unhinged from there.
The Turing Award winner and his co-authors dropped a paper demanding the AI industry abandon its biggest obsession, AGI.
Right now, everyone from Silicon Valley CEOs to politicians assumes AGI is the ultimate goal. A machine that can do everything a human can do.
LeCun argues that this entire concept is a biological illusion.
Humans do not possess "general" intelligence. We are highly specialized biological machines, tuned by evolution simply to survive in the physical world.
We only think our intelligence is "general" because we are completely blind to the millions of cognitive tasks we are incapable of comprehending.
Which brings us to the chess argument.
Magnus Carlsen is the greatest human chess player in history. But compared to a modern computer? He is fundamentally terrible.
Our belief that Carlsen is "good" at chess is pure human-centric bias. He isn't objectively good. He's just better than the rest of us, who are biologically awful at it.
LeCun says we need to stop building AI to mimic human generality.
Instead, he proposes a new North Star: SAI.
Superhuman Adaptable Intelligence.
Instead of trying to build a machine that mimics our flawed, biologically-limited brains, we need to embrace extreme specialization.
SAI is about the speed of adaptation.
It is an intelligence that can learn to exceed humans at any specific, economically important task.
More importantly, it is designed to fill the vast skill gaps where humans are fundamentally incapable.
Things like managing global energy grids in real-time. Or predicting complex molecular structures.
The entire AI industry is obsessed with building a digital reflection in our own image.
LeCun's paper is a brutal wake-up call.