Excited to share our new working paper: how much did the US standard of living really rise over the twentieth century — and when? With @bhwittenbrink
Using 5.1 million Sears catalog listings, we build a quality-adjusted price index for consumer goods, 1900–1990. Two headline findings: growth was much larger than official statistics imply, and its peak came before WWII, not after. (1/9)
New paper on data as a driver of automation and growth: https://t.co/WgVBhQJVJG
One of my favorite lines from the Nicomachean Ethics goes ‘for the things we need to learn before we do them, we learn by doing them’. The world is complex and messy… ‘men become builders by building and lyreplayers by playing the lyre’. How did AI systems become taxi drivers? How will they become office workers?
One view is that they’ll gain those skills by training on high quality data. But where will that data come from? What kinds of data will the economy accumulate? How quickly? Who gains and who loses? And if we train AI systems on tons of data so it becomes superhuman at every job, what will be left for us to do?
6 month prediction: agentic coders will use giant models to scaffold, set up, and plan; small model swarms for implementation, review, and verification.
Cursor, conductor, cognition, and other 3rd party harnesses will surge with this pattern, unless Fable-level models truly come down *dramatically* in price.
Few things converging to drive this:
- fable is legitimately great, but too expensive.
- small models are getting really good in short bursts (go try Gemma E2B. It’s insanely good for a 2B model.)
- ensembles are a real pattern with legs (even copilot(!) catches issues in fable PRs)
- ai cost controls are arriving in the enterprise
- Anthropic and OAI shifted enterprise costs to usage based
- teams want to collaborate, and Anthropic and OAI aren’t focused on that pattern
Again, Fable-tier models could get super cheap and fast and surprise us all. But the current crop of small models are *magnitudes* cheaper; a good harness could unlock their potential.
Super excited to share joint work with @axiommathai that kicks off a broader project of formalization in economics.
Aumann's celebrated theorem says we can't "agree to disagree."
But what does that actually mean – formally? 👀
If someone with a demonstrably excellent forecasting record makes predictions that seem outlandish to you, the correct response is to update your beliefs, not mock them.
🚨Transformers don't learn Newton's laws? They learn Kepler's laws!
Like us, transformers don't predict a flying ball via a differential equation, but by fitting a curve.
Moreover, reducing context length steers a transformer from Keplerian to Newtonian. Compression in play.
Finding myself going back to RSS/Atom feeds a lot more recently. There's a lot more higher quality longform and a lot less slop intended to provoke. Any product that happens to look a bit different today but that has fundamentally the same incentive structures will eventually converge to the same black hole at the center of gravity well.
We should bring back RSS - it's open, pervasive, hackable.
Download a client, e.g. NetNewsWire (or vibe code one)
Cold start: example of getting off the ground, here is a list of 92 RSS feeds of blogs that were most popular on HN in 2025:
https://t.co/dwAiIjlXet
Works great and you will lose a lot fewer brain cells.
I don't know, something has to change.
So Moltbook seems like "tokens pretending to be Reddit". But I don't know - why shouldn't an organization have a forum for AI agents to share information then propose structural changes or new skills? This is effectively a context management trick that makes a ton of sense to me.
New paper on bank runs with Correia and Luck:
"Bank Runs With and Without Bank Failure"
Questions:
- What are the determinants of runs?
- When do bank runs result in bank failure?
- Can runs trigger the failure of healthy banks and amplify small shocks into large crises?
- Are runs themselves the initial cause of financial distress or are they a symptom of deeper fundamental solvency problems in the financial system?
What we do:
- Apply LLMs to historical newspapers to uncover over 4,000 runs on individual banks in the pre-FDIC US banking system from 1863 to 1934. Capture the most famous runs (Bank of the US
- Merge data on runs and other bank-level events discussed in newspapers (suspensions, failures) to bank-level fundamentals (harder than it sounds!)
What we find:
(1) Runs are considerably more likely in weak banks, but can also occur in strong banks, especially in response to negative news about the real economy or the broader banking system.
(2) However, runs typically only result in failure for banks with weak fundamentals [see figure below]. Strong banks survive runs through various mechanisms, including interbank cooperation, equity injections, public signals of strength, and suspension of convertibility
(3) At the local level, poor fundamentals necessary for runs to translate into large declines in lending. Moreover, bank failures (with and without runs) translate into substantially larger declines in deposits and lending than runs without failures.
Overall takeaways:
- Poor fundamentals are key for whether runs pass through into failure and have severe consequences for the broader economy.
- The findings temper the view that small shocks can result in large jumps to bad equilibria via runs on demandable debt.
Full paper here. Comments welcome. Given the methodology and evolving AI tools, we expect to make refinements to the runs database over time. Any input is welcome.
https://t.co/mbqoFf5g03
A funny thing is happening: the more I build with agents, the less I want to use Python. I explore this in my latest "From Human Ergonomics to Agent Ergonomics"
https://t.co/ybTacPtuc6
The answer to this is likely the same as it always was: define your problem well and then allow an optimizer to find the best prompt/context/structure. You can @DSPyOSS CLI agents just as well as you could basic LLMs.
With the advent of workable AI agents, we are sort of back to square one on prompting/context engineering. Lots of opinions but there is no clear evidence on the right way to prompt agents, how they should be orchestrated together & how to organize and structure very long tasks .
Can't recommend this @BerenMillidge talk enough - super clear, interesting, and well put together. I'm really happy this frame is becoming increasingly popular in the discourse, although as Beren notes it's still highly neglected/underappreciated. https://t.co/RPQxryaS9F
The tl;dr is that the prevailing pessimistic view in rationalist discourse - AI monotheism, i.e. the singleton view of the world - is probably wrong. Instead competition/cooperation among many AIs (AI polytheism) is more likely and may well preserve human-like values rather than destroying them. Competition doesn't actually produce homogeneous fitness maximizers, and cooperation itself is a powerful competitive strategy, which is why pro-social values evolved in the first place.
It's great that liberal democracy is mentioned too - can't stress how important I think this is both today and in the future, and why it's relevant to the normative part of the 'alignment problem'. Once you internalize this, it becomes more evident why 'aligned vs not aligned' is an impoverished frame. Imo liberal democracy is not just an arbitrary ideology but actually one of the best evolved solutions to the game-theoretic problems of cooperation and cohabitation. (This of course doesn't mean it can't be improved!)
This frame resonates with recent work I've been exploring too:
- The economics side of things, e.g. Hayekian implications for knowledge, institutional economics etc. This relates to questions about the nature and role of the 'firm', scarcity, transaction costs - I explore one such angle in my Coasean Bargaining piece: https://t.co/BfI2mk9stH. This is also why I'm interested in Levin's work on multi-scale cooperation/alignment - principles and abstractions from other disciplines can be very useful here! @BenjaminLy61243 writes a lot about this.
- On the safety side, see also this paper some colleagues and I published recently: https://t.co/3DrGXFPthD - where we explore the implications of patchwork AGI hypothesis (AI polytheism!) for safety. To me this is a positive update: there are challenges to cooperation ofc (as we see with humans!) but this is something we can continually iterate and perfect over time rather than needing to solve on the first try. And as Beren notes, AIs might be much better at cooperation than humans (higher bandwidth communication, better capabilities, etc).
- It's not enough to just care about harm/catastrophe avoidance with AI; the other half of the coin is also working out what we are building towards. It's great that there's more interest lately in specifying positive post-AGI futures! This is important and not just feelgood sci-fi. And of course this is where a lot of the existing political/social sciences, philosophy, and economics literature/expertise becomes very relevant. More to come on this front!
Institutions and trust in the rule of law are important; they enable a cooperative equilibrium that has benefited the U.S. greatly. The current administration has demonstrated a repeated willingness to betray this trust through its handling of our allies overseas, its citizens’ rights via DHS, and the Fed. We should be doing everything in our power to prevent this cancer from spreading further.
The cultural brain hypothesis is a really useful model of the world -- jagged intelligence is good for society as a whole. Looking forward to the essay!
To @dwarkesh_sp point in the substack discussion, human intelligence is jagged as hell!
Here is a preview of some data from an essay post I'm writing for why organizations--which have evolved over thousands of years--make us forget about this jaggedness.
Can AI "learn" economic states, addressing the Lucas Critique?
With @alexolegimas we simulated data from an NK model, fit a transformer, and tested out of sample fit
It generalizes surprisingly well. We hope this stimulates discussion and future agendas
https://t.co/lXcJh9IkE9
As an AI faculty now surrounded by economists at Booth, the discussion on scaling and the Lucas critique has been entertaining to watch. IMO the spirit of @arpitrage's claim feels right but double descent isn't the best analogy. Instead, we should look to language understanding:
How much does intelligence cost? How concentrated is the AI market and is it winner take all? When prices fall, how does demand change and is there a Jevons effect? These questions matter, but actual market data has been hard to come by. We use data from Microsoft Azure and OpenRouter to find out.
Details below.