New technology arrives first, then a long awkward period where it's being funded by instruments designed for the past, then someone figures out the right structure, and then the real expansion begins. Who will solve it for AI?
https://t.co/38HgmHE1XL
Alright theory is this. We basically never see the benchmark we actually want, “cost to achieve a unit of economically productive work” because the people who build this stuff best come from either an academic background or safetyist one, where cost at scale fairly unimportant.
The common story about private market investing is that it’s about companies, entry multiples, process, and stage; and while true...
...I’d argue that private market investing has actually been about finding breakthroughs in how to think about the balance sheet.
The first chapter of private market investing was on the liability side, when Henry Kravis "invented" the LBO. By expanding the liabilities a single company owed to a lender, Kravis was able to take companies private for little cash down, and provide immense value to the equity, and the GP in doing so. That vintage of firm was KKR, Apollo, and the OG PE firms.
The second chapter was on the asset side, when Mitt Romney and Bob White focused on taking a firm from its “operating enterprise value” to its “strategic enterprise value”. By leveraging their leanings from serving other companies as consultants at Bain, Bain Capital pioneered the idea of providing value directly to the operating asset side of the business by bringing expertise, best practices, and industry-wide ways of doing things that could transform companies. That "playbook era" vintage gave us Vista Equity Partners, Thoma Bravo, and the current top dogs.
We’re now entering the third chapter - where technocapital leverage takes firms from their strategic enterprise value, to their contextual enterprise value - taking advantage of both sides of the balance sheet, and building on the breakthroughs that came before.
To fund this chapter though, we need to invent better coordination technologies: https://t.co/E2fDAWN3eX
Why does this logic stop at the level of the firm and not extend all the way to the employee, who, at some point will have unique knowledge of how to orchestrate agents that is illegible to tools like Foundry and whatever $MSFT sells you?
Where's their sovereignty?
Vibe Coding, Prompt Engineering, and Harness Design are all Managerial Primitives
The architecture of directing agents has seen an improvement in capabilities with a handful of breakthroughs:
- Chain-of-Thought (CoT) Reasoning
- Test-Time Compute
- Org Harnesses (e.g. @SierraPlatform's Pinecone)
- Evals
- Steering
- Autoresearch
- Loops (cc @satyanadella and @JayaGup10's posts)
- .MD files and skills (effectively standard operating procedures)
The irony is that all of these things are digital derivatives of people-focused managerial science.
You don’t really explicitly tell people what to do, you need to steer them, coach them, repeat things multiple times, have ways of checking their work, give them performance review gates, federate work to more or less capable people, and on, and on.
Instructing and managing humans at work is mostly vibe-coding. The entire art and science of management is a mix of practices that seek to deal with the probabilistic, moody, and emotional outputs that humans are known for.
The more explicit that you can outline a task for a human to follow, the less economically valuable it is.
The inverse is also true - what exactly would you tell a CEO/CMO/COO to do every single day if you were to program their every minute? There is hidden complexity in their role that you don’t understand, and trying to make their tacit knowledge explicit would miss the most important parts of their role; the parts that give you leverage and ultimately create enterprise value.
You can’t formally code the best CXO. So too is the case with vibe coding.
The highest value work - in terms of economic output (using compensation tables as a comparison) - is necessarily vague and hard to specify. The best way to get the optimal behavior out of an elite executive is to incentivize, not to direct.
It follows that the most important areas in the economy in which to deploy AI look much more like nudging models in a sandbox of incentives rather than a comprehensive set of skills, .md files, and other forms of formalistic cope. The harness is useful insofar as it can capitalize on the emergent “ghost” (h/t @karpathy) behavior of undirected algorithms.
Vibe-coding, prompt engineering, and harness design should be better understood as an evolution of management science, NOT programming.
The next agentic coding breakthrough is more likely to be sitting under some dust in the GSB library on Knight Way than in OAI's lab.
Taylorism’s quantification of process, and James O. McKinsey’s view of doing this across industries, are precursors to how best to think about encoding business leverage into fine-tuned model weights. A similar dynamic is likely to be the case for everything from self-driving cars to self-operating companies.
Silicon Valley is currently short business school academics.
You should probably be long.
Everyone at Sierra uses an internal agent called Pinecone to automate 90% of our coding, analytics, and busywork. I can't go back to any other way of working.
Pinecone:
* Runs an agentic harness in our internal cloud.
* Talks to all our tools (Slack, Github, Linear, GSuite, Clickhouse, Gong, etc.)
* Lets you pick any model. Especially useful this week!
* Can spin up a local version of all of Sierra to test its work.
* Is accessible over mobile, Slack, Github, and Web and is collaborative.
@neilrahilly and team cooked on this. In this post, they share lessons learned from building and scaling Pinecone to the whole company.
Reality has a surprising amount of detail.
The only way to get the economically valuable, out-of-distribution data that is constantly dynamic, changing, and worth training on - is to run the enterprise yourself.
There is heterogeneity inherent in the day-to-day reality of running a firm that will constantly be throwing off signals that’s worth training off of. In that way enterprise, science, and physical data are closer to video rather than static text.
There are no true “data” companies, because data without context - or the environment it's generated in - takes on no rivalrous properties. Compiling data like Scale or Mercor are half-attempts.
It’s better to think about data points as one chain in a sequence of a larger chain that we’re trying to feel our way in the dark to learn the shape of. Learning about the cell tells me nothing about the reality of the arm, anymore than knowledge about the arm tells me the full truth of the body.
For the labs to break through each level of scaling plateau, they’re going to have to push closer and closer to seeing the work happening in real time.
This is their entire product roadmap, and every breakthrough came through a new environment/surface innovation - NOT better algorithms: Chat > IDE w/ Claude Code > Now Claude Design and Claude Science.
If you are not the substrate upon which work is happening, you are training off of economically useless derivatives. The fidelity of the verified-rewards that you need to train off of is as important as your ability to get out-of-distribution data to RL on.
The best way to do this is to own the entire company that allows you to see all the tacit, explicit, and embodied knowledge interacting with each other. Territory, not map.
If you believe in scaling laws, then you also have to believe that we’re returning to the conglomerates of the ‘80s - DuPont, GE, Jack Welch - multiple domains, with multiple projects, touching different parts of the economy. All to be consumed and trained on. This demands different talent, skills, and recruiting than SV has historically excelled in.
Expect incredibly large companies to centralize compute, access to heterogeneous data, and compute applied to recursive intelligence to grow operations. We’re already seeing a hint of that with SpaceX.
Focus on data labeling, bits with no context, and collective data-pooling efforts at your peril.
A Stargate for Data
Labs are on a trajectory towards >$100B/year of data spend by 2030. As we begin the trillion-dollar compute project, we need to think about the equivalent civilizational-scale effort for the other core ingredient: data.
At the foundation of the scaling revolution is a simple empirical law: deep neural networks improve smoothly, near magically, as you scale two things in proportion — (1) the size of the model and (2) the amount of data you train on. And despite the scaling laws being brutally diminishing, we’ve successfully bitten the bullet of logarithmic scaling with exponentially larger clusters and datasets, and received incredible new capabilities in return.
But this exponential scaling is bound to hit some limits. Oddly enough, compute has compounded fairly smoothly without limit, with trillions flowing into hypercluster buildout. Instead, we’re starting to hit the limits of an exponential demand for data. Gone are the days of being purely in the compute-limited regime, where we had effectively infinite internet data but never enough GPUs, we’re now entering a data-limited regime.
Luckily, this limitation is coinciding with staggering improvements in AI capabilities. Incredibly, we seem to have a real line of sight towards automating a majority of knowledge work with the methods we have today. RL + pretraining, and the data for each, will be generally sufficient to achieve most economically valuable tasks, given some minimal algorithmic progress and continued compute scaling.
In a data-limited world, economic progress & scientific acceleration will be directly bottlenecked by our coverage in each domain. We need to see data collection as imperative, deserving the same civilizational ambition we’ve given compute.
The internet as a one-time subsidy
It’s underrated how much all progress in AI owes everything to the blessing of the internet, this one-time civilizational subsidy to deep learning, decades of unintentional accumulation of a perfect dataset: every book, blog post, image, video, paper, discussion, etc. all digitized and freely available. Without the internet, we’d likely see comparably minimal progress in AI today, and in fact, if you notice where systems currently underperform, it’s almost always a domain where web coverage is limited and data is private, expensive, non-digitized, or non-existent.
But we’re running out of it. There are only about 300 trillion tokens of useful public human text, and the internet doesn’t produce nearly enough new high-quality data to match what scaling demands — we’re soon to hit the limits of public data for pretraining. And though the advent of RL bought us reprieve — chain-of-thought RL needed a new form of untapped data, gradable math & coding tasks, also available online — we’re quickly running dry of hard tasks for RL as well.
Why do we need so much data anyways? Humans learn comparably in far less time, needing just one textbook where language models might need the equivalent of hundreds to learn a new topic. It’s possible we discover methods that are massively more data efficient — synthetic data, data efficient architectures, other exotic algorithms — but fundamental progress is slow and highly unpredictable, and the recipe we have just works today.
And, while I’m wary of getting too deep here, even arbitrary data efficiency can’t replace data that just doesn’t exist in the first place. There’s a massive amount of missing information on the web: the dark matter of the internet — tacit knowledge, undocumented processes, etc. — most of which was never published and lives only inside organizations, the physical world, or just in people’s heads. I’ll leave it here and say, for reasons far longer than I can fit in this post [1], it’s best to operate on the assumption that our insatiable desire for data will continue as it has for the last decade.
There will be >$100B/year in data spend by 2030
We’re not screwed yet, of course. Only a fraction of useful data in the world is on the public internet, the rest is stored inside private datasets, corporations, personal archives, universities, governments, and otherwise. Labs can and will continue to license these private datasets, or create them from scratch, like Anthropic’s book scanning project. And we’ll increasingly task human experts to manufacture new high-quality data, with a large fraction of hard RL training tasks already being sourced this way.
But collecting this data, unlike before, will be expensive. As the free internet dries up and demand for data rises, we should see labs investing equally in data as compute, likely spending a significant fraction of their compute budgets on data. As we see trillions spent on compute, we should also expect hundreds of billions spent on data (human data & collection budgets), given their equivalent importance. And, notably, data spend is already tracking this way: total data spend across vendors, not counting internal lab efforts, is already roughly $7 billion per year. It’s quite reasonable we’ll see >10x by 2030.
Data is the moat
Data becoming increasingly private will also majorly shift the competitive landscape. While compute is a commodity — everyone buys the same chips and builds the same clusters — data really isn’t. The big reason why frontier models have felt eerily similar to one another, until now, is they were trained on substantially the same internet (pretraining data variability across labs seems pretty low). As labs diverge onto more exclusive, manually collected corpora, I think models will begin to increasingly diverge.
OpenAI pulling ahead in mathematics and Anthropic in cybersecurity isn’t an accident. I really think laser-focused collection of high-quality midtraining tokens, custom RL tasks, environments, with dedicated research effort, has driven much of the visible progress in the last year. James Betker has an excellent blog about “the ‘it’ in a model is the dataset”: model architecture and compute buy you efficiency and order-of-magnitude performance, but ultimately, models, of any architecture, are such incredible approximators of their dataset that the core meat of a model boils down to just that, nothing else. Data is a major moat.
AGI long, ASI short
As I’ve tweeted before, I’m confident that, despite the narrative, the data labeling industry will continue to fuel great businesses and be an excellent AGI long, ASI short. The argument is just: By the time the AGI labs no longer need data, it’s probably over for everything else too [2]. In this frame, the last companies left should be the data companies, as the last speck of economically relevant data is sucked in. And these companies are already among some of the fastest-growing companies in history: Mercor, founded three years ago, is rumored to be doing $2 billion in revenue with something like a few million expert labelers under contract.
While these businesses are very non-stationary, what type of data is needed shifts constantly, I don’t think that diminishes their value. The long-tail of the economy is long, and the value isn’t diminishing as you extend farther into more obscure information: as models get more capable, the value of the marginal dataset goes up, not down. Automating a full job means covering its full distribution of tasks, tools, edge-cases, and long-horizon loops. There’s some O-ring logic to it: a dataset that buys a 1% bump can justify a previously unjustifiable collection cost when it’s the difference between a system that does 99% of a job and one that does all of it [3].
The competitive dynamics of the data industry are still evolving but as demand for data is increasingly niche, ultra high-quality, expert-generated, I think we’ll see real consolidation. Again, contra-narrative, we’ll probably see true competitive differentiation built on brand, quality control of data (which, from personal experience, can vary massively), as well as in network effects from the talent networks themselves over time. We’ve already seen rapidly shifting data type demand work in favor of incumbents, benefiting those with early knowledge of where the market is headed.
The binding constraint
It’s truly remarkable that we seem to have the recipe — pretraining + RL — to absorb most economically valuable work, despite being far from a lot of what we expected from “AGI”. The same way chess engines revealed we never needed general intelligence to solve chess, as we originally thought, we’ll soon realize that software, mathematics, and the vast majority of the economy (including physical, just running ~3 years behind!) are the same. If recursive self-improvement or some other algorithmic breakthrough arrives, that’s wonderful, but we really don’t have to wait for it. The binding constraint between here and an automated economy isn’t that, it’s data coverage: every app, workflow, edge case, process, etc. sitting in private stores or someone’s head.
Ultimately, while we make tremendous strides in more efficient model architectures, and clusters like Stargate equip us with zettaflop-scale compute, we really aren’t making rapid progress collecting the data we lack.
We’ll soon live in a world where we have the methods & compute to accelerate scientific progress or economic growth, but not the data. And we’re already there today: frontier models would surely be as good at accounting/many medical tasks/legal advice as they are at software engineering if we only had the same pretraining & RL coverage as we did for code.
I really want to drill this in: The speed at which we automate the economy is going to be directly rate-limited by our ability to collect data about it.
Worth noting that under this assumption, with data as defensible and directly proportional to economic & scientific progress, data should also be considered a national strategic asset like compute. Imagine what we’d do in a world where we had a Manhattan Project-effort for AI and needed to mobilize data collection as a limiting factor. We should be concerned about China, with greater state capacity and authoritarian economic control, being capable of mobilizing data collection at national scale, potentially compounding their economy and scientific output faster than us down the line.
A Stargate for data
I’m leaving my complete ideas for a future post, as this one is already far too long, so I’d really like to pose the question here. Stargate exists because we organized trillions of dollars, international strategy, gigawatts around compute as a fundamental ingredient. What would equivalent ambition look like for data?
Obviously, scaling data collection, a heterogeneous mass of information across the economy, isn’t going to be as clear as scaling compute, as a homogenous infrastructural effort. A core division will be first, coverage — all uncaptured knowledge sitting across the economy/science/physical world and all that simply isn’t recorded — and, secondly, sheer volume in the domains we already train on: more hard math tasks, more high-quality web text, way more coding data, more legal drafts, etc.
I have a post coming soon which breaks down my proposals. There’s a lot of room for creativity. Quickly, we’ll probably want to start with a deep census of what we have and what we’re missing, predict what the 2030 model will still be bad at and work backward to what we should be collecting today. You can probably license a large amount, leveraging high lab valuations to buy datasets or companies altogether. There’s an adversarial nature to a lot of this collection with firms, so there’s lots of engineering to do this correctly. We should go convince important companies to turn off deletion policies, even if we’re not buying from them yet. Data flywheels in consumer products will be massive. Confidential training, government legislation for grant-funded research, running companies at a loss for their data, etc.
We’re headed towards hundreds of billions in expenditure, national prioritization, and major data limitation on the horizon. We have a great opportunity to think creatively about what a megaproject for data would look like: How do we, deliberately this time, construct the next internet’s worth of data?
Footnotes:
[1]: I’ll probably soon publish my much longer post explaining my position on data efficiency and why the value of this data is still pretty high in most worlds regardless of new algorithms.
[2]: The “AGI freeroll” bet: heads you win, tails ASI flips the world upside down anyways.
[3]: We already see a glint of validation of this point, given the data market is strongly tilting towards ultra-high-quality agentic data, rather than unskilled labeling — niche expert workflows, live environments, and evaluations requiring increasingly obscure talent & knowledge — yet shows increasing, not decreasing, revenues.
For the first time, "context" is a machine-legible asset that is sitting on every firm in the world's balance sheet.
If you know that assets need a matching change in liabilities, that means every. single. company. in the world is currently mis-priced.
I don't care that Claude can pass the 8th grade math bench
I don't care that GPT 5.6 outperforms legal tasks on a fine-tuned eval that you're publicly sharing with me
I care that these benchmarks can solve problems that create unique value for my firm, and we have a widely accepted benchmark for that.
Free Cash Flow.
Token wrappers intentionally confuse the map (tasks that can be automated) from the territory (what knowledge is required to create enterprise value) because they need firms to cede more context.
Firms already know what hills to climb. Token wrappers don't want you to figure that out.
Increasingly clear that the labs have been hill climbing using the wrong map. It's all about tokenized economic complexity per $ compute (ie. the density of verifiably useful economic capability that's deployed in any given token, capability that can be combinatorially / repeatedly deployed to generate solutions to economically valuable problems). Capabilities borne from abstract task hill climbing don’t compress true economic complexity. True economic complexity is crystallized in firms, and cannot be separated from that context without coercion. Tokens per dollar as a metric is useless and not the right unit of work.
This era of "bench-marketing" by token-wrappers is ironic because the only benchmark for token economic productivity... is whether or not you generated free cash flow.
The evals to hill climb that actually create enterprise value will always be private knowledge.
If a firm knows what to climb, why would they publicly talk about that?
@aashaysanghvi_ With AI roll-ups it's less financial eng and more the willingness to take on liabilities that SV/SaaS has historically been unwilling to.
Owning a company, its management, and its outcomes are the primitive.
Most SV/SaaS/AI don't want to do that, to their peril.
The more general-purpose a technology, the more its value is captured far from where it’s built.
At peak output, Ford and GM combined were ~5% of US GDP.
The automobile’s value was captured miles away from Detroit (e.g Suburbs, McDonald’s, Walmart, C.H. Robinson, etc.)
AI is going to be an even more extreme example of this dynamic.