The next breakthroughs in biological AI will be shaped by the datasets we build now.
We’re planning an Impetus Grants @impetusgrants round to fund shared, AI-enabling biological datasets.
Learn more and contribute: https://t.co/pCQFD4hOSI
What would the infrastructure needed for longevity to maximize healthy years look like? And what could the field become once it exists?
Take antibiotics as an example. During our war against infectious disease, they became an industry when each discovery and infrastructural innovation became the ground for the next. Together, advances in discovering and producing them compounded into their golden age.
Longevity has not yet found it's versions of the advances that allowed antibiotics to become an industry. We are making progress, and new routes are emerging. Even more progress toward the infrastructure needed for broader industrialization will be essential to human flourishing.
Read it here: https://t.co/qrqA2xMVHR
The last great rise in human longevity came from industrializing the fight against infectious disease. The next will come from doing this for aging.
Since the 1960s, we have primarily been adding years to lives through the Sisyphean work of fighting disease one by one. This is the result:
https://t.co/th2n4oWdx1
That is Good, but we should ask for more. The next great project of longevity is therefore to organize medicine and drug development around that goal, following the precedent set by the fight against infectious disease: Build the tools and infrastructure to turn longevity into a compounding industry whose product is maximal healthy years of human life.
We wrote about turning longevity into medicine’s next great industry, link below.
The problem is that age-related diseases accumulate. By 65, about half of people have at least two chronic diseases; by 85, three quarters do. Take one away, and you are soon in the grips of the next.
Yet modern medicine has crystallized around developing treatments for them one by one, reinforced by a drug-development system that directs candidates toward single-disease indications where uncertainty is easier to contain.
The result is Sisyphean medicine, destined to leave most old people less-but-still sick.
Longevity is the war against aging as the main limit on healthy life. It is the bright alternative, where we can change the course of later life by altering the ways our bodies break over time.
That includes discovering interventions, but it also goes beyond them. Longevity is a field-building effort aimed at expanding what medicine can test and commercialize.
Erdinc Sezgin @Sciezgin was not in the aging field when he came to us with an idea: test whether membrane fluidity, stiffness, viscosity and molecular organization could be measured together to build a biophysical clock of aging and resilience.
We funded him to pursue it via @impetusgrants. He has since produced three papers building what he needs to answer that question, including better probes for labelling cell membranes, software that turns images into biophysical measurements, and evidence that membrane properties distinguish cell states.
The longevity field is richer with people like Erdinc joining from other fields and bringing new expertise with them.
Read the papers here:
https://t.co/6LZqFixVam
https://t.co/g0qBV7rjux
https://t.co/4jYiPO8CJ9
The growing value of long-term dataset building is extraordinary. Here’s another example:
UK Biobank made its dataset even more powerful by applying large-scale proteomic profiling to plasma samples collected over a decade ago. Researchers could then identify blood protein signatures associated with the risk of developing a wide range of diseases years later.
Published in @NatureMedicine, the study used longitudinal data to capture changes and outcomes over time that cross-sectional data alone cannot: https://t.co/Vu6BnDmpjW
Time is a non-substitutable input, as we've written about (link below). No amount of money could recreate UK Biobank by next month.
Pub out by @Xiaojie_Qiu on a new privacy-preserving single-cell foundation model used to identify candidate rejuvenation factors, funded by @impetusgrants. Great to see more work on making AI useful for longevity!
We are thrilled to share our latest work:
“Predictive Single-Cell Foundation Model for Gene Regulation and Aging with Privacy-Preserving Tabular Learning.”
We introduce Tabula, a foundation model for single-cell genomics that models cellular data directly in tabular form, rather than as artificial gene sequences, while enabling privacy-preserving, federated training.
Beyond improving predictive performance, Tabula recovers experimentally validated combinatorial gene-regulatory logic across multiple developmental systems. We also tested the framework experimentally with a custom inducible Perturb-seq system and found the reprogramming factors Oct4 and Sox2 lower the age score in aged fibroblasts as predicted, while Tabula-nominated candidates such as Cpe, Postn, and Olfm2 act orthogonally to reprogramming, expanding the space of tractable rejuvenation strategies
An early version of this has posted in biorxiv previously but was now dramatically updated by the incredible @JiayuanDing , Jianhui Lin, Ziyang Miao, @nilsmechtel, in collaboration with @Y_Ryan_Lu, @imweio@tangjiliang and many others.
Thanks for the support from @LaudeInstitute
and @arcinstitute
Details below. 👇
“progress depends on the interplay of techniques, discoveries and new ideas, probably in that order” - Sydney Brenner
With @ImpetusGrants, we've supported and put several new tools into everyone's hands: from a cytokine on-off switch and RNA sensors to perturbation-sequencing techniques, an RNA-writing method, microfluidics-based lifespan profiling, and a system for discovering mTOR inhibitors using drug-sensitized yeast.
Time and again, a new tool is what tips a field into its next advance.
Read them here:
https://t.co/LNn9Y8H3Dc
https://t.co/WFZKkefHuc
https://t.co/MEBWrbD5jz
https://t.co/K55yTgcZoN
https://t.co/6IcrsCk8Zp
https://t.co/Y6vei5cdNt
https://t.co/96oI9pF7nZ
https://t.co/rjpeqOQs0d
We build datacenters to meet anticipated compute demand. We should also build biological datasets to meet anticipated demand from abundant intelligence.
AI is quickly getting good at drug design and hypothesis generation, which means more hypotheses in line to be tested. But target validation hasn’t kept pace.
It is still hard, slow, and expensive, so drug discovery has crowded around the same already de-risked targets while most of the genome stays unexplored.
Like our founder @MartinBJensen mentioned at @PMWCintl, to break out of that and achieve real target abundance, causal validation of targets needs to work at high throughput.
AI hasn’t been able to tell us what targets to go after because we don’t have the right data. Biology operates at three layers: molecular, cellular, and physiological. Each layer has emergent properties which cannot be predicted by just looking at the layer below it. Most diseases are organ-level phenomena, and we don't yet have the in vivo datasets to predict them from cellular data alone.
This is where we are focusing next for Impetus Grants @impetusgrants: building the datasets that will give us more and better targets for aging drugs.
Contact us in the link below if you want to support this round.
This is true: biology is still craft-bound - especially at the physiological layer. We’re putting together the components that will transform aging biology from craft to industry:
We’ve organized around the need for validated aging biomarkers with the Biomarkers of Aging consortium @agingbiomarkers, we’re creating new clinical trial designs to test them (link below), and we’ve been thinking about new reward pathways for aging drugs in the market (https://t.co/xoF6qNfSgu). On top of that, millions in grants have been distributed to basic research on aging and tools with @ImpetusGrants.
Once the right pieces exist and interlock, output can start to compound. Like electrification and the semiconductor industry, whose pieces interlocked to give us universal tools that could be improved over and over.
Why we made @Bioticorg. I believe biology today stands roughly where construction and civil engineering stood 300 years ago. That needs to change. A Thread 🧵
This data scarcity applies to bio too.
The “bio needs more data” take is gaining traction. We agree, and we want to turn that into definite strategies for data collection.
We need datasets that reveal causal disease trajectories, how disease emerges, how interventions change it, and which measurements predict physiological benefit.
These are the next “Protein Data Banks”: foundational, AI-enabling biological datasets that may look unglamorous while they’re being built, but could become the substrate for the next AlphaFold-scale breakthroughs.
In longevity, we are enabling that with an @ImpetusGrants data for AI-focused round, read more here: https://t.co/pP5DenvlBr
A Stargate for Data
Labs are on a trajectory towards >$100B/year of data spend by 2030. As we begin the trillion-dollar compute project, we need to think about the equivalent civilizational-scale effort for the other core ingredient: data.
At the foundation of the scaling revolution is a simple empirical law: deep neural networks improve smoothly, near magically, as you scale two things in proportion — (1) the size of the model and (2) the amount of data you train on. And despite the scaling laws being brutally diminishing, we’ve successfully bitten the bullet of logarithmic scaling with exponentially larger clusters and datasets, and received incredible new capabilities in return.
But this exponential scaling is bound to hit some limits. Oddly enough, compute has compounded fairly smoothly without limit, with trillions flowing into hypercluster buildout. Instead, we’re starting to hit the limits of an exponential demand for data. Gone are the days of being purely in the compute-limited regime, where we had effectively infinite internet data but never enough GPUs, we’re now entering a data-limited regime.
Luckily, this limitation is coinciding with staggering improvements in AI capabilities. Incredibly, we seem to have a real line of sight towards automating a majority of knowledge work with the methods we have today. RL + pretraining, and the data for each, will be generally sufficient to achieve most economically valuable tasks, given some minimal algorithmic progress and continued compute scaling.
In a data-limited world, economic progress & scientific acceleration will be directly bottlenecked by our coverage in each domain. We need to see data collection as imperative, deserving the same civilizational ambition we’ve given compute.
The internet as a one-time subsidy
It’s underrated how much all progress in AI owes everything to the blessing of the internet, this one-time civilizational subsidy to deep learning, decades of unintentional accumulation of a perfect dataset: every book, blog post, image, video, paper, discussion, etc. all digitized and freely available. Without the internet, we’d likely see comparably minimal progress in AI today, and in fact, if you notice where systems currently underperform, it’s almost always a domain where web coverage is limited and data is private, expensive, non-digitized, or non-existent.
But we’re running out of it. There are only about 300 trillion tokens of useful public human text, and the internet doesn’t produce nearly enough new high-quality data to match what scaling demands — we’re soon to hit the limits of public data for pretraining. And though the advent of RL bought us reprieve — chain-of-thought RL needed a new form of untapped data, gradable math & coding tasks, also available online — we’re quickly running dry of hard tasks for RL as well.
Why do we need so much data anyways? Humans learn comparably in far less time, needing just one textbook where language models might need the equivalent of hundreds to learn a new topic. It’s possible we discover methods that are massively more data efficient — synthetic data, data efficient architectures, other exotic algorithms — but fundamental progress is slow and highly unpredictable, and the recipe we have just works today.
And, while I’m wary of getting too deep here, even arbitrary data efficiency can’t replace data that just doesn’t exist in the first place. There’s a massive amount of missing information on the web: the dark matter of the internet — tacit knowledge, undocumented processes, etc. — most of which was never published and lives only inside organizations, the physical world, or just in people’s heads. I’ll leave it here and say, for reasons far longer than I can fit in this post [1], it’s best to operate on the assumption that our insatiable desire for data will continue as it has for the last decade.
There will be >$100B/year in data spend by 2030
We’re not screwed yet, of course. Only a fraction of useful data in the world is on the public internet, the rest is stored inside private datasets, corporations, personal archives, universities, governments, and otherwise. Labs can and will continue to license these private datasets, or create them from scratch, like Anthropic’s book scanning project. And we’ll increasingly task human experts to manufacture new high-quality data, with a large fraction of hard RL training tasks already being sourced this way.
But collecting this data, unlike before, will be expensive. As the free internet dries up and demand for data rises, we should see labs investing equally in data as compute, likely spending a significant fraction of their compute budgets on data. As we see trillions spent on compute, we should also expect hundreds of billions spent on data (human data & collection budgets), given their equivalent importance. And, notably, data spend is already tracking this way: total data spend across vendors, not counting internal lab efforts, is already roughly $7 billion per year. It’s quite reasonable we’ll see >10x by 2030.
Data is the moat
Data becoming increasingly private will also majorly shift the competitive landscape. While compute is a commodity — everyone buys the same chips and builds the same clusters — data really isn’t. The big reason why frontier models have felt eerily similar to one another, until now, is they were trained on substantially the same internet (pretraining data variability across labs seems pretty low). As labs diverge onto more exclusive, manually collected corpora, I think models will begin to increasingly diverge.
OpenAI pulling ahead in mathematics and Anthropic in cybersecurity isn’t an accident. I really think laser-focused collection of high-quality midtraining tokens, custom RL tasks, environments, with dedicated research effort, has driven much of the visible progress in the last year. James Betker has an excellent blog about “the ‘it’ in a model is the dataset”: model architecture and compute buy you efficiency and order-of-magnitude performance, but ultimately, models, of any architecture, are such incredible approximators of their dataset that the core meat of a model boils down to just that, nothing else. Data is a major moat.
AGI long, ASI short
As I’ve tweeted before, I’m confident that, despite the narrative, the data labeling industry will continue to fuel great businesses and be an excellent AGI long, ASI short. The argument is just: By the time the AGI labs no longer need data, it’s probably over for everything else too [2]. In this frame, the last companies left should be the data companies, as the last speck of economically relevant data is sucked in. And these companies are already among some of the fastest-growing companies in history: Mercor, founded three years ago, is rumored to be doing $2 billion in revenue with something like a few million expert labelers under contract.
While these businesses are very non-stationary, what type of data is needed shifts constantly, I don’t think that diminishes their value. The long-tail of the economy is long, and the value isn’t diminishing as you extend farther into more obscure information: as models get more capable, the value of the marginal dataset goes up, not down. Automating a full job means covering its full distribution of tasks, tools, edge-cases, and long-horizon loops. There’s some O-ring logic to it: a dataset that buys a 1% bump can justify a previously unjustifiable collection cost when it’s the difference between a system that does 99% of a job and one that does all of it [3].
The competitive dynamics of the data industry are still evolving but as demand for data is increasingly niche, ultra high-quality, expert-generated, I think we’ll see real consolidation. Again, contra-narrative, we’ll probably see true competitive differentiation built on brand, quality control of data (which, from personal experience, can vary massively), as well as in network effects from the talent networks themselves over time. We’ve already seen rapidly shifting data type demand work in favor of incumbents, benefiting those with early knowledge of where the market is headed.
The binding constraint
It’s truly remarkable that we seem to have the recipe — pretraining + RL — to absorb most economically valuable work, despite being far from a lot of what we expected from “AGI”. The same way chess engines revealed we never needed general intelligence to solve chess, as we originally thought, we’ll soon realize that software, mathematics, and the vast majority of the economy (including physical, just running ~3 years behind!) are the same. If recursive self-improvement or some other algorithmic breakthrough arrives, that’s wonderful, but we really don’t have to wait for it. The binding constraint between here and an automated economy isn’t that, it’s data coverage: every app, workflow, edge case, process, etc. sitting in private stores or someone’s head.
Ultimately, while we make tremendous strides in more efficient model architectures, and clusters like Stargate equip us with zettaflop-scale compute, we really aren’t making rapid progress collecting the data we lack.
We’ll soon live in a world where we have the methods & compute to accelerate scientific progress or economic growth, but not the data. And we’re already there today: frontier models would surely be as good at accounting/many medical tasks/legal advice as they are at software engineering if we only had the same pretraining & RL coverage as we did for code.
I really want to drill this in: The speed at which we automate the economy is going to be directly rate-limited by our ability to collect data about it.
Worth noting that under this assumption, with data as defensible and directly proportional to economic & scientific progress, data should also be considered a national strategic asset like compute. Imagine what we’d do in a world where we had a Manhattan Project-effort for AI and needed to mobilize data collection as a limiting factor. We should be concerned about China, with greater state capacity and authoritarian economic control, being capable of mobilizing data collection at national scale, potentially compounding their economy and scientific output faster than us down the line.
A Stargate for data
I’m leaving my complete ideas for a future post, as this one is already far too long, so I’d really like to pose the question here. Stargate exists because we organized trillions of dollars, international strategy, gigawatts around compute as a fundamental ingredient. What would equivalent ambition look like for data?
Obviously, scaling data collection, a heterogeneous mass of information across the economy, isn’t going to be as clear as scaling compute, as a homogenous infrastructural effort. A core division will be first, coverage — all uncaptured knowledge sitting across the economy/science/physical world and all that simply isn’t recorded — and, secondly, sheer volume in the domains we already train on: more hard math tasks, more high-quality web text, way more coding data, more legal drafts, etc.
I have a post coming soon which breaks down my proposals. There’s a lot of room for creativity. Quickly, we’ll probably want to start with a deep census of what we have and what we’re missing, predict what the 2030 model will still be bad at and work backward to what we should be collecting today. You can probably license a large amount, leveraging high lab valuations to buy datasets or companies altogether. There’s an adversarial nature to a lot of this collection with firms, so there’s lots of engineering to do this correctly. We should go convince important companies to turn off deletion policies, even if we’re not buying from them yet. Data flywheels in consumer products will be massive. Confidential training, government legislation for grant-funded research, running companies at a loss for their data, etc.
We’re headed towards hundreds of billions in expenditure, national prioritization, and major data limitation on the horizon. We have a great opportunity to think creatively about what a megaproject for data would look like: How do we, deliberately this time, construct the next internet’s worth of data?
Footnotes:
[1]: I’ll probably soon publish my much longer post explaining my position on data efficiency and why the value of this data is still pretty high in most worlds regardless of new algorithms.
[2]: The “AGI freeroll” bet: heads you win, tails ASI flips the world upside down anyways.
[3]: We already see a glint of validation of this point, given the data market is strongly tilting towards ultra-high-quality agentic data, rather than unskilled labeling — niche expert workflows, live environments, and evaluations requiring increasingly obscure talent & knowledge — yet shows increasing, not decreasing, revenues.
Our Longevity Nexus members organize the field around ambitious research directions.
Bjorn Fraser Olaisen @BjornOlaisen is one of them. His new Aging Cell Perspective, based on the first Replacement in Aging workshop at ARDD 2025, helps turn an emerging area of longevity science into a field-shaping roadmap.
The Perspective gives researchers a shared way to think about replacement-based ageing interventions. It maps how replacement, regeneration, damage-removal technologies, and aging biomarkers could work together in the pursuit of systemic rejuvenation.
It also helps define the right problems to tackle, giving researchers a clearer path for making progress. Here, that means clarifying when replacement is the relevant intervention and asking what questions follow if longevity science moves beyond conventional therapeutics toward rejuvenation.
Read the Perspective here: https://t.co/PC8j6R7k9N