Defensibility in AI Data: Lessons from Ads
(Warning: LONG POST)
Companies selling data / expertise / RL environments to AI labs are growing at extraordinary speed, with several reaching tens to hundreds of millions in annualized revenue within a year.
But venture investors are unsure how to value these companies, given customer concentration, limited recurring revenue, and the constant treadmill of new data needs from the labs, which leads to every "product" (i.e. data feed / RL environment) having a very finite lifetime.
I see many parallels with advertising in the early 2000s.
Back then, dozens of ad networks sprouted up as the number of websites ballooned (thanks to blogging platforms, cheap hosting and CMSs like WordPress). These networks acted as middlemen; they bought unsold inventory from thousands of small and mid-sized publishers, bundled it into audience segments based on behavioral, vertical or contextual targets, and resold it to advertisers who couldn't negotiate deals site by site.
These ad networks solved a real problem, since most advertisers couldn’t manage relationships with thousands of websites, while publishers couldn’t build sales teams to reach every advertiser. Networks made fragmented supply easy to buy. Revenue grew exceedingly fast, and several got to tens to even hundreds of millions in revenue, which back then was very significant.
Over time, a problem emerged: many networks had access to the same inventory. Publishers worked with multiple networks to fill their inventory. Advertisers spread budgets across them but struggled to track where their money went. Revenue grew pretty without necessarily making the businesses harder to replace.
To solve this fragmentation, ad exchanges such as Right Media, DoubleClick Ad Exchange, and Microsoft's AdECN emerged around 2005–2007. They created marketplaces where buyers and sellers could trade impressions directly, laying the groundwork for real-time bidding and the programmatic advertising ecosystem that dominates today.
As these publisher-side ad exchanges (and on the advertiser side, Demand Side Platforms) matured, access to inventory became less differentiated. The question became: what does this network contribute that the buyer can’t get elsewhere? Ultimately, most ad networks shut down, got swallowed up and there are very few, I can think of that generated much equity value. Most of the equity value accumulated to the exchanges and the DSPs that centralized buying and selling.
AI data companies face a similar question.
Today, AI labs need experts, demonstrations, evaluations and increasingly complex tasks. Finding contributors, verifying their capabilities and delivering quality at speed is valuable work.
But if the same experts work for five vendors, the lab specifies the task, and the lab owns the output, what is the supplier accumulating?
The lab gets a better model. The supplier gets paid and starts looking for the next project.
That can be a very good revenue business. Unfortunately, this doesn't mean it's building durable equity value or is a true technology platform.
I see three ways to build durable equity in this space:
1. Own differentiated supply: Google and Facebook were the biggest winners in advertising because they both owned their inventory and didn't rent it. Data vendors should similarly have a product, community or exclusive relationship that continually generates useful data. The ability to produce the next valuable dataset matters more than possession of the last one.
2. Become part of the customer’s workflow: Doubleclick (Google paid 1% of its market cap for it) was deeply embedded in both agency and publisher workflows. The Trade Desk ($10B public company) is the platform used by many agencies to buy media. Similarly, the data company needs to become indispensable to the lab: move one step earlier in the funnel and help the lab identify model weaknesses, design tasks, evaluate results and decide what data to buy next. The deeper you participate in those decisions, the harder you are to replace.
3. Build self-improving environments: Invest in technology that turns model failures into new training tasks, verifies outcomes and adjusts difficulty as models improve. Each training cycle should make the environment better at identifying and teaching what the model still gets wrong. The durable asset is the software that generates and evaluates the next useful task. That requires access to performance feedback; if all the learning stays inside the lab, the supplier keeps producing data without building a compounding technology advantage. This is something you can absolutely charge a recurring fee for.
#1-3 collectively also likely mean that data vendors should consider specializing in a handful of verticals where they have unique advantages. I believe that large horizontal data provider will be replaced by vertical specialists who own differentiated supply for that vertical, run customer workflows for that vertical and build self-improving RL environments. AppLovin ($150B market cap) is the largest stand alone ad-tech company today. They do exactly ONE THING - drive app installs - and they do it damn well. They own the end to end value chain and have built incredibly differentiated supply as well as a self-improving loop that delivers downloads at ever improving cost-per-install. If they had tried to do anything more, they would likely not have seen similar success.
The question every data founder should ask themselves: after delivering your first $100M of revenue, what will you own that makes the next $100M easier to win and harder for a competitor to take? The answer should be more than a larger roster of experts or stronger relationships with the labs. It should be unique supply, deeply embedded workflows, or technology that gets better with every training cycle.
The demand boom gives you the revenue to build those assets. Whether you do will determine how much equity value remains as sourcing data becomes easier and (inevitably) commoditized.
PS: If you're a founder working on the above, please DM or email me at gokulr at gmail. I would LOVE to speak with you and brainstorm!
Lots of founders are crashing out
A year ago AI felt like an incredible opportunity.
Most of us were able to ship faster, remove admin, and get time back.
Lots of engineering founders were even able to explore marketing for the first time: SEO, organic social, outbound.
Building complex systems that felt extremely productive.
But since then, a lot has happened.
Here’s some of the themes I’m seeing:
AI marketing isn’t working. At first it feels impressive and it looks good but for most it’s not driving results. This is made worse by the fact everyone has access to the same tools. Inboxes are flooded, buyers are fatigued, the economy is flat.
Engineers are fried. Yes you can ship more but for many the flow state has gone. It’s a new way of working and it’s not for everyone. The speed leaves many of us exhausted before 11am.
There’s a feeling of what’s the point? Is my product’s next feature going to be redundant in a year, or a month? Can’t AI do what my business does, better?
SaaS valuations have collapsed. From 3-4x revenue to 1x. There’s a sinking realisation that a large number of SaaS companies will go to 0 in the next few years. Building to exit seems almost crazy right now.
There’s far fewer buyers and far more sellers. These include vibe coders cloning products without the care or craft, contributing to the distribution challenges from people doing it the right way.
Lots of people are trying to make money by selling shovels. This just adds to the frenzied energy.
Build in public stopped being fun. A few makers realised that the new game is attention, and now everything feels insincere and stunt-driven. They are influencers not founders. Changes to the X timeline compounded the issues.
There are exceptions and there are still moats remaining, but the challenges feel existential
Databricks CEO @alighodsi went off on @a16z pod about enterprise AI adoption:
"They're just so far behind in the adoption curve of actually automating things and getting value out of this stuff."
Ali says most companies are still just using chatbots. There's hardly any agentic transformation.
Why is that?
"The models are smart enough, but they just don't have the context that exists inside of any organization."
"They have not been in every meeting. They don't know what's in everybody's heads. They don't know all the processes."
"If you just fused that and gave that context into the AI models...there's so much productivity gains you could get for any organization on the planet."
Context Creation is the biggest opportunity in AI right now.
My god this is such a good speech that every SWE needs to hear. You know what? Every person should hear it
Keep the happy memories, eyes on the reality, be excited about the future. That’s the best that anyone can do
just got off the phone with a new engineer hire at a portco with a very wise insight
“AI has been replacing my job since I got a CS degree, yet I’m busier than ever - we can all just be more ambitious”
this is the same insight that the inimitable @JensenHuang has for us all
People outside the AI labs should have a real say in how this technology develops, and a clear way to judge if it's happening safely.
Standards should help prevent the concentration of power, including by making sure new companies and open-model companies can compete.
They should also help countries and companies compare evidence and learn from failures.
We think the US should lead this effort. Here is our proposal:
https://t.co/2FT9a2m4l0
"The average person has fewer tasks" is right. I think we also overestimate the degree to which people value "productivity". The sweet spot is in the magic of the experience. Does it solve a pain, and does it feel easy doing it?
I think Silicon Valley is making a fundamental mistake with personal agents… optimizing entirely for the outcome and forgetting that, for a lot of things, the process is a very core part of experience.
Travel is the obvious example. People will almost always never want “book me a trip.” They want to browse, compare, daydream, change their mind, send options to friends, and eventually book.
And the average person probably has far fewer recurring tasks worth delegating than the AI industry seems to think.
That’s why personal agents feel more like a feature that gets absorbed into existing products than a standalone category.
A few thoughts on the current state of venture capital.
When the Music Is Playing
In July 2007, a few weeks before the credit markets seized up, Chuck Prince, then the CEO of Citigroup, gave an interview to the Financial Times. The line everyone remembers is this one: "As long as the music is playing, you've got to get up and dance." He was mocked for it for years afterward, and he lost his job a few months later. But I have come to think he was saying something honest. He wasn't claiming the music would play forever. He was admitting that he couldn't sit down while it was still going, and neither could anyone else in his seat.
I've been thinking about that quote a lot lately, because right now is the most disorienting period in venture capital I can remember, and I have been doing this for a while.
Here is what makes it disorienting. It's not that things are bad. Some things are spectacular. We have companies in our portfolio growing faster than anything I have seen in my career, and I don't say that lightly. At the same time, we have companies with no revenue, no product, and a founding team you could fit in a conference room raising billions of dollars at valuations of $10 to $50 billion. Both of these things are true at once, and if you try to reason about them with the same framework you will drive yourself crazy.
Two ideas have helped me make sense of it. Neither is mine.
The first is reflexivity, which George Soros has been writing about since the 1980s. In most of life, perception follows reality: the weather is what it is, and your opinion of it changes nothing. In markets, it runs the other way too. Prices change what participants believe, and what participants believe changes the prices. The feedback loop can run for a long time, and while it's running it looks exactly like progress.
Here is how reflexivity is playing out in AI. Full disclosure: Menlo is an investor in Anthropic, so read the following with that in mind. People watched a frontier lab go from a $4 billion valuation to $18 billion, then $60 billion, then $180 billion, then $380 billion, and now something close to a trillion. They drew the obvious conclusion: that is what a neo lab looks like. So the next neo lab gets priced off that path, not off anything it has built. Then it gets marked up in a subsequent round, and the markup itself becomes the proof. Look at Thinking Machines. Look at Reflection. At that point valuation has stopped being an output of the metrics and has become the metric. Nobody is discounting cash flows. They are discounting the last round.
Soros is very clear about one thing, and it's the part people skip: you cannot know when or how a reflexive process ends. You only know that it does. Every one of them has.
The second idea is Chuck Prince's, and it explains why smart people keep dancing even when they can see the loop for what it is. As far as I can tell, there are two groups on the dance floor.
The first group got in early. Firms like ours were in some of these AI companies before the numbers got silly, and the paper gains are enormous. When you are sitting on gains like that, you start to feel like you're playing with house money. I have been around long enough to know that house money is the most dangerous kind, because you don't respect it the way you respect money you had to earn.
The second group missed the early rounds and knows it. Their LPs know it too. So they are trying to make up for lost time by writing very large checks very late, which is the one strategy almost guaranteed to turn a missed opportunity into a real loss.
House money on one side, FOMO on the other, and reflexivity feeding both. That's the whole story. Everyone has a reason to keep dancing, and the reasons are different, which is why nobody can talk anyone else off the floor.
So what do you do? The instinct in our business is to answer with company identification: just pick the right neo lab and you'll be fine. I think that's the trap. When price has become the signal, being right about the company is not enough, because you can be right about the company and still be wrong about the price by a factor of ten. The public-market investors I admire figured this out a long time ago. They spend as much time on how much to own as on what to own.
The winners in venture over the next decade will be the firms that treat portfolio composition and position sizing as seriously as they treat sourcing. How much of the fund is in companies whose valuation rests on the last round rather than on revenue? What happens to the portfolio if the reflexive loop breaks next year instead of in five? Those are not exciting questions. They are the ones that will matter.
The music will stop. It always does. Dance if you must, but know where the chairs are.
@ccatalini@wu_jane Metric manipulation is also going to be a big problem. One thing I've found already is agents and models at scale can design and convince with metrics better than even great data scientists, even if its not accurate.
Fascinating to me to think about what quantities are worth measuring. Designing high signal metrics is really hard to do well, even with all information available to you. Communicating about how those metrics change and what it means practically is even harder. Huge opportunity space.
The best and most talented people I know have a hard time promoting themselves. There's humility and pride in craft that keeps them from being loud, even when it would benefit all of us. Hard problem.
Shamelessly promote yourself.
The world is full of incompetent people who aren't ashamed to promote themselves.
So if you're competent, it's your obligation to promote yourself.
There are two ways AI progress could go very badly and that we must avoid.
First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities.
Second, we could end up in a world with too much concentration of power. If an extraordinarily powerful AI is used by one person or company to impress their worldview onto everyone else, the results could be extremely dystopian.
Avoiding these two threats requires walking a narrow middle path; for example, one country could gain too much power. Another example is one lab ending up with too much power.
Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing.
We also need to accelerate and spread the benefits of AI, such that they are diffused broadly across countries, communities, and companies. This requires a frontier ecosystem in which both closed and open-source models can thrive.
And for firms, it’s imperative that they retain full control over their unique and tacit knowledge. Every organization should be able to build its own continuous learning loop/hill climbing machine, without becoming dependent on any one model provider, and have the ability to embed its own knowledge into models and weights they control.
So, in this context, we welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal. We also welcome ideas like "embedded evaluators" and the broader efforts to develop the mechanisms to make this more than just talk.
The key is that this cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia.
This is the approach we are taking: broad access and choice at every layer of the AI stack; enterprise control of learning loops and models; and the “Code of Conduct” that underlies our own first party MAI models that we’ll publish tomorrow for public consultation.
finding super technical founders with killer commercial/biz instinct is so hard and rare
most just want to build product but building a prod is very diff from building a business