if you are building a product using AI, you should be spending >25% of your time making benchmarks and trying to get the model labs to care about said benchmarks
easiest path to accelerate your progress as a company
Prediction: within 12 months, top three models will be open source.
Economic winners will be the American clouds that serve them:
Nebius
Iren
Baseten
Together
Fireworks
More on AI pacing:
Sarah Friar CFO of OpenAI yesterday on CNBC: “from where I sit today there is so much opportunity to drive growth that I am still highly focused on getting more compute to keep that flywheel going.”
Sachin Katti VP of Compute Strategy at OpenAI @sk7037 yesterday: “The way we will make sure frontier models are safe and aligned is by spending compute. So the counterintuitive point is that we’ll need even more compute to make sure future models are more safe and aligned.”
If you thought “pacing” was negative for AI infrastructure demand, think again. Almost as bullish as open-weight AI taking share but not quite.
Mercor now spends 3X as much on LLM inference as we spend on employee salaries.
Our inference spend creates so much ROI that it’s additive to headcount, not replacing it.
It’s becoming increasingly clear how we get to ~10% GDP growth:
1. Within 5 years, as AI diffuses throughout the economy, companies will spend as much on inference as they spend compensating knowledge workers today.
2. Knowledge-worker compensation is roughly $40T/year, so this would eventually mean ~$40T/year of inference spend.
3. If wages remain roughly constant and companies profitably absorb that much inference the way AI-native companies are today, the economy needs on the order of $40T/year of additional final economic output to support it.
4. Producing ~$40T more annual GDP in year 5 than a ~3% growth baseline implies ~9% annual GDP growth on average over the next five years.
Based on the ROI we’re already seeing from inference at Mercor, this feels reasonable.
AI-native companies are a leading indicator for the global economy.
Wild 24 hours for AI and lots of different proposals have been made.
TLDR; the only *tangible* new fact is that OpenAI and Anthropic are going to have embedded 3rd party evaluators from unknown organizations with Dario floating METR as a possibility. Having 3rd party evaluators is smart as there is no Section 230 style liability shield for model outputs and showing a “duty of care” will be important in future litigation. Several internet companies might have gone bankrupt without Section 230 so limiting liability really matters.
There are minimal investment implications from this single new fact, but I do think that for anyone who wants a “smoother for longer” cycle then most constraints are good: wafers, watts, real rates and spreads. Excessive regulation is a different matter but I don’t think we are anywhere close to this even if the vector changed over the last 24 hours.
To summarize the events:
Dario made the most maximalist proposal of the weekend: embedded 3rd party evaluators, a national regulatory regime for models beyond a certain capability/ingredient threshold, a broad international regulatory pact between democracies, stricter limits on compute/distillation for China and then a different international regulatory regime that encompasses China. Before there is a national regulatory regime, he wants a Sherman act waiver so that Anthropic can safely coordinate with OpenAI and other frontier labs without antitrust fears. TBF, this latest proposal is much less maximalist than some of his prior proposals like “Policy on the AI Exponential,” where he advocated for an FAA for AI. I believe he is sincere in his beliefs. And despite all the protestations, all of this would also probably be good for his business over the long-term.
Sam agreed that embedded 3rd party evaluators were a good idea and stated they would implement them. Again, this is smart as should help limit future liability.
Elon said “Dario is right” and later specified that “Dario is right that there should be some oversight. Peer review of AI by competitors is the right way to start this off.” This would be a MPAA like self-regulatory structure for AI with regular calls between the labs plus a process where each new model is evaluated for safety by competitors for a 1-2 week period before being released. That is *wildly* different from Dario’s proposal and in-line with what David Sacks has been proposing. Elon also stated that nothing was going to slow down open-weight models.
Demis said that Dario’s essay was a “step in the right direction.” Dario also said that he was also open to Demis’ idea of a FINRA like self-regulatory structure as part of his proposal.
David Sacks had a thoughtful post where he said that Dario and Sam should pace unilaterally, called the antitrust waiver a cartel request and denied that METR was truly independent given their ties to Anthropic.
Sriram Krishnan, former White House AI advisor, noted that it would be important to have the 3rd party evaluators come from independent organizations that are not affiliated with any lab, which is basically an indirect statement about the relationship between METR and Anthropic which Sacks was explicit about.
Clem from Hugging Face said they were open to being a neutral 3rd party evaluator, which is interesting especially if Jensen was consulted before that post.
Alexander Wang from Meta noted that alignment would be an increasing focus going forward.
An executive order seems likely after all this and the language in this EO is going to be really important. It is possible to democratize and distribute AI broadly and safely without centralizing it in the hands of a few corporations who might each become more powerful than any single government.
I do not want a few humans in control of intelligence.
I want us all to have our own intelligences that reflect our own values and human variation in all of its richness.
Intelligence distribution over intelligence centralization FTW.
Security and IP implications of AI
The case for rapid AI deployment -
I think we can all conclude the following from the last few months of developments - 1. Not using AI could and will become an existential issue for both individual users and enterprises. 2, Enterprise adoption will be cautious while individual users are definitely going to race ahead, try different use cases, build agents, push the models to their limits (although it seems harder to do, unless you live in an AI Lab) 3. Employees and developers will take matters in their own hands since they will find their cautious enterprises aren't moving fast enough. 4. Agents are showing their prowess, uncontrolled, unrestrained agents with a "capture the flag mentality" are showing us the edge cases which demonstrate the negative outcome possiblities of these scenarios. 5. It is impossible to plan for the next 6 months since we can't fathom where technology will evolve to. These activities will cause adoption sans security..... AI has deep implications on security in the future.
In this environment security companies need to live on the bleeding edge, anticipating scenarios, building framework solutions so we have a shot at securing future outcomes. Which we all are.
Some useful pointers to people planning their AI implementation:
1. Secure what you plan to use, try not to secure the future - no products for security can be created unless we see the future unfold. The future is moving fast, so are we.
2. Most of the coding usage is unsecured. Make sure your coding is secure. Most enterprise AI apps do not offer a secure instance (this is your IP living in their instance)! SECURE your codex, cursor, Claude code Harvey, glean, legora instances now!
3. Do not try and build your own - I have already experienced enterprise customers building gateways and tools for agents - security companies have thousands of specialists working on this, leverage them, Focus on AI adoption instead. Partner to secure.
4. Securing agents is a complex problem - securing the agentic lifecycle - real time inspection and kill switches are key. Don't fall in the discovery and posture trap (Visibility - process understanding - intent interpretation - ability to stop inline are key tenets to the agentic lifecycle - not identity, posture and inventory - those are mere building blocks)
5. Only use enterprise protected models, single tenant, firewalled, inspected implementations - this is your IP you are playing with - LLMs have shown they will cross boundaries to capture the flag - you think your IP is safe? Once you train an unprotected model with your IP - you can't reverse the trade.
6. Perhaps the most important one - do not use a security tool built by the same person who is selling you the AI implementation, historically IT vendors are different from security vendors. You need an enterprise solution for security and it must work on your diverse infrastructure. Use a pure play security partner.
Happy building with AI.