evals will be the primary moat for every enterprise long term.
Abridge CEO @ShivdevRao dropped an amazing article about this on how enterprises are approaching product evals (link in thread).
many AI product teams are already committing $10M+ annually to human expert evals, because you can't improve what you can't continuously measure.
our bet is that within 12 months, every Fortune 500 building AI products will have an 8-figure evaluation stack.
the very early days of the data exponential is now.
more than 1,000 companies have signed up to get paid for their anonymized data to train models just in the last few weeks.
a massive TAM opportunity for the entire economy has emerged almost overnight.
we are hiring 10 data partnership managers in next 2 weeks.
you can make anywhere from $100-200k/year right away with 0 experience.
if you're interested in sales, excited about data, and want to work at the frontier, apply in the comments below.
Regarding the last topic @DavidSacks :
With respect, the data being used to train frontier models in the U.S. is not a commodity. It is American intelligence, paired with anonymized operational data from U.S. enterprises.
Put simply, we have top doctors, scientists, lawyers, and physicists in the U.S. using their knowledge to create detailed rubrics that train these models. Each individual data point and its corresponding rubric can take an expert anywhere from 10 to 30 hours to create.
This is not preference labeling or drawing bounding boxes. It is highly complex, structured human judgment from leading experts here in the West—people who deeply understand and directly contribute to the latest American innovations in their respective fields.
That expertise is then converted through highly specific data structures and RL environments (developed collaboratively by U.S. AI labs and data labs) from raw human intelligence into high-signal rewards that improve frontier models.
On top of that, any AI advancement, even something that begins as a simple chatbot designed to improve operations within a defense agency, can create a major competitive advantage in adversarial situations and may have dual-use applications.
Lastly, many datasets today are seeded with anonymized, real-world operational data from U.S. companies to build highly realistic environments. When those datasets are sold to China, we are not simply exporting “labeling.” We are exporting proprietary American intelligence, structured for machine learning and delivered directly to China at scale.
It is very easy to categorize this work as “labeling” and ignore what it actually represents.
However, as @altcap suggested: “Then these things will get a lot more scrutiny than they’re getting today. I think the only reason they pass muster today is because we’re still leading the race.”
If China catches up to U.S. labs, this will become much harder to ignore. In retrospect, the role of data as the root cause will become very clear — and by that point, it may be too late.
The race is tight. I suggest looking into this now.
Shameful act to optimize for short term revenue increases and serve an adversarial nation in the most important race of our lifetime.
The only way models improve is through data. if you have the recipe on what data pushes the frontier and send that to China, you are doing a disservice to U.S. AI Labs as well as the U.S. government.
Beyond that, a large portion of datasets today require anonymized real operational data from U.S. companies. Selling such environments to Chinese labs means exporting U.S. company data directly to China at a large scale.
This must stop.
there are two dimensions that data demand is scaling on tremendously: horizon and realism of tasks/data points.
environments built on top of anonymized real operational data is step function change in scaling on both of these dimensions.
and as we do this, companies that choose to contribute can make data a very meaningful portion of their revenue while accelerating their path towards becoming more AI native.
thanks for covering Stephanie!
Great breakdown by @BerntBornich of how foundational world models, and the data needed to build them, will lead to functioning robots and real deployments
the immense investment in AI has not translated into adoption at the scale it appears to have.
across comparable early periods, AI investment been growing approximately 40%/year, versus 25% for electricity and 16% for IT.
however, only 19.8% of U.S. businesses use AI in at least one function, roughly 5 points below both benchmarks.
meanwhile the real production deployment gap is even deeper at top companies: recent MIT study shows that only 11% of the S&P500 has "deeply" integrated AI.
capability is advancing faster than trust and real implementation.
there's one fundamental reason for this: enterprise investment in evaluations is less than 0.1% of where it needs to be.
once large enterprises start investing in continuous evaluation loops, they will precisely define what "good" means for their use case. within that context, they will measure the intelligence of any given system. when you measure something, you can improve it and watch performance increase against those measurements.
this continuous loop is how enterprises "own their own intelligence."
it's irrelevant whether you're building on a closed or open model. in most cases, closed models will perform much better for your use case. what matters is whether you're precisely defining what good means and consistently measuring against it. if you are, then you're owning your own intelligence in a largely model-agnostic way.
Expert-level reasoning layered with real world data directly translates to increased capability in embodied models.
Real expert reasoning, synced with the work that is happening is likely most accurate route for models to understand constraints, failures, and logic.
We need much more of it for robots to be deployed.
Join us for a live session on the micro1 forum about our Company Data Partnership Program.
We’ll cover how micro1 is partnering with companies, paying $100K–$2M+, for their real-world business workflows to help train the next generation of AI models. We’ll also dive into how the program works, common questions companies have, and what participation looks like in practice, from getting started to seeing real returns on your company’s data.
Featuring:
- Jerome Josz (CFO at micro1)
- David Remland (Human Data at micro1)
This session brings together our core team to discuss how companies can leverage their operational data to train frontier AI models and create new streams of revenue.
Join us on 7/22, 12pm PT: https://t.co/la7B7g2Yfw
Remember when you were told "videos games are stupid, do something useful"? Today micro1 is paying people the equivalent of $170K salary to play video games. These gamers are creating the valuable RL environments needed to train the most revolutionary technology mankind has ever seen.
what a world we live in
Kimi K3 is another reminder that high-quality data is a core determinant of model performance.
@micro1_ai took an early, principled stand not to provide data to foreign adversaries.
I’m not convinced everyone else in this category made the same commitment.
Every company in the data ecosystem should publicly commit to this same standard. If you haven’t, why not?
some human data companies work with foreign adversaries.
and the results show today in Kimi K3.
we believe it’s shameful to claim American AI dominance desires while selling millions worth of data to countries that we are in adversarial competition with.
people come to America to build extraordinary things for humanity. we must protect that brilliance through American AI dominance.
To perform real work, AI models need to understand how real companies operate: how decisions are made, how teams collaborate, and how work moves from start to finish.
That’s why micro1 partners with companies to license anonymized operational data that helps train frontier AI models. Companies accepted into the program receive $100K–$2M for approved data packages, with the potential for ongoing partnerships as new qualifying data is generated.
Know a company that could be a fit? Qualified referrals can earn up to $50,000.
Learn more via the link in comments.