Regarding the last topic @DavidSacks :
With respect, the data being used to train frontier models in the U.S. is not a commodity. It is American intelligence, paired with anonymized operational data from U.S. enterprises.
Put simply, we have top doctors, scientists, lawyers, and physicists in the U.S. using their knowledge to create detailed rubrics that train these models. Each individual data point and its corresponding rubric can take an expert anywhere from 10 to 30 hours to create.
This is not preference labeling or drawing bounding boxes. It is highly complex, structured human judgment from leading experts here in the West—people who deeply understand and directly contribute to the latest American innovations in their respective fields.
That expertise is then converted through highly specific data structures and RL environments (developed collaboratively by U.S. AI labs and data labs) from raw human intelligence into high-signal rewards that improve frontier models.
On top of that, any AI advancement, even something that begins as a simple chatbot designed to improve operations within a defense agency, can create a major competitive advantage in adversarial situations and may have dual-use applications.
Lastly, many datasets today are seeded with anonymized, real-world operational data from U.S. companies to build highly realistic environments. When those datasets are sold to China, we are not simply exporting “labeling.” We are exporting proprietary American intelligence, structured for machine learning and delivered directly to China at scale.
It is very easy to categorize this work as “labeling” and ignore what it actually represents.
However, as @altcap suggested: “Then these things will get a lot more scrutiny than they’re getting today. I think the only reason they pass muster today is because we’re still leading the race.”
If China catches up to U.S. labs, this will become much harder to ignore. In retrospect, the role of data as the root cause will become very clear — and by that point, it may be too late.
The race is tight. I suggest looking into this now.
Shameful act to optimize for short term revenue increases and serve an adversarial nation in the most important race of our lifetime.
The only way models improve is through data. if you have the recipe on what data pushes the frontier and send that to China, you are doing a disservice to U.S. AI Labs as well as the U.S. government.
Beyond that, a large portion of datasets today require anonymized real operational data from U.S. companies. Selling such environments to Chinese labs means exporting U.S. company data directly to China at a large scale.
This must stop.
happy birthday to the great United States of America. 🇺🇸🦅
I moved here from iran when I was 10 years old. every day since, i’ve felt deeply grateful for the opportunity this country gave me, and for the chance to contribute to its continued prosperity.
what an incredible country. the greatest experiment the world has ever seen.
Today we're publishing LongExtractBench, a benchmark commissioned by @reductoai and independently validated by micro1.
We evaluated seven production document extraction systems across the same 225 complex enterprise documents. The benchmark was intentionally difficult: documents averaged 358 pages and contained roughly 88,700 ground-truth fields each. Every system was evaluated using the configuration documented in the benchmark methodology.
Key findings:
• Reducto Deep Extract was the only system to successfully complete all 225 documents.
• Direct frontier LLM baselines achieved substantially lower completion rates on long, complex documents.
• In this benchmark, dedicated extraction platforms achieved higher completion rates than the direct frontier LLM baselines.
• Recall was the clearest differentiator. Precision remained high across systems, but recall ranged from 33.8% to 99.6%, highlighting which systems consistently captured the information contained in long, complex documents.
The full report includes the benchmark methodology, limitations, and reproducibility resources. Check out the report and results in the comments below.
We're hosting a forum to discuss micro1's Company Data Partnership Program this Thursday at 9am PT.
micro1 Founder & CEO, @aliansarinik , and Strategic AI lead, Soliman Aniss, will share insights on how we're partnering with 50 companies, paying $100K–$2M+, for their real-world business workflows that will help train the next generation of AI models.
Join us to learn how the program works, common questions companies have, and what participation looks like in practice.
Register here: https://t.co/ikrRdPogPy
Today, we’re committing $5,000,000 to launch the micro1 Company Data Partnerships Referral Program.
For every company you refer, you can earn up to $25,000. Simply introduce a company, have them identify you as the referrer during onboarding, and once they enter into a paid data partnership with micro1, you’ll receive your referral payout.
If you know a company that wants to turn its operational data into a recurring revenue stream while accelerating its adoption of AI through micro1's Data Partnership Program, we’d love an introduction.
visit /data to get started
Introducing the Realm Financial Reasoning benchmark, our new evaluation of frontier AI on reasoning in finance and spreadsheet-grounded analysis.
Tasks are built around the actual work product that practitioners deliver, from IFRS reconciliation workbooks and hedge-fund backtests to VC term sheet analyses and treasury cash-flow forecasts. Each task drops the model into a sandbox with the same source materials a human analyst would open: named-range Excel workbooks, broker PDFs, earnings call transcripts, monetary-policy decisions.
Here's what the results showed (Pass@3):
-GPT-5.5: 0.456
-Claude Opus 4.7: 0.398
-Gemini 3.1 Pro: 0.349
The three models score similarly, and none clears 50% on tasks that demand a judgment call. The back and middle office are defensible today, but on capital allocation questions current frontier models should be treated as research accelerators, not final decision-making support systems.
Full report linked in the comments.
Introducing Prospera: a benchmark that tests AI agents on real federal tax returns, designed by our research team in collaboration with CPAs and industry-leading tax professionals.
A complete federal return requires dozens of source documents, hundreds of interdependent calculations, and no room for errors. We evaluated GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro with no hints on which forms to file, scored against 20+ expert-authored criteria per return.
Here’s the Results (Pass@3):
-GPT-5.4: 28%
-Gemini 3.1 Pro: 18%
-Claude Opus 4.6: 16%
To put those numbers in context, the tasks in Prospera weren't obscure edge cases. Filing a federal tax return is something millions of Americans do every year, yet 44% of evaluation criteria failed across all models.
Full report linked in the comments.
micro1 Cortex: the human intelligence layer for high-performing AI agents.
Most agents look good in demos, but break in production.
- Benchmarks don’t reflect real workflows
- Internal testing misses edge cases
- Failures are hard to diagnose
- Performance degrades as you scale
Cortex leverages expert human judgement to evaluate, train, and continuously refine agents so they deliver exceptional performance in real-world workflows and drive measurable outcomes in production.
the micro1 robotics lab:
real world data for intelligent models that co-exist in the physical world.
we’re in-the-wild across 75 countries in 6,000+ unique environments collecting data. diverse movements, objects, and settings.
the future of AI is as human as you can imagine. join us to start training robots today (link in comments).
This Wednesday at 11am PT, Stanford Professor, Omer Reingold, and member of our technical staff, Nima Yazdani, will be on the micro1 forum to discuss bias and fairness in AI systems as they move from research into real-world deployment.
Register here: https://t.co/OrgTfRsvIb
He grew a company 35x in one year to $250M+ at 25. He’s the same kid who came to the US from Tehran at 10 without knowing a word of English.
This is the untold story of Ali Ansari and micro1, which just cracked the top 10 on the Lean AI Leaderboard with $250M+ revenue and 80 employees.
If you are building a lean AI company or want to sell to the top AI labs, read on for the full playbook.
Ali built 2 companies before college.
At Berkeley, he launched a software dev agency and started hiring international engineers.
The interviews alone were eating 30–40 hours/week, so he built a tool to automate them, using GPT (one of the first AI recruiters in 2022).
That insight eventually became micro1.
For 2 years, it grew steadily with two business lines (an AI interviewer SaaS and an engineering marketplace) with happy customers.
Then a data vendor approached Ali with an unusual request: hire 700 engineers to train AI models.
That one conversation changed micro1’s trajectory.
Ali realized they had accidentally built what every major AI lab desperately needed: a system to find, vet, and manage domain experts at scale across industries, at volume and fast.
He then made a decision most founders would never have the nerve to make:
He killed both working businesses and bet the entire company on going direct to the labs, with no safety net or guarantee that it would work.
But the bet paid off, resulting in 35x growth, as they went from $7M ARR to ~$250M in one year.
None of it came easy, and this level of growth became possible only after Ali solved the hardest problem in this space:
How to sell to AI labs where buyers are deeply technical and part of tight communities where reputation travels fast.
So I spent 10+ hours going deep into the decisions behind how micro1 built, sold, and scaled within the AI ecosystem and turned it into an actionable playbook for founders who want to sell to researchers.
Inside, you'll get:
• The 3-stage sales sequence Ali uses to close research deals like OpenAI and xAI
• How he got into Stanford research circles with zero connections (and how that helped him close deals)
• The proof of concept strategy: The dos and don’ts when researchers are evaluating you
• How Elon Musk accidentally handed micro1 their biggest sales breakthrough
• The net expansion playbook for enterprises and Fortune 500 companies (and what is converting fastest)
• The full AI stack that powers micro1's recruiting engine
• The incentive philosophy Ali rebuilds individually for every core team member every quarter
Originally, I put this together as a resource for founders I work with directly.
But the ideas and insights are too valuable not to share, so I'm giving it away publicly.
Founders who crack AI sales at this hyperscale usually keep it close, but Ali shared every piece of it.
So if you are building a business around frontier AI and research, grab this right away.
It will save you months of costly relationship mistakes (Link in the first comment).
Ali and team, welcome to the Leaderboard!
On our next micro1 Forum on March 6 at 11:30am PST, micro1 founder and CEO, @aliansarinik and micro1's VP of AI, @AndrewLeeMaas, will chat about why human judgment is becoming the most important layer in modern AI systems.
From training models to evaluating agents, the future of AI is ultimately shaped by the people who guide it.
Register here: https://t.co/kDqC1iUnea
.@APompliano is exactly right. while the scope of work of some jobs may suffer in the mid-term, entirely new categories of work are continuing to emerge. AI will ultimately create a lot more jobs than it displaces. and those displaced, in almost every case, will be evolutions of the same job.
Ali Ansari (@aliansarinik) runs the current fastest growing company in the world (for all nine figure run rate companies).
$200 million ARR growing 30% month-over-month.
This new industry of human data pipelines will shock you.