From the 2017 VLDB paper to Senior SWE-Bench and beyond, Snorkel's vision has stood the test of time. Big shoutout to the team that spent many nights and weekends making this happen and proving that vision once again. https://t.co/5Aj3HNlnJd https://t.co/KcglTRprGC
I'm excited to announce @SnorkelAI's $350M Series E at $3.5B, led by @insightpartners and @S32_VC.
We've grown 18x+ in the last 12 months since launching our Data-as-a-Service offering, passing $375M ARRR this week.
As AI advances to superhuman capabilities, AI data & environment development must advance with it - and basic staffing and crowdsourcing approaches are not enough.
AI progress now requires deep research and technology work that combines human expertise with specialized AI in compounding ways. @SnorkelAI is building the RSI data engine and frontier data lab for this next phase.
We're honored to have the support of existing investors Addition, @lightspeedvp, @GreylockVC, @GVteam, P7, Factory, @WellsFargo, Walden Catalyst Ventures, and new investors @ThirdPointLLC, @MarchCPs, @BlumbergCapital, @AllegisCapital, @Frontlinevc, and @standard_vc.
–
@SnorkelAI started as a research project a decade ago at @StanfordAILab.
Our thesis was simple: AI progress would become increasingly data-centric – and therefore data development should be studied as a true research and technology problem, not just a staffing and crowdsourcing one.
Today, as AI capabilities verge on superhuman, building the data and environments to safely measure and train AI is becoming too hard for even the smartest human experts to do alone.
Only humans and AI agents, collaborating together in compounding ways, can meet the accelerating needs of the frontier, and keep humans in the driver’s seat of AI progress for decades to come.
At @SnorkelAI, we are building the data lab to define the shape of this new “Data 2.0” frontier, and the new paradigms of human-computer interaction needed to advance it.
Our key focus is building the RSI engine for data, where specialized AI models accelerate and improve human expert output, and in turn, scaled human supervision is used to continuously evaluate and improve these models – creating a powerful compounding loop to keep pace with an accelerating RSI frontier.
With this round of funding, we are also doubling down on our commitments to support data development for open benchmarking and evaluation (more news here soon!); an increasingly diverse ecosystem of general and specialized intelligence; and a path to safe, well-aligned AI built on robust training and evaluation data.
Data development will guide and drive the next stages of AI – and must do so in a human-centric, AI accelerated, open, diverse, and safe way. We are excited to support this mission in the next decade of research ahead at @SnorkelAI.
More thoughts here: https://t.co/Dzk6olqmAc
@ajratner@SnorkelAI@insightpartners@S32_VC From the 2017 VLDB paper to Senior SWE-Bench and beyond, Snorkel's vision has stood the test of time. Big shoutout to the team that spent many nights and weekends making this happen and proving that vision once again.
@still_boneless Your system is flawed, sir! You omitted the OGs who arrived across the Bering Strait several thousand years before 1607. Your A should be a B and so on, assuming meaning to an inherently arbitrary hierarchy, of course. Reason itself speaks against you!
The "evaluation gap" is the silent killer of AI progress. We’re shipping models faster than we can reliably test them. This $3M Open Benchmarks Grant from @SnorkelAI and partners is exactly what the ecosystem needs to increase rigor in frontier AI. 📈🚀
Our ability to measure AI has been outpaced by our ability to develop it, and this evaluation gap is one of the most important problems in AI.
Today we're launching Open Benchmarks Grants — a $3M commitment to fund open benchmarks for frontier AI and close the evaluation gap.
Grateful to be partnering with @HuggingFace, @togethercompute, @PrimeIntellect, Factory HQ, @harborframework, and @PyTorch to back the teams building these benchmarks! 🚀
Word games yield valuable insights when evaluating LLMs. We built the SnorkleWordle benchmark to test models on 100 rare English words—and the results are 🔥
Excited to see @SnorkelAI continues to lead the way in developing specialized, high-quality data, the key ingredient needed for expert-level agentic AI!
Agentic AI will transform every enterprise–but only if agents are trusted experts.
The key: Evaluation & tuning on specialized, expert data.
I’m excited to announce two new products to support this–@SnorkelAI Evaluate & Expert Data-as-a-Service–along w/ our $100M Series D!
---
Snorkel Evaluate is our new data-centric agentic AI evaluation platform for specialized, mission-critical enterprise settings where vibe checks and out-of-the-box metrics driven by simple LLM prompts are not enough.
Snorkel Expert Data-as-a-Service is our white glove service for expert-level AI datasets, powering frontier LLM developers in areas like expert knowledge, reasoning, agentic action and tool use, and more!
Both built on top of @SnorkelAI’s Data Development Platform, using our programmatic technology to drive higher-quality expert data, faster– for getting specialized AI to real production value.
If you’re building enterprise AI and want to partner around the key ingredient in AI today–the data–book a demo and let's talk! https://t.co/w0J8izpn8p
Finally, see thread for details on 🧵👇
- 📽️ A walkthrough of Snorkel Evaluate and Expert Data-as-a-Service on an agentic AI enterprise task
- 📅 An upcoming event on Enterprise Agentic AI with innovators from @Accenture @BNY @Comcast@Stanford@QBE & others
- 📊 An upcoming series of benchmark datasets and model artifact releases
👀 Want early access to the full agentic AI dataset? Retweet this post and we'll send you the link!
There's a shocking fact about AI that nobody tells you: You can catch up to the public AI research frontier in just 2 weeks. Yes, really.
I've built a $150M annual revenue startup over the last 8 years and If I were to start a company today, I’d drop everything and go all-in on AI.
But like many busy software builders, I felt lost—overwhelmed by the noisy, crowded and fast-moving modern AI landscape. And I wasn’t alone.
So I spent my entire holiday diving deep into AI research—reading 30+ papers, watching hours of lectures, analyzing trends, and catching up to the research frontier.
✨ Here’s what I learned:
- You don’t need months (or years) to catch up.
- You don’t need a PhD or decades of ML experience.
- You need fewer than 20 papers and 2 weeks to understand the major breakthroughs shaping AI today.
It's because the technology is extremely nascent and most techniques that came before are no longer relevant:
- ChatGPT is barely 2 years old and Transformers are only 7 years old.
- Most game-changing discoveries happened within the last 4 years, driven by a few breakthrough ideas, scaling laws, and efficient matrix multiplication.
The biggest secret?
Many groundbreaking AI papers with thousands of citations are surprisingly simple and applied, like adding "let's think step by step" to the prompt, or simply asking the LLM over and over again to improve its answer (Self-Refine).
I realized there are tons of founders and builders in the same boat—wanting to dive deeper into AI but unsure where to start.
I've created an essential AI Guide that helped me catch up, in just 2 weeks, to the frontier of public AI research to figure out where the next opportunities and gaps were:
- Curated list of only the most important papers
- Simple explanations of key concepts
- Clear pathway to understanding the frontier of modern AI
It’s perfect for:
- Founders expanding into AI
- Builders wanting to innovate at the frontier of AI
- Investors looking to separate the signal from the noise
👇 Want the full guide?
- Like and Share this post
- Comment "AI Guide"
- I'll send you the complete guide
(ps, I’m also teaming up with @VishalVasishth, co-founder of @obviousvc with @ev (focused on large-scale societal impact companies like Twitter, Medium, Beyond Meat), to host a small meetup to discuss what's working and needs to be solved in the AI stack in SF. Message me if you're interested)
After seeing “Hidden Figures” some think we did these calculations manually. Not so. Not in our head, not by hand, not with a slide rule. We used what were for the time high speed computers. Development on Univac 1108. Then into the IBM 360 at NASA’s Real Time Computer Complex.
Just a few hours ago, the Centers for Disease Control and Prevention, the CDC, announced that they are no longer recommending that fully vaccinated people need to wear masks – whether inside or outside.
BREAKING: More than 600,000 Georgians have requested their mail ballots for the January 5 runoff elections.
Help elect @ReverendWarnock and @ossoff to the U.S. Senate by requesting your ballot today ➡️ https://t.co/xCyh7BhY3o. Happy voting and let’s get it done... again. #gapol