Today, LLMs are no longer built from human data alone. They rely on other LLMs to generate training data, filter corpora, evaluate outputs, provide rewards, and guide development decisions. So how many models and datasets is a modern LLM built on?
• OLMo 3 → 89 model + 183 dataset dependencies
• Nemotron 3 → 273 model + 560 dataset dependencies How did we find it out? We built ModSleuth. 🧵
LLMs are no longer created w/ human data alone. They rely on other models to generate & filter data, evaluate outputs, & guide dev work.
So what is a modern LLM built on? Olmo 3 → 89 model + 183 dataset dependencies; Nemotron 3 → 273 + 560
We made ModSleuth to trace this. 🧵
I’m at #COLM2026 this week! I’ll be presenting Branch-Adapt-Route on Thursday afternoon (poster session 6).
Come say hi and reach out if you want to chat!
Imagine you fully post-trained "YourModel v1". Then, you've got better data — math, code, tool use, safety — and you want to improve it.
Today, that usually means retraining the whole model.
But what if new data could be added modularly, with a fixed cost each time?
Could long-context architectures “find all contradictions” in a science literature? Not yet! 🧵
We study a new class of "high-complexity” tasks whose difficulty scales quadratically with corpus size (as opposed to linearly), reversing common LCLM decisions! (block-sparse attention, hybrid models…)
Today, LLMs are no longer built from human data alone. They rely on other LLMs to generate training data, filter corpora, evaluate outputs, provide rewards, and guide development decisions. So how many models and datasets is a modern LLM built on?
• OLMo 3 → 89 model + 183 dataset dependencies
• Nemotron 3 → 273 model + 560 dataset dependencies
How did we find it out? We built ModSleuth. 🧵
One day I tried tracing all of Olmo's dependencies manually. A few hours later, I realized I can't do it and gave up. Then @sadhikesaven and @CoderBak ModSleuth 🔥
Turns out Olmo and Nemotron have hundreds of dependencies that are super deep, recursive, and not easily visible. I'm glad I gave up early 😅
Spoiler: I thought this would be a one-week Claude Code project. It was not.
The hard part wasn't information extraction (which Claude Code is good at). The hard part was something much trickier. Check out the paper to learn more!
(And yes, if a model release says it used Claude Code, ModSleuth will trace that too... which means the model depends on Claude Code, which has its own dependencies, and ModSleuth itself depends on Claude Code 🤯)
LLMs are no longer created w/ human data alone. They rely on other models to generate & filter data, evaluate outputs, & guide dev work.
So what is a modern LLM built on? Olmo 3 → 89 model + 183 dataset dependencies; Nemotron 3 → 273 + 560
We made ModSleuth to trace this. 🧵
Today, LLMs are no longer built from human data alone. They rely on other LLMs to generate training data, filter corpora, evaluate outputs, provide rewards, and guide development decisions. So how many models and datasets is a modern LLM built on?
• OLMo 3 → 89 model + 183 dataset dependencies
• Nemotron 3 → 273 model + 560 dataset dependencies How did we find it out? We built ModSleuth. 🧵
LLMs are no longer created w/ human data alone. They rely on other models to generate & filter data, evaluate outputs, & guide dev work.
So what is a modern LLM built on? Olmo 3 → 89 model + 183 dataset dependencies; Nemotron 3 → 273 + 560
We made ModSleuth to trace this. 🧵
Across 4 open-source releases, ModSleuth recovers 1,060 source-verified dependencies, with chains up to 8 hops deep. This graph also surfaces findings that are hard to find manually:
• License-relevant multi-hop paths
• Train-evaluation coupling
• Mismatches between papers, cards, and code
Today, Human Archive is announcing our $8.2M seed round to model human embodied intelligence.
Despite decades of research, we still barely understand ourselves. Our goal is to learn how humans interact with the world, and over the past 6 months, our team’s made enormous progress toward that alongside leading AI labs.
learn more @TechCrunch
https://t.co/faLhyVBjl1
Human Archive (YC W26) is a research lab modeling human embodied intelligence, and we’re hiring globally across 10 roles in hardware, software, and operations.
We build cameras and sensors, deploy them globally at scale, and train models to better understand how humans interact with the physical world. Our goal is to replace manual labor, increase global abundance, shift human effort toward creativity and exploration, and advance how we understand the brain, human cognition, prosthetics, and rehabilitation.
By joining now, you’ll join the founding team and work on the most important data project in human history.
Software:
Machine Learning Engineer (SF)
Research Engineer (SF)
Hardware:
Head of Hardware Engineering (SF + China)
Embedded / Electrical Engineer (China)
Firmware Engineer (China)
Mechanical Engineer (China)
Head of Operations (China)
Operations:
Infrastructure Engineer (India)
Software Engineer (India)
Operations (Globally)
Apply here: https://t.co/7zwypcGjA9
Human Archive (YC W26) is a research lab modeling human embodied intelligence, and we’re hiring globally across 10 roles in hardware, software, and operations.
We build cameras and sensors, deploy them globally at scale, and train models to better understand how humans interact with the physical world. Our goal is to replace manual labor, increase global abundance, shift human effort toward creativity and exploration, and advance how we understand the brain, human cognition, prosthetics, and rehabilitation.
By joining now, you’ll join the founding team and work on the most important data project in human history.
Software:
Machine Learning Engineer (SF)
Research Engineer (SF)
Hardware:
Head of Hardware Engineering (SF + China)
Embedded / Electrical Engineer (China)
Firmware Engineer (China)
Mechanical Engineer (China)
Head of Operations (China)
Operations:
Infrastructure Engineer (India)
Software Engineer (India)
Operations (Globally)
Apply here: https://t.co/7zwypcGjA9
MoEs are everywhere in frontier models, and they are deployed as a monolith system.
But many applications only need a narrow slice of capabilities, e.g., math, code, biomedical, etc.
So what if "modularity" is actually the missing opportunity for MoEs?
Today, we're releasing EMO: an end-to-end pretrained MoE where modularity emerges naturally, enabling selective use of experts!
How do you add new capabilities to a fully post-trained language model, without retraining from scratch, or losing what it already knows?
We're excited to introduce Branch-Adapt-Route (BAR): train independent experts, merge them into an MoE, and upgrade them as needed.
Imagine you fully post-trained "YourModel v1". Then, you've got better data — math, code, tool use, safety — and you want to improve it.
Today, that usually means retraining the whole model.
But what if new data could be added modularly, with a fixed cost each time?
Last year, we introduced FlexOlmo, a novel way to train parts of a model independently then combine them later.
BAR builds on that idea for a harder problem: how to keep improving a model without having to retrain each time. 🧵
More broadly, BAR suggests a new way to build and improve upon LMs: not one monolithic pipeline that must be re-run for every update, but a modular system where experts can be trained, added, and upgraded independently.