Through conversations with @andrewho03 and others at OpenAI and frontier labs, one thing has become clear: a data company lives or dies by its ability to understand what good data looks like. At Fleet, we’ve upstreamed computer-use and domain-specific capabilities into mainline models, and carefully studied perf gains through model post-training and scalable oversight. Designing good RL data is deeply nonobvious, and Andrew and I first connected over exactly this during late-night dead hangs at the gym (his hang time is crazy!).
We are in the early innings of a new era of data (cc @willdepue), one where human ops doesn't scale across an expanding & uneven capability frontier. Better models create a Jevons-style effect — models make knowledge work cheaper, we attempt more of it, and the remaining work becomes harder and more contextual. The coding “slop” Andrew mentions comes as an evolution of how software eng happens now vs a year ago, and this pattern will repeat far beyond coding; moving the reliability gap from 30% to 90% creates harder and more creative data problems in every domain. The surface area of "general" in AGI is fractal.
Most datasets miss economically useful work because real workflows are dynamic, deeply contextual, and difficult to capture faithfully in gradable, semi-synthetic environments. Andrew is tackling this problem first in scientific workflows, and I'm excited to see him bring this rigor to biology. At Fleet, we are tackling other frontier domains and have built the research foundation and platform needed to turn real workflows into reliable training signal. If you’re at a lab that cares deeply about data taste—or an engineer or researcher who wants to build it—come work with us. My DMs are open.