Through conversations with @andrewho03 and others at OpenAI and frontier labs, one thing has become clear: a data company lives or dies by its ability to understand what good data looks like. At Fleet, we’ve upstreamed computer-use and domain-specific capabilities into mainline models, and carefully studied perf gains through model post-training and scalable oversight. Designing good RL data is deeply nonobvious, and Andrew and I first connected over exactly this during late-night dead hangs at the gym (his hang time is crazy!).
We are in the early innings of a new era of data (cc @willdepue), one where human ops doesn't scale across an expanding & uneven capability frontier. Better models create a Jevons-style effect — models make knowledge work cheaper, we attempt more of it, and the remaining work becomes harder and more contextual. The coding “slop” Andrew mentions comes as an evolution of how software eng happens now vs a year ago, and this pattern will repeat far beyond coding; moving the reliability gap from 30% to 90% creates harder and more creative data problems in every domain. The surface area of "general" in AGI is fractal.
Most datasets miss economically useful work because real workflows are dynamic, deeply contextual, and difficult to capture faithfully in gradable, semi-synthetic environments. Andrew is tackling this problem first in scientific workflows, and I'm excited to see him bring this rigor to biology. At Fleet, we are tackling other frontier domains and have built the research foundation and platform needed to turn real workflows into reliable training signal. If you’re at a lab that cares deeply about data taste—or an engineer or researcher who wants to build it—come work with us. My DMs are open.
The Redpoint InfraRed 100 is now live.
These are the companies building the infrastructure that powers everything happening in AI right now, from world models and agent runtimes to the sandboxes, databases, and security tools agents depend on.
Congratulations to this year's honorees!
Read the full 2026 InfraRed Report: our state of the union on AI and cloud infrastructure 👉 https://t.co/Y1y94ZwI5B
Excited to cohost "Forecasting as a New Frontier of AI" with Haifeng & co. We have an incredible lineup of speakers, sponsors, and (soon) papers.
See you shortly in Korea!
Introducing Catalyst, the agent layer for all of finance.
Turn any natural language idea into a live strategy: research, backtesting & execution.
Don’t get left behind. Waitlist open, join now.
Fun fact, when @andrewthezhou and I started fleet, a context hub called "FleetContext" was the first product we built to extend the capabilities of frontier models
a year later came @Context7AI and now we have chub!
Excited to share new benchmarking work from @fleet_ai & friends.
We challenge frontier models to draw!
Surprising, across the entire frontier, models are really bad. The ways they fail can teach us about how AI perceives our world 🧵
Gemini 3 Deep Think can help make things. 🧠
Here's our side project: We sketched a laptop stand and Deep Think coded that into an interactive prototyping tool. We used that tool to generate a STL file, which we sent to @fleet_ai. And now I have a new laptop stand!
What will you build? See the final photos below!
Today we are highlighting Fred - a former customer turned founding engineer
Hear about his journey from Global Head of AI at Macquarie to Founding Member of Technical Staff, in his own words:
In middle school, with money earned performing magic shows I opened my first investment account.
At Stanford, captivated by the markets, I realized a paradigm shift was on the horizon.
Today, I am excited to announce our $5.7M raise, led by @novaholdings.
With the developments in AI, I knew I could no longer wait to start @vigillabs. At Vigil, we are building the engine to understand markets in realtime, designed for our own traders.