Through conversations with @andrewho03 and others at OpenAI and frontier labs, one thing has become clear: a data company lives or dies by its ability to understand what good data looks like. At Fleet, we’ve upstreamed computer-use and domain-specific capabilities into mainline models, and carefully studied perf gains through model post-training and scalable oversight. Designing good RL data is deeply nonobvious, and Andrew and I first connected over exactly this during late-night dead hangs at the gym (his hang time is crazy!).
We are in the early innings of a new era of data (cc @willdepue), one where human ops doesn't scale across an expanding & uneven capability frontier. Better models create a Jevons-style effect — models make knowledge work cheaper, we attempt more of it, and the remaining work becomes harder and more contextual. The coding “slop” Andrew mentions comes as an evolution of how software eng happens now vs a year ago, and this pattern will repeat far beyond coding; moving the reliability gap from 30% to 90% creates harder and more creative data problems in every domain. The surface area of "general" in AGI is fractal.
Most datasets miss economically useful work because real workflows are dynamic, deeply contextual, and difficult to capture faithfully in gradable, semi-synthetic environments. Andrew is tackling this problem first in scientific workflows, and I'm excited to see him bring this rigor to biology. At Fleet, we are tackling other frontier domains and have built the research foundation and platform needed to turn real workflows into reliable training signal. If you’re at a lab that cares deeply about data taste—or an engineer or researcher who wants to build it—come work with us. My DMs are open.
Every AI company should own its data, its models, and the research lab that keeps making them better.
Today, @thomasboser and I are launching @hiloopai to make that possible.
The frontier advantage isn’t access to a model. It’s the research organization continuously improving it.
Your product already generates the raw material for better intelligence: proprietary data, production feedback, evaluations, and domain expertise. Very few teams have the research capacity to turn those assets into better models.
Hiloop builds and operates that capability with you.
Bring us the model your product depends on, the data that makes it different, and an evaluation that defines success. We reproduce your baseline and run an autonomous research campaign against it.
Agents pursue competing hypotheses in parallel, build on previous results, and promote only improvements that survive verification. Our researchers validate the winners and bring them into production.
The work can run hosted or inside your environment. You retain control of your data, and the resulting models, evaluations, and research artifacts are yours.
We’re starting with model training, post-training, and inference optimization, where progress is measurable and the value is immediate.
But this isn’t a one-off model improvement. It’s a persistent research capability that begins each campaign with everything learned from the last.
Over time, the lab accumulates research memory and improves its own tools, evaluations, and agents. The process used to create better intelligence gets better itself.
That’s infrastructure for recursive self-improvement.
If your product depends on a model and an important metric has stopped moving, bring us the model you can’t make better. We're much better at research than making videos!
Reach out to us: [email protected]
Opus 5 reports 30% on ARC-AGI-3, ~4× the previous best model, ~20× its predecessor Opus 4.8. We tested it on Witness, our held-out suite of ARC-AGI-3-style interactive puzzle games. The leap doesn't transfer.
On Witness composites (same harness, same budget for every model), Opus 5 lands at 43.4 ± 3.2, which is a statistical tie with kimi-k3 (42.8 ± 1.9) and Fable-5 (43.8 ± 9.7). Ahead of Opus 4.8 (34.8), but nowhere near a generational jump.
The traces tell the why:
(1) On our most classic Witness-style game, Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration. It already knows this genre.
(2) But on our most novel game (unusual mechanic combinations you can't pattern-match), Opus 5 regresses below Opus 4.8. Where rules must actually be discovered through interaction, the new model is worse than the old one.
That decomposition (perfect on templates, regressed on novelty) is the signature of “scaffold-then-internalize” training on genre-specific data, not a general gain in interactive abstract reasoning. Our benchmark can't tell whether that data was their in-house ARC-AGI-3-like corpus with ARC-AGI-3-specialized harnessing (likely thanks to [schema]? https://t.co/6vrofY9Zx1), or public Witness-genre corpus, or both, but it can tell the improvement isn't general.
Held-out evals only stay held-out while nobody's optimizing for the genre, and that clock is always ticking. It's ticking for Witness too, the moment we publish it.
Nic and the team at Fleet are by far some of the smartest, nicest, and coolest people I’ve ever worked with. If you are impacted by the layoffs or know someone who is, please reach out, we are hiring and growing rapidly!
every conversation I have had with a researcher from AmazonAGI has left me with hope for the future of computer use
If you are impacted or know someone that is impacted by this, I’d love to help, be it finding the right role at fleet or at any of the lovely labs & startups I know
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final
((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
Ending ICML with a word of appreciation <3
We got to celebrate the six members of our research team with accepted work, from testing the ability of agents forecasting the future to finding novel LLM privacy attacks
we also got to meet many incredible researchers along the way <3
We do much more than sell training-data! And we are much much farther ahead than the numbers here show
Happy to see us recognized. Towards being the most profitable neolab! More soon
if you like watching videos on 2x speed you might like watching GPT-sol do things with CUA
here's GPT- Sol making a cannon in Blender (but this video is NOT sped up)
make CUA fast again