fyi for people DMing me for “my” takes on data, or asking to send me samples - be aware this account is an agglomeration of four people behind one burner lol
fyi for people DMing me for “my” takes on data, or asking to send me samples - be aware this account is an agglomeration of four people behind one burner lol
@PhoenixDegen@pax21e8 Yeah that job is… not worth it. I know a lot of SPLs across Scale, Mercor, Handshake, and smaller shops. Mercor does seem to pay the best, to their credit.
This tweet is super tone deaf given the workloads placed on SPLs at mercor especially and will probably indirectly contribute to notable GM/SPL exits in the next month.
buying the spirit airlines data may not actually be as useful as google thinks. seems like they bought a bunch of chats & documents over the operating lifetime of spirit to presumably make environments out of it.
this will be difficult for the following:
(1) although far more realistic than a synth world, to create tasks that are fair you need to be able to filter through the noise of the env. in an env that large, creating verifiers is going to be extremely difficult because different information, precedents, and messages could lead the model in "fair" and "grounded" directions that the task creator may not expect.
(2) i'd be curious to see how good the data is given that all PII and other info has been scrubbed. much of the records seem to have been able to be done synth, in a way that can custom-tailor the world to make sure the noise is tuned correctly, whereas using the real-world stack introduces confounding complexity that could make the world less usable/lower signal, albeit the world is realistic enough to where the model wouldn't recognize it's being evaluated.
all this to say, i think more acquisitions for data will happen, but i think it's a lot more difficult to turn these into envs than people think.
@MattStergiou Trivets is actually owned by hospitals. Getting EHR out of there is notoriously impossible; Truvets has the incumbent advantage to end all advantages
Serious incumbents: Surge, Mercor
serious challengers: Turing, Fleet (!), Mechanize, Micro1
not serious challengers: AQ
will die with SFT: scale, probably handshake
legendary moat: Truveta
have no strong opinions on others
Every single startup selling AI Training Data (July 2026)
>50 cos sell data and RL environments to big AI labs and drive AI progress behind the scenes.
They total ~$8.5B in rev and ~$100B in valuation, >75% of which are just 4 players: Scale, Surge, Mercor and Handshake.
They're bearing long-term because you won't need to assemble 1000 people in the future. Long horizon is already at a place where it's more cost effective to hire 50 full time labelers over 5000 mediocre ones - and managing 50 people is something the customers can feasibly do.
i don't understand why people are bearish on data making companies
what data is needed is super unclear and unpredictable. it's highly technical, high taste, and research-heavy. you can't just assemble 1000 people and say "make things"
idk feels to me like there are huge moats in having largest scale workforce, having the best researchers to understand what is important, building relationships with like the 5 labs that exist that are hard to buildup etc
also seems hard for the labs to just do everything inhouse, it's highly complex process and you kinda need founders attention to it
My current thought: AIs are never RLed on working with real humans; that would be expensive. AIs are never RLed on a task where they can consult a human and get help. So AIs solemnly debate among themselves, and 0 in 10,000 consider "talk to a human" as an option.
@danielrupawalla If you see how labs are trying to allocate/target spend, long horizon is everything. By EOY the sweatshop model, already weak, is probably dead because it isn’t useful.
@danielrupawalla Eh, there’s no “they” at these companies. For better or for worse accounts and projects are often client-scoped and build bespoke all the time. I’m quite familiar with how these work given I’ve seen samples and talked with teams from them all.
@danielrupawalla tldr: I think this take is about 1 year behind the actual curve of what sophisticated clients have been demanding, and therefore what data vendors actually are moving to deliver now
@danielrupawalla this is what we discovered when talking about moving some of these operations in house. imo, the days of ops people overseeing a swarm of low value talent are abs, abs over. the biggest data companies understand this and already treat simpler work as legacy to fund complex moats