This is the drum I’ve been beating for the past year - stealing enterprise data through contractors the way Mercor, buying the codebases of dead startups the way Handshake and others are - is cope for the fact that nobody has designed the right business model to incentivize companies like Ramp (that have some of the best people for a given domain, who use models day-to-day and know their failures inside-out) to partner with a third-party that can translate this feedback into RL envs.
Such a partnership is needed because paying people full-time to generate RL envs creates envs that don’t reflect the frontier, because these creators grow disconnected from the real world. And because it’s more economically productive for now for SWE’d and other domain experts to actually work in their own domain full-time vs data labeling. E.g the best lawyers will earn more from being a lawyer than a data labeller.
This is why I call this feedback ‘exhaust’ - it’s a by-product of doing the real work that is very economically valuable. Designing a partnership model where one resists the urge to pitch their investors that they’ll replace the other party is key.
This is both a technical breakthrough (automating the above process) but also a business model and culture breakthrough. This is what I want to do for the next decade of my life because it’s the only way to access the highest quality RL data and to avoid adverse selection in data quality. If this sounds interesting, I’d like others to bounce ideas off of too.
Bearish. Thesis of the company is that an alternative to Cowork/Codex needs to exist because those two are too confusing for non-engineers, yet the video is obviously targeting towards founders.
3 months ago i resigned from @OpenAI. today we are launching Energy
everyone will soon use AI for all work on their computers, but very few do today even though AI is good enough
our users do days of work in hours. we'll onboard the world to working with ai
@catboosted This assumes that it’s hard to scale humans and quality, and anyone who’s truly in the know knows that there’s atleast one company that’s already solved this.
Introducing PG-LLM, a benchmark testing if general-purpose LLMs can predict protein variant effects.
Across 217 tasks, Claude Opus 5 (Max) leads all tested LLMs at ρ = 0.406 and outperforms 49 of 95 specialized protein predictors when ranking 50 variants.
Unit economics don’t matter if your models suck ass, your product sucks ass, and there’s no demand for the unit - which is what will happen if DM doesn’t spin out and why @matt_slotnick is right.
GDM procurement teams for data are bad, they bleed talent to every other lab right now, internal politics are so strong, and there’s no Elon to clean house.
@philjacobson@nachkari You can also get a private chef to cook you and your team meals for like $28 a meal.m fresh daily (which is what we do right now).
Never before in the history of capitalism have founders and their investors been so oddly joyful that parts of company are made obsolete by a supplier integrating vertically upwards. Stockholm Syndrome.