We taught an agent a fixed set of verified expert corrections, then changed only how they were delivered. Force-fed into every prompt, they made the agent worse than the untaught model. Consulted on demand behind an applicability guard, they made it better. A tenth of the composite score, from delivery alone, crossing zero.
The guard's value tracks the base model's ability to judge relevance: it converts corrections into a gain on GPT-5.5 and only rescues GPT-4.1 from a collapse back to parity. Guarded recall should get better as models improve.
Full post → https://t.co/1x6qltWZJO
excited to announce that Phyvant is joining @a16z@speedrun in SF this summer.
@nihalgunu and i met in high school playing basketball and doing speech and debate. our team was ranked 4th in the country. somewhere between tournaments we fell into a habit we've never broken: pick something neither of us understands and go all the way down.
extremophile bacteria in antarctic ice. why chess engines still can't explain human intuition. federated learning across hospitals holding the same rare autoimmune case.
that habit took us further than we expected. we published in robotics and astrophysics journals. we worked at NASA and Mercor. we've done space and army funded research.
we also spent time inside some of the largest private companies in the world, watching how work actually gets done at that scale. the thing you notice, over and over, is how much of it lives in expert's heads.
today we're leaving Stanford and Purdue to work on @phyvant full time.
thank you to the entire @speedrun team. special shoutout to @JoshLu and @justmazer for being our first believers, @_CallMeMacy and @custo_lejla for the enterprise sales help, and @Chen for the sharp eye on this post and our website. And @SamiraBehrouzan and @TheRioDeGennaro for the photo help.
we help enterprises own their intelligence. if that problem sounds familiar at your company, let's chat.