meet @mostik_ai!
what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, which I can't say much about while the competition is still running. and this, which I can.
everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning?
we do this by enabling models to communicate in latent space. through our protocol, hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched.
how do we know it works? we tested it on a setup where a 753B model reads the problem, and a 4B edge-class model writes the answer. with this approach, we get results 80% as accurate as the frontier model, but at 20x faster performance.
we're committed to preventing frontier model lock-in and are already partnering with inference providers to accelerate open-weight adoption. we've done this between 15 of us, in four months, 12 PhDs and a Fields medalist, backed by @generalcatalyst
WIRED has the first external account of the company and the work: https://t.co/tP8nItCsDl
full writeup, the setup, and all the numbers: https://t.co/C9NZ5vtV1V
12 PhDs and a Fields Medalist in the mountains arguing about latent communication, and it happened to be my birthday 🎂
taking suggestions on how to top that 😂
Reading what AI doesn't say
(this topic genuinely matters to me, and it happens to be my birthday today, so if you feel like it, a repost would be a nice gift 🎂)
Something shifting in AI architecture that I don't think gets discussed enough through a safety lens. Models are starting to pass information to each other directly through internal representations, hidden states, activations, and computation vectors (still mostly research right now, but the infrastructure will follow). That matters for how we audit reasoning, since most of what we can currently check comes from what shows up in words.
Right now, the main tool we have for auditing AI reasoning is reading chain-of-thought. Korbak et al. (2025) spend a whole paper arguing this window is already "fragile" for single-model reasoning. If models start reasoning through vectors passed model-to-model instead of text, understanding that latent layer becomes the only way to keep any visibility into what's actually happening.
Thus I think that what seems underexplored is understanding the math of these latent spaces, and it might be the same research direction as learning how to monitor them.
Models appear to converge toward similar geometric structure regardless of architecture (Huh et al. 2024). If that's right, probes for safety-relevant features might generalize: deception patterns, goal representations, misalignment signals studied once and applied across models rather than re-derived per architecture.
Whether this holds for safety-critical features specifically is open, so is what adversarially robust latent decoding would even look like. Both feel more urgent than the current research investment suggests.
@IntuitMachine Hoping the AI with access to my memories never finds the part where I peed myself at summer camp laughing at my own joke. Actually wait...
@sriramk Greeks built an early steam-powered device almost 2000 years before the industrial revolution. It was treated as a curiosity, but that didn’t make steam power useless, it just meant society hadn’t yet found the right way to turn it into something.