Our newly accepted #COLM2026 paper explains how & why LLMs fail at large optimization tasks (1000s of variables). We show how to fix these issues while being efficient with token usage. 🧵
If you miss us, message me or send an email! We’d be happy to talk after the conference as well. My wonderful coauthors: Ilias Diakonikolas, Mingchen Ma, @fredsala!
Excited to be presenting an oral and poster at #ICML2026 about hybrid models and their capabilities! There have been many empirical results, but far too few theoretical ones explaining their expressivity. We show there are tasks with a separation! Come chat! Oral 6A, poster 4622.
Nested models let you train a whole family of submodels at once. What if you could use them all at once, too?
Block triangular weights enable this structure.
It gives us token-adaptive routing, self-speculative decoding, and more.
Introducing: Fully Nested Transformers (1/9)