After the Fable drama, I really get the incentive behind local open-weight models. I'm into the local direction myself (GLM/Qwen on Mac/5090), but the gap is just too big: a single slot doing ~100 tps under 128K versus hundreds of agents running in 1M context. The first is a cute demo, the second is real productivity. So the day frontier labs pull the plug on their inference APIs, our productivity takes the hit and local models aren't ready to catch us. Inference compute is a strategic advantage, and it's undervalued.