@torgeirlysen 1000% agree,, in retrospect i learned the most (both abt myself and the subject) during my MS when i took 3 PhD level courses along w my regular course
brutal but i dont regret it in the slightest
It’s in torchtitan. It’s literally implemented in deepspeed. Guys why build this ourselves it is up streamed in huggingface accelerate. Guys just use megatron. We can plug this into monarch, we are using fairscale. It’s all already done in lightning. Let’s just clone torchtune. We can drop this into axolotl, should take a few days max. I don’t understand why we can’t just build on mosaicml composer. I mean nanotron is right there. We can spin this up today with Nemo.
What if a model could adapt its compute budget on the fly?
Today, we face a rigid trade-off: either maintain costly fleets of specialist models, or rely on dynamic methods that fail to scale to large pre-trained LLMs.
We introduce Nested Subspace Networks (NSNs) to break this.