one intuition behind this is that linear attn can act as a learned, data-dependent positional bias. in fact they conclude that rope is not only redundant, but hurts the performance
paper: kimi linear (https://t.co/i6RdQLEsTU)
previously i thought linear attention was about trading off some long-context performance for major efficiency/computational gains. turns out a mix of linear attn and mla can outperform pure mla on long context
@vvvincent_c Very nice breakdown! I was too skeptical when I forst saw the PR, as it didn't even include logs at the time. Btw, the "does it scale" part is a great idea.
@Grad62304977 It can still be dynamic between conversations though? Like effort level set on the start of the conversation. Still much less flexible than CoT, but probably worth noting.
@ethanmclark1 Not true that representation prediction is the only one doing self-supervised pretraining and then SFT on top. Cosmos uses the exact same recipe on a high level.
New NanoGPT Speedrun WR at 79.2 (-0.5s) from @sisovicm, by moving the MLP up projection to FP8 on the forward pass, and increasing step count by 5 to offset the loss increase. The march towards lower precision continues! https://t.co/HSaA2PR6gv
@classiclarryd@yzhang_cs Wow didn't expect such speedups to be possible still! Looks like residual mixing keeps on giving good results, and is maybe even still underexplored.
@iamgrigorev Very cool! Why were you hesitant, and how are the loss curves looking compared to mHC, or do you not have clean comparisons since mHC proved too complex.
@eliebakouch It's interesting to compare to Kimi K2.6, also a post train extension of K2.5: Terminal-Bench 2.0 66.7 and SWE-Bench Multilingual 76.7. While the scores are better they aren't dramatically so.
@Dorialexander Although to be fair, dense MLP could also be viewed as having "soft" routing, so I could also see this MoE effect being smaller than expected.
@Dorialexander Very interesting, so it could be that MoE is more robust to variance in the training distribution because of training dynamics.
The router groups similar tokens, and thus similar sequences, to the same experts, so anomalous modes are localized to just a few experts.