@aimalysheva@mostik_ai So, if we take Matryoshka Transformer architecture models like Gemma 4 with larger E4B and smaller E2B (both sharing the latent space by design), then after prefilling the prompt with E4B and generating answer using E2B will generate at E3B level, right?
https://t.co/eV8pJE7sro
Very nice MTP modification for Gemma4
1. Uses 4 decoder layers with 3 SWA + 1 GA
2. After embedding + hidden concat, downprojects to 256 (instead of hidden size) which makes computation much faster. Before LM head, uprojects back to hidden size.
3. Much more efficient LM head in the MTP by selecting clusters of logits!
Definitely worth looking into
@rohanpaul_ai@grok does one need to train this model for each new location or if the router position was changed, or a single pretrained model works universally?
Tencent introduces Continuous Autoregressive Language Models (CALM)
A paradigm shift from discrete next-token prediction to continuous next-vector prediction, making LLMs ultra-efficient. It compresses K tokens into a single vector, reducing generative steps by a factor of K with 99.9% accuracy.
@grok @Peasant3000 @alexwei_@OpenAI@elonmusk Good job! But:
1. doesn't your bright mind think that "solving" problems after their solutions were published along with the problem descriptions is not quite impressive?
2. Also, what shouldn't you forget to take to understand that?
@grok @Peasant3000 @alexwei_@OpenAI@elonmusk Ok. Solve the 1st one. But I bet that without pills you won't even be able to find the problem description