The Kimi K3 architecture figure for yesterday's big open-weight model release, along with some observations and thoughts.
1. Yes, it looks relatively complicated, but it's essentially a scaled-up production version of their Kimi Linear model they released last year (scaled up from 48B -> 2.8T; K3 is by far the biggest open-weight model right now)
2. The one new component compared to Kimi Linear is the LatentMoE. I omitted it in the figure below since it's already very crowded, but that's essentially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are curious). The idea here is to compress (down-project) large linear layers similar to multi-head latent attention.
3. Kimi K3's overall trend (similar to Nemotron 3, DeepSeek V4, and others) is also towards better inference efficiency. That is, there are many components that replace existing components with efficiency-tweaked versions. I.e., MoE -> LatentMoE, regular attention -> multi-head latent attention and Kimi Delta Attention. (I also have short tutorials and write-ups in my gallery if you are curious about additional details).
4. The one component change that is not an efficiency tweak is attention residuals. Like DeepSeek V4 improved the residual path with mHC (manifold-constrained Hyper-Connections), attention residuals are a way to improve the residual path, but it works a bit differently. I.e., mHC made the residual path wider. Attention residuals (also already part of Kimi Linear) connect the residuals across layers; the connection itself uses an attention score for an important/contribution weight. According to the report, it improves the validation loss and downstream performance (a bit) consistently and adds about 4% in training cost and 2% in inference cost.
5. Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead. (Again, this is inherited from Kimi Linear). In other architectures, the recent trend was towards RoPE in local attention layers (like sliding window attention) and NoPE in the global layers. There were a few architectures that only used NoPE everywhere, but this is the first frontier-level one as far as I know.
6. Kimi K3 now also has native multimodal support, which is great!
There are several other interesting training tidbits in the technical report, but that's it from the architecture front so far. A really great release overall.
@thisisashukla@prathoshap Training data is all from youtube .... but very sure there is more than 5 hrs worth of chanting.
Have to take permission if have to build and release it to public.
I am moslty interested in the chanting style ... voice can be from anyone.
@thisisashukla@prathoshap Not any model ... this is prerecorded.
I was trying something so that me and my family could learn that.
Especially karaoke style. Easy to repeat and memorise.
@santalum_aurum English translations act as a vital gateway, sparking the initial interest that drives global learners to study original Sanskrit. Furthermore, modern scholar-led translations allow insiders to correct colonial-era Eurocentric biases and accurately protect the text's true intent.
@prathoshap : How can we train the model to render in Challakere brothers chanting style?
Is there a way for us to tweek and tune that?
Any documentation would be greatly appreciated for further finetuning on different styles.
Thank you for this great work.
A 15-year-old dream has come true today. I started a PhD with the dream of creating a system that chants any Sanskrit shloka perfectly.
And here I am opening sourcing 𝐕𝐚𝐠𝐝𝐡𝐞𝐧𝐮 - 𝐀 𝐯ṛ𝐭𝐭𝐚 (𝐦𝐞𝐭𝐞𝐫) 𝐚𝐰𝐚𝐫𝐞 ś𝐥𝐨𝐤𝐚-𝐭𝐨-𝐜𝐡𝐚𝐧𝐭 𝐭𝐞𝐱𝐭-𝐭𝐨-𝐬𝐩𝐞𝐞𝐜𝐡 (TTS) 𝐬𝐲𝐬𝐭𝐞𝐦 𝐟𝐨𝐫 𝐒𝐚𝐧𝐬𝐤𝐫𝐢𝐭. This is the world's first vrutta-aware, open-source TTS for Sanskrit Chanting.
@prathoshap@AnujKum82046422 Translation -
You bring up a daughter with lot of love and care. Then you give her away to a in-laws with money and gold. What did you expect in return? Nothing .... Similarly Virtue is its own reward. It is good for maturity of the mind. - Mankutimma
@SanjeevSanskrit Excellent summarization !!!
You cannot settle every question of dharma by appealing to one text one rule, or one authority.
Instead dharma always depends on
consequences,intent,circumstances,
preservation of the larger moral order, and Sanatana Dharma guides us there.
@Conor_D_Dart None of these are happenning.
Remember google is a listed company not some start up trying to win the model race that may last for few weeks at max.
Aquiring customers who jump between models like frogs will be at the bottom of their interest list.
Although they pretend its not.