Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
@zerotomasteryio@AndreiNeagoie Went from programming newbie to understanding computer architecture thanks to Zero to Mastery's amazing Full-Stack & DS&A courses! Unlocked new career paths & entrepreneurial dreams. Ready to code the future! #TechTransformation#LearningJourney
1/ In 2014, I (@AndreiNeagoie) was in your shoes 👟.
I wanted a career change.
I wanted to get into tech.
So I taught myself to code using only free resources and got hired in 5 months.
You can too.
This is how I'd do it if I had to start from scratch in 2024.
Let's go 👇🧵