The efficiency bar just moved again. IFM says K2-Horizon-36B-A4B is matching models that are over 20× larger, and AA ranks it #4 out of 142 comparable models, while only activating around 4B parameters per token.
Worth being precise here: 36B total parameters, 4B active, 25 on the AA index, and #4 out of 142 in its class. It’s not #1, but the interesting part is getting that kind of result with only 4B active parameters.
If you were testing a model like this, what would you look at first—the benchmark results, how the KV cache holds up with long contexts, or whether the training-stability claims actually check out?
At 1.1M steps, it’s only about 1.2–3.7 points behind the 32B dense model on MMLU, GSM8K, HellaSwag, and HumanEval. Dense still comes out ahead, but the real question is whether that small gap is worth using 8× more active compute
MoE has been living in the FFN for years. MoVA basically takes that same idea and applies it to the value vectors in attention. Same concept, just a different part of the model.
Cutting gradient clipping by 50% might not sound like a big deal, but it can be the difference between leaving a training run overnight and having to babysit it.
The FFN has been the main place we’ve been scaling sparsity for a while. Was there a specific reason attention was left alone, or was it just harder to make it work there?
It’s a different way to think about sparsity: instead of adding more experts to the FFN, you’re putting the experts into the value computation. I’m curious to see how far this approach can actually go.
Vibe-coding changed how we build software. Pexo just brought that same energy to video creation. No complex prompts. No new tools to learn. Just share your vision, message Pexo like a friend, mark up the video and tell it what to change — like leaving a comment in a doc. Visuals, motion graphics, music, captions, voiceover, even a real person on screen — all made with Pexo. This is what video creation should have felt like from the start.
Our founder actually appears on screen in this video — and here's the thing, he didn't film with a camera, he didn't edit anything. Every frame, every motion graphic, the music, the captions, the voiceover, all of it came from a conversation with Pexo. Mark what you want changed directly on the video, just like commenting on a doc. If you vibe-coded your product, this is how you vibe-create the story around it. 🎬⚡
Every visual. Every motion graphic. The music. The captions. The voiceover. Even the founder on screen. None of it filmed. None of it manually edited. All of it built through a conversation with Pexo — no prompts, no new tools, just tell it what you want changed like you're leaving a comment. This is what AI video actually looks like when it's done right. 🎬⚡
The biggest barrier in AI video has never been the output quality — it's been the interface. Complex prompts, unfamiliar tools, no intuitive way to iterate. Pexo solves that by making video creation feel like a conversation. Share your vision, comment on what to change, refine naturally. If vibe-coding made building software accessible to everyone, Pexo is doing the same for video.