Want to build your own coding-agent UI?
Ante's mini-tui example is one Rust file: streaming replies, tool activity, y/n approvals, and Esc to interrupt.
The agent runs underneath.
Start there and make the interface yours
What stands out to me is how much they’re squeezing out of just ~4B active parameters.
You’re getting performance surprisingly close to a 32B dense model, without blowing up the KV cache, while also making training more stable.
It feels less like just another MoE tweak and more like a different direction for scaling sparsity.
If you’ve ever had to size a serving cluster around KV-cache usage, you’ll know why the “no extra KV-cache cost compared with GQA” part is probably the most important detail here.
Each token only uses 4 out of the 64 value experts, so about 6.25% of the pool is active. The value vectors keep the same dimensions, which means the KV cache stays about the same size as standard GQA.
The pre-training loss curve is pretty close to their 32B dense model, even with only ~4B active parameters per token. The dense model is still ahead, but the gap isn’t that big
A lot of architecture papers propose ideas that don’t play nicely with FlashAttention. This one is different—it keeps FlashAttention, GQA, and sparse attention all in play.
no telegram. no alpha group. no private channel before launch.
everything about $SEAT is posted in one place: @100seatsxyz.
if you're not following it, you're not early. you're just not informed.
K2-Horizon-36B-A4B shows how AI scaling is evolving beyond simply increasing parameter counts.
With 36B total parameters but only ~4B active per token, it aims to deliver strong reasoning while keeping inference more compute-efficient.
The standout piece is MoVA — introducing sparsity into the value path of attention, alongside the traditional MoE architecture in the FFN.
A fascinating direction for building larger, more capable AI systems without scaling compute requirements linearly.
The efficiency game is moving fast.
K2-Horizon-36B-A4B is reportedly delivering performance comparable to models 20× larger, according to IFM.
Artificial Analysis places it #4 among 142 comparable models, while only ~4B parameters are activated per token.
A strong example of how smarter architecture can push AI performance without scaling compute linearly.
MoE introduced sparsity into the feed-forward layers, and MoVA takes that concept further by bringing expert routing into the value computation of attention.
What’s interesting here is the new dimension of scaling: rather than simply adding more FFN experts, @IFM_AI is exploring how model capacity can expand through the attention mechanism itself.
A promising architectural direction for building more capable models while keeping compute efficient.
Putting @MiniMax_AI H3 through its paces - seamless omni reference, dynamic motion graphics, and crisp typography control all in one workflow. 🎬
Think you can cook up something insane? Enter the MiniMax H3 x Picsart Challenge!
• $50,000 in prizes 🎁
• Submissions end Sep 23 ⏳
• Join now: https://t.co/TehYAhtxKm
cc: @Picsart
Claude can now build your entire mobile app, like an Apple-level $350K dev, in minutes and FOR FREE.
What used to require a full team and weeks of work can now be done with just a few well-thought.
Want the full guide??👇
Comment "Need" . I'll send you Everything