Introducing K2 Horizon: a connected fleet of six foundation models ranging from 0.9 billion to 375 billion parameters.
- Frontier performance: Across coding and agentic tasks, K2 Horizon delivers top-tier performance in every size class—with the 0.9B, 3.7B and 7B models setting new state of the art at their respective scales.
- Radical openness: K2 Horizon represents the largest fully open-source model launch in AI history. The fully open code, training data and recipes are a significant step forward in transparency.
Launch page: https://t.co/gg0k803SbL
Tech blog: https://t.co/g35L5xMGdS
Hugging Face: https://t.co/3Lb28JhyG9
@nvidia 's $12.9B @huggingface deal is more validation of the critical importance of open source at the frontier. But open weights ≠ open understanding — real openness means the data, code & methodology too, so claims can be verified, not just trusted.
Today we're showing that at scale: K2 Horizon, 6 fully open models (0.9B–375B).
Read the full technical breakdown: https://t.co/n5Qs3xFTmA
Open models are essential to expanding access to AI and accelerating innovation around the world.
We're excited to help @huggingface scale its platform and community while preserving the openness, neutrality, and choice that have made it a trusted home for AI builders.
🤗💚
K2 Horizon 375B A23B, a new open weights model from UAE's MBZUAI, scores 47 on the Artificial Analysis Intelligence Index, with relatively strong agentic performance and a 30 point jump over its predecessor
K2 Horizon 375B A23B is an open weights Mixture-of-Experts model with 375B total and 23B active parameters from @IFM_MBZUAI, MBZUAI's Institute of Foundation Models. It scores 47 on the Intelligence Index, alongside models such as MiniMax-M3 (45, also a MoE with 23B active parameters), and a large upgrade from its predecessor K2 Think V2 (17, 70B dense model). It leads nearby open weights models on agentic evals and has a low hallucination rate, but trails on knowledge and the hardest reasoning evals. K2 Think V2 ranks among the most open models on our Openness Index; MBZUAI is updating the supporting documentation and code for K2 Horizon and we expect to add it to the Openness Index soon.
Key takeaways:
➤ Strong on agentic tasks, weaker on knowledge and deep reasoning. MiniMax-M3, a recent model that is close to it on the Intelligence Index, makes the cleanest comparison: K2 Horizon 375B A23B leads on GDPval-AA, our real-world knowledge work benchmark (Elo 1430 vs 1380), and on τ³-Banking (34.2% vs 15.3%), but trails on GPQA Diamond (87.3% vs 92.9%) and Humanity's Last Exam (32.0% vs 39.0%)
➤ Low hallucination rate, driven by abstention rather than knowledge. K2 Horizon 375B A23B attempts only 40% of AA-Omniscience questions, declining the remaining 60% rather than guessing. The result is a 26% hallucination rate, among the lower rates we have measured, while accuracy is 18%, essentially unchanged from K2 Think V2
➤ A new architecture over its predecessor. K2 Horizon 375B A23B is a 375B parameter Mixture-of-Experts model with 23B active, succeeding the 70B dense K2 Think V2, and extends context from 262K to 512K tokens. Its 23B active parameters match MiniMax-M3 (428B total, 23B active)
Key model details:
➤ Architecture: Mixture-of-Experts, 375B total parameters, 23B active
➤ Context window: 512K tokens
➤ Multimodality: Text input and output only
➤ Pricing and availability: Yet to be announced
➤ Licensing: Open weights (license details to be announced)
At the Institute of Foundation Models at @mbzuai, sovereign compute helps our researchers move faster from ambitious ideas to AI breakthroughs. Learn more about the infrastructure supporting our work in @core42_ai's latest case study.
🔗https://t.co/Y95MleND5H
Advance AI. Stay sovereign.
Introducing Sovereign AI Stories; a new series showcasing how organizations are turning sovereign AI into real-world capability.
Our first story features @mbzuai, the world’s first graduate research university dedicated entirely to AI.
🔎 The challenge
MBZUAI needed the scale and flexibility to move quickly from experimentation to large-scale training and inference, without compromising customer managed data zones.
⚙️ The Core42 Solution
Core42 provides MBZUAI with access to heterogeneous AI infrastructure and the operational capabilities needed to run complex AI workloads securely at scale.
This gives research teams the flexibility to use the right computing environment for each workload, while maintaining control over where critical data, models and research operate.
🚀 The impact
Faster experimentation. Shorter training cycles. Greater flexibility across AI workloads.
The environment has supported the development of models including Jais, Jais Climate and K2, while helping MBZUAI retain control of strategically important research and intellectual property.
→ Read the full case study: https://t.co/94atiMg3t4
#SovereignAIStories #Core42 #ArtificialIntelligence #MBZUAI #SovereignAI #AIResearch #UAE
Goal maps objectives into paths. Identity covers resources and capabilities—not consciousness. Configurator chooses how to reason.
GIC is research, not a finished IFM product.
Where does the organizing intelligence live?
https://t.co/7eZPKX17zw
https://t.co/3uUoZEd2Hs
Most AI agents are organized from the outside.
Prompts set goals. Tools constrain methods. Workflows set the sequence. Scripts decide when to stop, retry, or escalate.
The intelligence lives in the model—and in the scaffold around it.
An agentive system shifts that organization inward—not away from goals, tools, or people, but toward internal mechanisms that decide what must happen, what the system can do, and how much reasoning the task deserves.
GIC proposes three parts: Goal, Identity, Configurator.
A pixel can reveal a pattern—but not the history or sentiment behind it.
As MBZUAI President and University Professor Eric Xing (@ericxing) noted at the Embedded Vision Summit, world models are video models that capture context across generated visual sequences. That context turns observation into understanding—and helps AI reason about what comes next.
Learn more: https://t.co/JctZFrPsYQ
A visually plausible future is not enough. A useful world model should help an agent compare consequences before it acts.
For a more formal view of that shift from next-state prediction to next-world simulation, see WR-Arena:
https://t.co/k1DnWW0dNQ
AlphaGo offers a clean way to think about what world models are for.
Start from a current state. Test possible actions. Simulate the futures they create. Compare them against a goal. Choose.
That is a thought experiment before action.
Eric Xing’s (@ericxing) full LxMLS 2026 talk: https://t.co/YRHJa48V3z
As MBZUAI President and University Professor Eric Xing (@ericxing) noted at the Embedded Vision Summit, intelligence requires more than reading the instructions—it requires learning the physics and feeling the environment.
That’s the leap from seeing the world to understanding how it works.
Explore our work at https://t.co/JctZFrQ0Oo
Explore our world model, PAN, at https://t.co/vadchtkBla
You don’t learn to swim from the shore.
And AI can’t understand the world by only observing pixels.
It must learn how environments behave, how actions change outcomes, and how consequences unfold.
Understanding requires interaction.
LSH-MinHash dedup may be throwing away millions of valid training documents.
The culprit: treating overlapping duplicate buckets as transitive equivalence, then keeping just one document per connected component.
Local overlap ≠ global duplication. Here’s the fix 🧵
The method keeps the existing LSH-MinHash candidate-generation pipeline.
We replace only Stage 3:
connected-component unionization → bucket-native layered greedy clustering.
It reaches at least 99.43% of the strongest computed upper bound across all full-corpus rows.