Today, we are releasing a new version of K2 (K2-V2), a 360-open LLM built from scratch as a superior base for reasoning adaptation, while still excelling at core LLM capabilities like conversation, knowledge retrieval, and long-context understanding.
K2 fills a major gap: highly capable models with no transparency. Instead of releasing only weights, we’re sharing the full training story — dataset recipes, mid-training checkpoints, logs, code, and evaluation tools. That’s 360-open.
What’s inside:
• 70B dense transformer engineered as a reasoning-enhanced base model
• Native 512K context (extendable via RoPE scaling)
• Mid-training reasoning phase
• Strong tool-use scaffolding
What we’re open-sourcing:
• 250M+ reasoning traces (math, planning, multi-step logic)
• Full pre- & mid-training data compositions
• All mid-training checkpoints
• Training logs, code, Eval360
Performance:
• GPQA-Diamond: 55.1% mid-training → 69.3% after SFT (strongest fully open 70B model)
• KK-8 Logic Puzzles: 83% — competitive with DeepSeek-R1 & OpenAI o3-mini-high
• ArenaHard V2: 62.1% — close to Qwen3 235B
• Outperforms Qwen2.5-72B and approaches Qwen3-235B despite being smaller and fully transparent.
🔗 The Model:
https://t.co/gsjRUwfnvN
🔗Technical Report:
https://t.co/oFZQuLQaNg
🔗Blog:
https://t.co/zQdpmLgEUt
K2-Think 32B, built on Qwen2.5
Scores (pass@1, avg of 16 runs):
- AIME’24: 90.8
- AIME’25: 81.2
- HMMT’25: 73.8
- Omni-HARD: 60.7
- LiveCodeBench v5: 63.97
- GPQA-Diamond: 71.1
It is trained with long CoT SFT and RL with verifiable rewards on the Guru dataset, then improved at inference through a Plan-Before-You-Think scaffold and Best-of-3 sampling, which also shortens outputs by 6–12%. Deployment on the Cerebras Wafer-Scale Engine achieves ~2,000 tokens/sec (32k ≈ 16s) versus ~200 tokens/sec (32k ≈ 160s) on H100/H200. Safety-4 averages 0.75, strong in refusal and conversational robustness, weaker on cybersecurity and prompt-extraction.
The model, training code, inference code, and full tech report are openly available, with the complete reasoning system also served as a live API and web portal at Cerebras-level speed.
𝖯𝖺𝗋𝗍 𝗈𝖿 𝗍𝗁𝖾 𝗌𝖺𝗆𝖾 𝖪𝟤/𝖫𝖫𝖬𝟥𝟨𝟢 𝖾𝖼𝗈𝗌𝗒𝗌𝗍𝖾𝗆 𝗍𝗁𝖺𝗍 𝖺𝗅𝗌𝗈 𝗍𝗋𝖺𝗂𝗇𝖾𝖽 𝗍𝗁𝖾 𝟨𝟧𝖡 𝗈𝗉𝖾𝗇 𝖪𝟤 𝖣𝖨𝖠𝖬𝖮𝖭𝖣 𝗆𝗈𝖽𝖾𝗅
A big milestone for advanced AI reasoning. K2 Think is more than just a model — it’s a full system with powerful capabilities and impressively fast inference. Try it out yourself! https://t.co/xUBSSsRDyH
Introducing K2 Think - a breakthrough in advanced AI reasoning.
Developed by MBZUAI’s Institute of Foundation Models and @G42ai, K2 Think delivers frontier reasoning performance at a fraction of the size of today’s largest systems.
Smaller. Smarter. Open to the world.
Available now: https://t.co/BeARU0JHF0
#K2Think #AI #OpenSource #MBZUAI #G42 #Innovation
The upcoming launch of the "K2 Think" model will mark a significant step in advancing artificial intelligence from the
UAE to the world, reflecting our leadership's vision for technological progress. K2 Think will combine the efficiency of smaller models with world-class performance and superior inference speeds, surpassing larger systems to set a new global benchmark for open-source reasoning
Developed through the partnership of @MBZUAl and @G42ai, it showcases the strength of cooperation between academia, public institutions, and the private sector in turning national aspirations into reality.
We congratulate the UAE and its leadership on this achievement, as our nation continues to enhance its global standing across all fields.
The MBZUAI IFM and the LLM360 team's first day at @iclr_conf, come to visit our new Institute of Foundation Models! Booth D04 in Hall 2!
We’re looking forward to meeting researchers and engineers to introduce them to @mbzuai .
Does cleaner data always yield better models in LLM alignment? Our findings challenge this notion, revealing that even high-quality, difficult examples, can hinder alignment, reducing performance by approximately 10 win rate points.
More insights and findings:
Paper: https://t.co/rzVFH5tnbq
Github: https://t.co/sWkC3Vwh8u
ICLR25 + NAACL25 acceptances in one night—what a surprise!
NAACL: NAT: Enhancing Agent Tuning with Negative Samples [https://t.co/24fCaLfEoz]
ICLR: ToolGen: Unifying Tool Retrieval and Calling via Generation [https://t.co/eJnVQt5c9R]
Both works are agent-focused and lead by @realReasonWang, showcases his exceptional talent and dedication. Huge congrats!
We are excited to share our tutorial at @coling2025 on Safety Issues for Generative AI.
We’ll explore:
🔍 LLM Jailbreaking and red-teaming
🤖 Multi-modal & agentic AI Safety
🛡️ AI automating and Scheming
...
📅 20 Jan, 9:00-12:30
📍 Capital Suite 7, ADNEC, Abu Dhabi
🌐 Website: https://t.co/5Db5o3EujJ
This Tutorial is organized in collaboration with @han_xudong@Emad_A_Alghamdi@ShomLinEd@monojitchou@zjf_heart@eltimster
Looking forward to seeing you there!
#COLING2025 #AISafety #GenerativeAI #AI
5/ 🙏Big thanks to @mbzuai and all the researchers and institutions for their invaluable support of Libra-Leaderboard. Your collaboration and contributions are key to making this project a reality. The names of our supporters are featured in the image below. We look forward to continuing this journey together in advancing AI safety.🤖💡
4/ Safety Arena is an innovative platform designed to make AI safety accessible to non-experts, helping raise public awareness of AI risks. Users can interact with LLMs through a chat-based interface and apply built-in adversarial attack methods to their input prompts. The platform is closely integrated with Libra-Leaderboard, where user feedback directly impacts the evaluation scores of LLMs.
Try it out: https://t.co/xzv6U1MKIC
ToolGen: Unified Tool Retrieval and Calling via Generation
Integrate tool knowledge directly into LLMs by representing tools as a unique token.
This allows the LLM to generate tool calls and arguments, enabling seamless tool invocation and language generation.
"Experimental results with over 47,000 tools show that ToolGen not only achieves superior results in both tool retrieval and autonomous task completion but also sets the stage for a new era of AI agents that can adapt to tools across diverse domains."
I like this approach because it prevents the need to call an external system and can probably be better tuned for more challenging tool-calling scenarios. It might also be an interesting method to combine with CoT and these large reasoning models.
Our latest work on Tool Learning and Agent
🚀 ToolGen: Unified Tool Retrieval and Calling via Generation
Integrating tool knowledge directly into LLMs by representing tools as unique tokens, enabling seamless tool invocation and language generation. 🔧🤖
💡 ToolGen achieves superior performance in both tool retrieval and autonomous task completion.
We believe it will set the stage for a new era of AI agents ! #AI #LLM #ToolRetrieval #AIResearch #NLP #ML
📝Please fill in your information to get a free pass before they’re gone-only 3 days left to register!
⬇️Check the comments for the link to our questionnaire.
Let’s meet and talk about innovation, AI, and opportunities!
#LibrAI#AI#GITEX#FreePass#GITEX2024#ExpandNorthStar
Excited to share our 6 papers at #ACL2024NLP
We focus on benchmarking and evaluation. Welcome to catch up with us.
12 Aug @ 17:15 Poster Session 2
1️⃣ CMMLU: Measuring massive multitask language understanding in Chinese
13 Aug @ 12:15 Poster Session 3
2️⃣ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
3️⃣Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
4️⃣A Chinese Dataset for Evaluating the Safeguards in Large Language Models
13 Aug @ 16:00 Poster Session 5
5️⃣Demystifying Instruction Mixing for Fine-tuning Large Language Models
6️⃣EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models