As per Artificial Analysis' updated Evals, our K2-Horizon-400B-Moe model still beats Thinking Machine's Inkling model.
Pretty respectable for a lab that was set up a little over a year ago, with a decent chunk of the K2-series effort being led by interns!!!
Couldn't be prouder to be part of the @IFM_AI family ♥️
This is amazing work! It's the first release of this scale from IFM so I expect rough edges when the models are actually used, but it's open source so if you don't like it you can fix it yourself!
Great resource for studying training dynamics and interpretability as well.
K2 Horizon 375B A23B, a new open weights model from UAE's MBZUAI, scores 47 on the Artificial Analysis Intelligence Index, with relatively strong agentic performance and a 30 point jump over its predecessor
K2 Horizon 375B A23B is an open weights Mixture-of-Experts model with 375B total and 23B active parameters from @IFM_MBZUAI, MBZUAI's Institute of Foundation Models. It scores 47 on the Intelligence Index, alongside models such as MiniMax-M3 (45, also a MoE with 23B active parameters), and a large upgrade from its predecessor K2 Think V2 (17, 70B dense model). It leads nearby open weights models on agentic evals and has a low hallucination rate, but trails on knowledge and the hardest reasoning evals. K2 Think V2 ranks among the most open models on our Openness Index; MBZUAI is updating the supporting documentation and code for K2 Horizon and we expect to add it to the Openness Index soon.
Key takeaways:
➤ Strong on agentic tasks, weaker on knowledge and deep reasoning. MiniMax-M3, a recent model that is close to it on the Intelligence Index, makes the cleanest comparison: K2 Horizon 375B A23B leads on GDPval-AA, our real-world knowledge work benchmark (Elo 1430 vs 1380), and on τ³-Banking (34.2% vs 15.3%), but trails on GPQA Diamond (87.3% vs 92.9%) and Humanity's Last Exam (32.0% vs 39.0%)
➤ Low hallucination rate, driven by abstention rather than knowledge. K2 Horizon 375B A23B attempts only 40% of AA-Omniscience questions, declining the remaining 60% rather than guessing. The result is a 26% hallucination rate, among the lower rates we have measured, while accuracy is 18%, essentially unchanged from K2 Think V2
➤ A new architecture over its predecessor. K2 Horizon 375B A23B is a 375B parameter Mixture-of-Experts model with 23B active, succeeding the 70B dense K2 Think V2, and extends context from 262K to 512K tokens. Its 23B active parameters match MiniMax-M3 (428B total, 23B active)
Key model details:
➤ Architecture: Mixture-of-Experts, 375B total parameters, 23B active
➤ Context window: 512K tokens
➤ Multimodality: Text input and output only
➤ Pricing and availability: Yet to be announced
➤ Licensing: Open weights (license details to be announced)
Introducing K2 Horizon: a connected fleet of six foundation models ranging from 0.9 billion to 375 billion parameters.
- Frontier performance: Across coding and agentic tasks, K2 Horizon delivers top-tier performance in every size class—with the 0.9B, 3.7B and 7B models setting new state of the art at their respective scales.
- Radical openness: K2 Horizon represents the largest fully open-source model launch in AI history. The fully open code, training data and recipes are a significant step forward in transparency.
Launch page: https://t.co/gg0k803SbL
Tech blog: https://t.co/g35L5xMGdS
Hugging Face: https://t.co/3Lb28JhyG9
I've been unsettled lately. Reading messages or papers feels like dissociating. Everything seems a bit alien, even if it's completely human. I've had a realization: When our simulations finally exited the Uncanny Valley, they brought the Uncanny with them. https://t.co/fdrKq6MkbM
Opus 5 is frankly unusable when trying to understand a concept. The amount of jargon filled in its responses is astonishing and I have to create a skill.md just to take care of this
Open research is meant to be built on. IFM created TxT360 to give the AI community an open, high-quality, and flexible foundation for LLM pretraining.
TxT360 is the first dataset to globally deduplicate 99 CommonCrawl snapshots alongside 14 commonly used non-web sources—making large-scale, diverse pretraining data more accessible.
Its impact is already visible across the ecosystem: @AMD used TxT360 in the pretraining corpus for its latest model, Instella-MoE-16B-A3B, and thanked the LLM360 team in the acknowledgments.
Explore TxT360: https://t.co/9p3pC6vo7K
I guess you could say this about using LLMs in general.
There's a fine line between letting AI do everything for you and using AI to accelerate your work while sharpening your thinking, and I don’t think we’ve figured out how to do the latter yet
i used to do this a lot. but then i felt like i was getting progressively dumber at writing.
i think there's some value in forcing your own brain to sharpen your thinking, instead of offloading it to AI
what i found helpful is to manually type out my own understanding in distilled bullet points after reading the LLM output
(re: OpenAI and Anthropic IPO news, a preprint)
Scaling up training reliably improves LLMs, but it also increases training and inference costs, leading to massive capital expenditure by AI firms. How can we understand what level of LLM scaling is justified economically? 🧵⬇️
When you map today’s #LLMs across performance, openness, and scale, the landscape becomes unmistakable: U.S. frontier models dominate performance, but remain completely closed; Chinese open-weight systems occupy a large semi-open middle band, K2-V2 is one of the only models that reaches the high-openness region while still competing on real capability.
The arrival of the K2 model series from @mbzuai Institute of Foundation Models is poised to change the LLM landscape to a better place – these models are delivering comparable performance to the best open-weight models, and they are “360-open”, with data, checkpoints, logs and methodology, fine-tuning recipes all made public alongside the model weights, so developer now have alternatives to these highly popular open-weight LLMs, but stay put on safety and reproducibility without compromising performance.