Uno researchers report language models generating tokens up to 3× faster while preserving output quality. Method uses lightweight diffusion adapters to propose multiple tokens in parallel, then validates with the base model. https://t.co/miTVRQTND7
It's currently trending as the #1 paper of the day on HF!
Thanks for sharing our work, @rohanpaul_ai
Folks, pls upvote if you liked our work 🤗
https://t.co/mIQUMkejc0
cc:@_akhaliq
Build on K2 Horizon.
• Self-host all six models with vLLM or SGLang, weights on Hugging Face
• Run locally with Ollama
• Use it in OpenCode and OpenClaw
• Deploy through our API platform
Details in IFM developer doc: https://t.co/LKxz8YhhQI
Download weights: https://t.co/3Lb28JhyG9
Access API Keys: https://t.co/gD3NU3GImI
Diffusion LLMs have two limitations relative to AR models:
(1) Lower quality, and
(2) Slower inference at large batch sizes.
We address this "Uno"
> Retains the AR architecture of LLMs
> Each layer has two sets of weights: AR weights and Diffusion weights
> Diffusion weights enable parallel sampling from the AR distribution losslessly
Results:
💥Faster than all speculative decoding methods: DFlash and EAGLE-3
🔥 Beats ALL diffusion LLMs: Mercury 2, Diffusion Gemma, Llada
Links to the paper, models, code below
Introducing K2 Horizon: a connected fleet of six foundation models ranging from 0.9 billion to 375 billion parameters.
- Frontier performance: Across coding and agentic tasks, K2 Horizon delivers top-tier performance in every size class—with the 0.9B, 3.7B and 7B models setting new state of the art at their respective scales.
- Radical openness: K2 Horizon represents the largest fully open-source model launch in AI history. The fully open code, training data and recipes are a significant step forward in transparency.
Launch page: https://t.co/gg0k803SbL
Tech blog: https://t.co/g35L5xMGdS
Hugging Face: https://t.co/3Lb28JhyG9
We are proud to present a full fleet of LLMs, radically open sourced. This is a fleet of 6 model sizes you can choose, with all model training details, code, data recipe, intermediate checkpoints. If you are interested in deployable agentic models, or curious about how models work, the K2 Horizon fleet is for you.
Really proud of our K2 Horizon release!
One message I hope people take from this work is that truly OSS models matter, and “open” does not have to mean weak.
True openness should go far beyond downloadable weights. It should include the training data, code, recipes, and sufficient details for others to understand how a model was built, reproduce it, adapt it, and improve upon it. This level of transparency is essential if we want AI progress to remain scientific, collaborative, and broadly useful.
K2 Horizon is a connected family of six foundation models spanning 0.9B to 375B parameters. Across coding and agentic tasks, it delivers top-tier performance in every size class, with the 0.9B, 3.7B, and 7B models setting new SotA at their respective scales. To me, this is powerful evidence that OSS models can also be frontier models.
My students and I have focused mainly on driving the small-model cohort: 0.9B, 3.7B (~4B), and 7B with @MaxMa1987's team. Small models are not an afterthought or merely smaller versions of large systems. They are an important research frontier in their own right.
Small models can bring capable intelligence directly to phones, laptops, robots, sensors, and other edge devices. Local deployment can reduce latency and cost, make AI useful with limited connectivity, and, critically, help protect privacy by keeping sensitive user data on the device rather than sending it to the cloud.
A special shoutout to my @RutgersCS student Junlin Chen (@Chen94751623484). As a senior undergraduate, Junlin has already spearheaded many essential parts of the research and engineering behind this small-model cohort. That requires technical depth, persistence, sound judgment, and the ability to own difficult problems end-to-end. I’m tremendously proud of what he has accomplished.
Open can be strong. Small can be powerful. And talented young researchers can lead work at the frontier.
Congratulations to the entire team @IFM_AI. I’m excited to see what the community builds with K2 Horizon.
#K2Horizon #OpenSourceAI #SmallModels #EdgeAI
Introducing K2 Horizon: a connected fleet of six foundation models ranging from 0.9 billion to 375 billion parameters.
- Frontier performance: Across coding and agentic tasks, K2 Horizon delivers top-tier performance in every size class—with the 0.9B, 3.7B and 7B models setting new state of the art at their respective scales.
- Radical openness: K2 Horizon represents the largest fully open-source model launch in AI history. The fully open code, training data and recipes are a significant step forward in transparency.
Launch page: https://t.co/gg0k803SbL
Tech blog: https://t.co/g35L5xMGdS
Hugging Face: https://t.co/3Lb28JhyG9
Small models shouldn’t mean small capabilities.
Proud to have worked across the K2 Horizon family—and to have been one of the main contributors to the 900M, 3.7B, and 7B models.
A true team effort.
Come check them out! 🚀
#OpenSourceAI#LLM#Edge