We are releasing Star Elastic - turn ONE reasoning LLM into MANY sizes with a single post-training run.
360× cheaper than pretraining a family of models.
7× better than SOTA compression.
Split reasoning capability.
Plus elastic budget control that beats the accuracy-latency frontier.
Paper: https://t.co/kOdTSZ1jHb
HF models: https://t.co/1SWU6O7xsE
Thread 👇
RLVR is powerful — but how do you train with multiple rewards effectively? 🤔
🎯GDPO (not GRPO) is coming.
We introduce Group reward-Decoupled Normalization Policy Optimization (GDPO), a new multi-reward RL algorithm that consistently improves per-reward convergence over GRPO across a wide range of tasks. (1/n)
🚀 Excited to share ToolOrchestra, an end-to-end RL training framework for orchestrating tools and agentic workflows.
Everyone’s building agent workflows these days — connecting tools, APIs, and LLMs like LEGO. 🧩
But here are our findings:
👉 Just prompting the agent workflow won’t cut it. It’s not how you build the best agent.
👉 Without learning, workflows plateau fast.
It’s time to bring RL fine-tuning 🔥back into agent development.
(1/n)
👀Your small LMs (SLMs) are… not that fast?
🚀At NVIDIA Research, we release 𝐍𝐞𝐦𝐨𝐭𝐫𝐨𝐧-𝐅𝐥𝐚𝐬𝐡 (NeurIPS 2025), a hybrid SLM family designed around real-world latency and trained from scratch with 1B/3B sizes, achieving SOTA accuracy, latency, and throughput.
🌟𝐍𝐞𝐦𝐨𝐭𝐫𝐨𝐧-𝐅𝐥𝐚𝐬𝐡 𝐡𝐚𝐬 𝐛𝐞𝐞𝐧 𝐢𝐧𝐭𝐞𝐠𝐫𝐚𝐭𝐞𝐝 𝐢𝐧𝐭𝐨 𝐓𝐑𝐓𝐋𝐋𝐌 𝐟𝐨𝐫 𝐩𝐫𝐨𝐝𝐮𝐜𝐭𝐢𝐨𝐧-𝐠𝐫𝐚𝐝𝐞 𝐢𝐧𝐟𝐞𝐫𝐞𝐧𝐜𝐞 with up to 41K tokens/second on a single H100 GPU! Try it following the instructions in our HF repo.
Will share more details at NeurIPS’25 (poster on Thursday, 11am–2pm)!
𝐏𝐚𝐩𝐞𝐫 𝐋𝐢𝐧𝐤: https://t.co/MTt6fwxiFC
🤗 𝐇𝐅 𝐦𝐨𝐝𝐞𝐥𝐬:
Nemotron-Flash-1B: https://t.co/t5CRT76W1b
Nemotron-Flash-3B: https://t.co/SBgsjWR7QB
Nemotron-Flash-3B-Instruct: https://t.co/N4tUPCEyaI
NVIDIA just dropped Universal Deep Research!
You can use it to design your own AI research assistant.
Universal Deep Research (UDR) is the first truly customizable research agent that breaks free from hard-coded limitations.
Unlike existing research tools that force you into predetermined workflows, UDR lets users create, edit, and refine completely custom research strategies without any training required.
The system comes with example strategies (minimal, expansive, intensive) but the real power is in the customization.
🚀 GenMol is now open‑sourced: you can now train and finetune on your data!
It uses masked diffusion + a fragment library to craft valid SAFE molecules, from de novo design to lead optimization.
#GenMol#DrugDiscovery#Biopharma
Small Language Models (SLMs) are the future of Agentic AI, claim @NVIDIA researchers.
Moreover, they offer a method for converting existing agent systems from using LLMs to SLMs that could work in practice.
Here are the details:
Can deep neural networks learn algorithmic abstractions? Our Finitary Abstraction Comprehension Toolkit (FACT), accepted to NeurIPS 2022, presents a dataset and benchmark to address precisely this. @p_belcak https://t.co/ZA2TwS1b9U