Training transducer ASR models is bottlenecked by many factors (speed, GPU memory, datasets etc), but another reason is that till recently there didn't exist a public RNNT loss function for PyTorch, which was easy to use.
Here comes https://t.co/K5sFRcHrdH
Introducing ToolGrad: an efficient framework for generating tool-use datasets. By generating ground-truth tool-use chains before prompts, it achieves a nearly 100% pass rate and improved LLM tool-use performance. More: https://t.co/YFwzfX9Rfc
Really excited to announce that at IOI 2026, our model surpassed gold and for the first time on an IOI problem set, the highest human score!
Special thanks to the IOI Technical Committee for helping us run our benchmark and making our team feel welcome. More details below! 👇
25 years ago, I earned a national silver medal but fell short of competing at IOI.
This year, I returned with our NVIDIA team and Nemotron, we achieved gold and outscored the top contestant!🥇
Congrats to all IOI 2026 contestants 👏, and thanks to the IOI Technical Committee!
Congrats to our researchers for exceeding the gold medal threshold on the International Olympiad in Informatics (IOI) 2026 problem set 🥇
Our fine-tuned Nemotron model scored 535.4 out of 600, as graded by the IOI team — higher than the top-scoring human participant.
The team competed unofficially in Uzbekistan, where the International Technical Committee supervised the human contestants. The model had no internet access and faced the same time limits and submission constraints, using the same platform in parallel with the official competition.
Read more in the technical report: https://t.co/4paFvrIV2B
Exciting day for NVIDIA and @huggingface.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI.
Thank you @ClementDelangue for coming to me.
NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗
https://t.co/q8Om2Xc5ye
Happy to share that our paper SHERLOC was accepted to EMNLP 2026 Main!
Code repair agents spend about half their budget finding the bug. But a file path is not a diagnosis.
SHERLOC gives them the where, the why and the fix idea before they patch. 🧵
We discussed a very cool paper in the HF journal club on Direct On-Policy Distillation 🎯
https://t.co/9XNsE8bIms
The motivation is pretty simple: RLVR gives large gains on reasoning tasks, but doing fresh on-policy RL on every large model is expensive because rollout generation dominates the cost.
The paper asks whether we can instead do the expensive RL on a much smaller model and transfer what it learned to a larger one.
The naive approach would be to just distill the post-RL model, but this has an obvious problem that you're not only distilling the improvements from RL, you are also distilling all the limitations of the smaller model.
The clever part of Direct-OPD is to distill the policy shift instead:
→ Take the small model before and after RL.
→ Compute the log-probability ratio between the two policies. This tells you which tokens/actions RL made more or less likely.
→ Treat this ratio as a dense implicit reward and use it to train the larger model on its own on-policy rollouts.
With Direct-OPD, the authors end up with some pretty impressive efficiency results:
→ Direct RL on a 7B model costs roughly 320 hours in their comparison.
→ RL on a ~1.5B model takes ~160 hours, followed by only ~4 hours of Direct-OPD.
→ So you get roughly the same post-training pipeline for ~164 hours instead of ~320 hours.
They also show a nice result on AIME 2024, where Qwen3-1.7B improves from 48.3% → 58.3% with around four hours of Direct-OPD on 8×A100s, and in some settings it outperforms step-matched direct RL.
I would have liked to see the paper show how far the weak-to-strong transfer works at larger model scales: can a 1.5B model improve a much larger one like 30B?
In any case, I think the framing is very neat and hope you enjoy our dicussion of it!
AI security advances when the industry builds in the open, together.
We're introducing the Open Secure AI Alliance with industry leaders to develop new techniques and tools to safeguard software and agents.
By sharing models, tooling and research in the open, we can broaden the community of defenders.
Learn more about the founding members’ contributions: https://t.co/A16oqxs5Ty
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
We're publishing over 365,000 open and agentic RL Environments for SWE, terminal, and search agents
The open research ecosystem has produced many great datasets for the three main agentic domains - software engineering, terminal use, and web research - but every one of them ships with its own harness, its own image conventions, its own grading scripts, and its own failure modes.
We integrated them all.
23 tasksets behind one API, one sandbox lifecycle, one command.
365,000+ tasks in total,
~198,000 software engineering tasks across 20+ languages
~28,600 terminal tasks
~137,600 search tasks
Ready for evals and RL training on Prime Intellect infrastructure, with validated and cleaned dataset re-uploads where the originals needed fixing.
We're happy to announce that our Nemotron-based system scored 30 out of 42 points at the International Mathematical Olympiad (IMO) 2026, held last week. The IMO team officially graded the solutions and certified that the score was equivalent to a Gold Medal for a human participant.
1/5
Mind Lab's 2M-token long-context RL project is now open source.
It took only 8 GPUs compared with the thousands-of-GPU scale for prior 1M-context work.
Ultra-long-context research is no longer a privilege of giant labs.
It’s now open to everyone.
🔗: https://t.co/5ymBcrL0yZ
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
Banning open-source AI would be a historic mistake — and a self-inflicted wound to U.S. AI leadership.
I've been in this field long enough to remember when GPT-2 was called "too dangerous to release." That didn't age well — and it drew heavy criticism from researchers even at the time. 1/5
The Nemotron family just passed 100M downloads!
Huge thank you to the community building with us and showing what’s possible with open models. Cheers to OSS 🍾
Introducing Sakana Fugu: A full multi-agent orchestration system accessible via a single model API.
Our ‘Fugu Ultra’ model matches the performance of Fable and Mythos, delivering frontier capability without the risk of export controls.
Try it: https://t.co/hhO6qTawgb 🐡