🧵 1/3
Our new open-source model, Nemotron-Labs-3-Puzzle-75B-A9B, is out 🎉
We compressed Nemotron-3-Super-120B-A12B into a smaller, faster derivative optimized for interactive deployment.
Announcing NVIDIA Nemotron 3 Super!
💚120B-12A Hybrid SSM Latent MoE, designed for Blackwell
💚36 on AAIndex v4
💚up to 2.2X faster than GPT-OSS-120B in FP4
💚Open data, open recipe, open weights
Models, Tech report, etc. here:
https://t.co/CAYpP1iK3i
And yes, Ultra is coming!
We are excited to release Llama-Nemotron-Ultra! This is a reasoning ON/OFF, dense 253B model. Open weights and post-training data. https://t.co/dCTacylBR8 We started with llama-405B, changed it via NAS pruning then followed by reasoning-focused post-training: SFT + RL in FP8.
Very excited about the release of the Llama Nemotron Super 49B model 🚀 #GTC25
Using distillation-based NAS (Puzzle) we achieved 5X throughput gain!
After SFT and RL, this model tops reasoning benchmarks among open 70B models
"One bad apple can spoil the bunch 🍎", and that's doubly true for language agents!
Our new paper shows how monitoring and intervention can prevent agents from going rogue, boosting performance by up to 20%. We're also releasing a new multi-agent environment 🕵️♂️
How can we interpret LLM features at scale? 🤔
Current pipelines use activating inputs, which is costly and ignores how features causally affect model outputs!
We propose efficient output-centric methods that better predict how steering a feature will affect model outputs.
New preprint led by my student @GurYoav with dream team @Roym4498, Chen Agassy, and Atticus Geiger 🧵1/
What's in an attention head? 🤯
We present an efficient framework – MAPS – for inferring the functionality of attention heads in LLMs ✨directly from their parameters✨
A new preprint with @AmitElhelo 🧵 (1/10)
📢 New Benchmark: SUPER for Setting UP and Executing tasks from Research repositories
Reproducibility is crucial in science. We introduce SUPER to evaluate LLMs' capabilities in autonomously running experiments from research repositories. ⬇️
https://t.co/U47r3F3UO5
Can AI agents solve realistic, time-consuming web tasks such as “Which gyms near me have fitness classes on the weekend, before 7AM?"
We introduce AssistantBench, a benchmark with 214 such tasks.
Our new GPT-4 based agent gets just 25% accuracy!
https://t.co/S3ok3Zmi0T
1/7 🚨 What do LLMs do when they are uncertain? We found that the stronger the LLM, the more it hallucinates and the less it loops! This pattern extends to sampling methods and instruction tuning. 🧵👇
@megamor2@JonathanBerant@OriYoran
🇲🇽 Excited to share our work was accepted to #NAACL2024 main conference!! 🇲🇽
ICL has been hypothesized to perform GD implicitly in its parameters. But is there good evidence for that? 🧐
Depends what you mean exactly!!
Can we leverage pre-existing coding abilities of LLMs to improve semantic parsing and compositional generalization?
🚨 Our new paper shows dramatic improvements when LLMs are prompted with Python rather than DSLs, along with helpful domain descriptions!
https://t.co/2aWfTIrasH
Watch and Share with the world.
A special project by @N12News.
This video contains footage taken by the young partygoers at the Nova Music Festival prior to the 7.10 terror attack.
You’ll only see a handful of the 260 victims and dozens of those abducted or still missing doing what they came there to do; Party, dance, live. Hours later, the festival became a blood-filled scene of unspeakable crimes.
Eyal Waldman is an Israeli billionaire, high-tech magnate (founder of Mellanox)
He built R&D centres in the West Bank & Gaza Strip to employ Palestinian developers in order to build better Israeli-Palestinian relations.
Hamas murdered his daughter Daniel at the music festival
Hello colleagues and fellows. Over the past few days I was shocked to learn that people in our community don't share what I consider to be basic human values. Please help me restore faith in our community by signing this.
https://t.co/zTyQMPizhr