50+ pages of prose and formulas, zero experiment, but code in the github.
Is this a slop-grenade, an unvalidated theoretical proposal, or a rushed flag-plant to stand out in 50k+ iclr submissions?
Banning open-source AI would be a historic mistake — and a self-inflicted wound to U.S. AI leadership.
I've been in this field long enough to remember when GPT-2 was called "too dangerous to release." That didn't age well — and it drew heavy criticism from researchers even at the time. 1/5
NVIDIA has released Nemotron 3 Super, a 120B (12B active) open weights reasoning model that scores 36 on the Artificial Analysis Intelligence Index with a hybrid Mamba-Transformer MoE architecture
We were given access to this model ahead of launch and evaluated it across intelligence, openness, and inference efficiency.
Key takeaways
➤ Combines high openness with strong intelligence: Nemotron 3 Super performs strongly for its size and is substantially more intelligent than any other model with comparable openness
➤ Nemotron 3 Super scored 36 on the Artificial Analysis Intelligence Index, +17 points ahead of the previous Super release and +12 points from Nemotron 3 Nano. Compared to models in a similar size category, this places it ahead of gpt-oss-120b (33), but behind the recently-released Qwen3.5 122B A10B (42).
➤ Focused on efficient intelligence: we found Nemotron 3 Super to have higher intelligence than gpt-oss-120b while enabling ~10% higher throughput per GPU in a simple but realistic load test
➤ Supported today for fast serverless inference: providers including @DeepInfra and @LightningAI are serving this model at launch with speeds of up to 484 tokens per second
Model details
📝 Nemotron 3 Super has 120.6B total and 12.7B active parameters, along with a 1 million token context window and hybrid reasoning support. It is published with open weights and a permissive license, alongside open training data and methodology disclosure
📐 The model has several design features enabling efficient inference, including using hybrid Mamba-Transformer and LatentMoE architectures, multi-token prediction, and NVFP4 quantized weights
🎯 NVIDIA pre-trained Nemotron 3 Super in (mostly) NVFP4 precision, but moved to BF16 for post-training. Our evaluation scores use the BF16 weights
🧠 We benchmarked Nemotron 3 Super in its highest-effort reasoning mode ("regular"), the most capable of the model's three inference modes (reasoning-off, low-effort, and regular)
Announcing NVIDIA Nemotron 3 Super!
💚120B-12A Hybrid SSM Latent MoE, designed for Blackwell
💚36 on AAIndex v4
💚up to 2.2X faster than GPT-OSS-120B in FP4
💚Open data, open recipe, open weights
Models, Tech report, etc. here:
https://t.co/CAYpP1iK3i
And yes, Ultra is coming!
30M downloads and counting for the NVIDIA Nemotron family on @huggingface 🤗
We're grateful for the incredible community that has made this possible.
Get started with Nemotron: https://t.co/4AtDcnOCXS
🚀 Introducing Nemotron-3 Nano 30B-A3B:
• Fully open (weights + training data + recipes)
• 1M context
• Fast inference with a hybrid MoE architecture
• Top-tier reasoning & agentic performance
Hard not to like this one.
Stay tuned for Super & Ultra models, coming in 2026!
🚀 Nemotron 3 Nano 30B-A3B is here! Open weights + open data + open source.
AA Intelligence Index: 52 (@ArtificialAnlys )
✅ 1M‑token context
✅ up to 3.3× higher throughput vs similarly sized open models
✅ stronger reasoning/agentic + chat
Details + links in the thread 🧵
The first step on the roadmap to pluralistic alignment is an evaluation framework to quantify the ability of LLMs to adhere to custom behavioral policies in a multi-turn setting. Come check out our eval framework at the Multi-Turn Interactions Workshop on 12/6 @ #NeurIPS2025
1/Excited to share the first in a series of my research updates on LLM pretraining🚀.
Our new work shows *distilled pretraining*—increasingly used to train deployable models—has trade-offs:
✅ Boosts test-time scaling
⚠️ Weakens in-context learning
✨ Needs tailored data curation
Nemotron-H: A family of Hybrid Mamba-Transformer LLMs.
* Hybrid architecture means up to 3X faster at the same accuracy
* Trained in FP8
* Great for VLMs
* Weights and instruct versions to come soon.
https://t.co/h3dLuDuiUz
(2/2)
It provides important improvements over our previous work in Aegis 1.0 (https://t.co/JK1HDng0fd). The dataset and model weights will be open-sourced soon.
If you’re curious about this topic, stop by and say hi! Looking forward to connecting with everyone!
#NeurIPS
Hi folks!
I’m presenting “Aegis 2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails” at the Safe Generative AI workshop (https://t.co/30dgBm9ybV) at NeurIPS 2024 in Vancouver!
(1/2)
#NeurIPS2024#SafeGenerativeAI
SQL injection-like attack on LLMs with special tokens
The decision by LLM tokenizers to parse special tokens in the input string (<s>, <|endoftext|>, etc.), while convenient looking, leads to footguns at best and LLM security vulnerabilities at worst, equivalent to SQL injection attacks.
!!! User input strings are untrusted data !!!
In SQL injection you can pwn bad code with e.g. the DROP TABLE attack. In LLMs we'll get the same issue, where bad code (very easy to mess up with current Tokenizer APIs and their defaults) will parse input string's special token descriptors as actual special tokens, mess up the input representations and drive the LLM out of distribution of chat templates.
Example with the current huggingface Llama 3 tokenizer defaults:
Two unintuitive things are happening at the same time:
1. The <|begin_of_text|> token (128000) was added to the front of the sequence.
2. The <|end_of_text|> token (128001) was parsed out of our string and the special token was inserted. Our text (which could have come from a user) is now possibly messing with the token protocol and taking the LLM out of distribution with undefined outcomes.
I recommend always tokenizing with two additional flags, disabling (1) with add_special_tokens=False and (2) with split_special_tokens=True, and adding the special tokens yourself in code. Both of these options are I think a bit confusingly named. For the chat model, I think you can also use the Chat Templates apply_chat_template.
With this we get something that looks more correct, and we see that <|end_of_text|> is now treated as any other string sequence, and is broken up by the underlying BPE tokenizer as any other string would be:
TLDR imo calls to encode/decode should never handle special tokens by parsing strings, I would deprecate this functionality entirely and forever. These should only be added explicitly and programmatically by separate code paths. In tiktoken, e.g. always use encode_ordinary. In huggingface, be safer with the flags above. At the very least, be aware of the issue and always visualize your tokens and test your code. I feel like this stuff is so subtle and poorly documented that I'd expect somewhere around 50% of the code out there to have bugs related to this issue right now.
Even ChatGPT does something weird here. At best it just deletes the tokens, at worst this is confusing the LLM in an undefined way, I don't really know happens under the hood, but ChatGPT can't repeat the string "<|endoftext|>" back to me:
Be careful out there.
@chris_j_paxton Aside from the ceiling/saturation comments already on here, I think this might have to do with the fact that the 70B is distilled from the 405B (considering that might be the primary difference between Llama 3 70B and Llama 3.1 70B)
New dataset https://t.co/YuUio9IlLL and paper https://t.co/YEntSqNwrc on AI safety that proposes an ensemble approach to content safety #ai_risks#ai_safety
We explore NeMo Guardrails, an open-source toolkit developed by @nvidia for easily adding programmable guardrails to LLM-based conversational systems. We dive into the implementation details on how to add NeMo Guardrails to an RAG pipeline built with RecursiveRetrieverSmallToBigPack, an advanced retrieval pack from @llama_index. We experiment with the following rails:
✅ Input rails
✅ Dialog rails
✅ Execution rails
✅ Output rails
We also compare NeMo Guardrails with Llama Guard. The conclusion? NeMo Guardrails is a thoughtfully and artfully crafted LLM security toolset (and framework). It is a much more comprehensive LLM security toolset, offering a broader set of programmable guardrails to control and guide LLM inputs and outputs, including content moderation, topic guidance, hallucination prevention, and response shaping. Definitely consider adding NeMo Guardrails to your next RAG pipeline!
https://t.co/r2dqmYalt6
I touched on the idea of sleeper agent LLMs at the end of my recent video, as a likely major security challenge for LLMs (perhaps more devious than prompt injection).
The concern I described is that an attacker might be able to craft special kind of text (e.g. with a trigger phrase), put it up somewhere on the internet, so that when it later gets pick up and trained on, it poisons the base model in specific, narrow settings (e.g. when it sees that trigger phrase) to carry out actions in some controllable manner (e.g. jailbreak, or data exfiltration). Perhaps the attack might not even look like readable text - it could be obfuscated in weird UTF-8 characters, byte64 encodings, or carefully perturbed images, making it very hard to detect by simply inspecting data. One could imagine computer security equivalents of zero-day vulnerability markets, selling these trigger phrases.
To my knowledge the above attack hasn't been convincingly demonstrated yet. This paper studies a similar (slightly weaker?) setting, showing that given some (potentially poisoned) model, you can't "make it safe" just by applying the current/standard safety finetuning. The model doesn't learn to become safe across the board and can continue to misbehave in narrow ways that potentially only the attacker knows how to exploit. Here, the attack hides in the model weights instead of hiding in some data, so the more direct attack here looks like someone releasing a (secretly poisoned) open weights model, which others pick up, finetune and deploy, only to become secretly vulnerable.
Well-worth studying directions in LLM security and expecting a lot more to follow.