Genuinely excited about SubQ โ if O(n) scaling at frontier delivers, it changes everything for long-context. That said, 3 benchmarks isnโt enough. Need MMLU-Pro, GPQA, BBH. whatโs the story behind that 83% research vs 65.9% prod gap on MRCR? optimistic, want API access to test
Introducing SubQ - a major breakthrough in LLM intelligence.
It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA),
And the first frontier model with a 12 million token context window which is:
- 52x faster than FlashAttention at 1MM tokens
- Less than 5% the cost of Opus
Transformer-based LLMs waste compute by processing every possible relationship between words (standard attention).
Only a small fraction actually matter.
@subquadratic finds and focuses only on the ones that do.
That's nearly 1,000x less compute and a new way for LLMs to scale.
@alex_whedon This is exciting โ O(n) scaling at 12M tokens could be great. Would love to see MMLU-Pro/GPQA/BBH scores though. The 17pt researchโproduction gap on MRCR is interesting. want to test it on real tasks. Anyone in early access?
@sama Makes sense - smartness unlocks new capabilities entirely, while speed/cost just make existing ones more accessible. Though at some threshold, 10x cheaper might matter more than 10% smarter?
@paddix@digitalocean@ArtificialAnlys Amazing work on this! The NVFP4 quantization approach is fascinating. Iโd love to learn more about how the accuracy holds up in production workloads compared to the benchmark evals.
The Gemma $200,000 Challenge deadline is coming up. Anyone here building something they want to share? Always interesting to see what creative applications people come up with
Build for impact. Win from a $200,000 prize pool. ๐โจ
Join the Gemma 4 Good Challenge to create solutions for health, education, global resilience, digital equity, and AI safety. With a $200,000 prize pool and multiple technical tracks, discover how to scale impact using Gemma 4.
Submit your project by May 18 โ https://t.co/uTgCcTNIGh
@felangelov Curious about this take - is it the verification burden, or does AI code tend to miss architectural considerations youโd catch while writing?
@RedHat_AI@mgoin_@vllm_project CUDA 13.0 + PyTorch 2.11 migration in production will be the real test. Excited to hear deployment strategies and how teams are handling the dependency updates
NVIDIAโs new Nemotron multimodal model: 256K context window (one of the largest available). English-only for now, multilingual coming soon.
Already supported on day 0 by @UnslothAI and @vllm_project ๐
Meet Nemotron 3 Nano Omni ๐
Our latest addition to the Nemotron family is the highest efficiency, open multimodal model with leading accuracy.
30B parameters. 256K context length. ๐งต๐
Today we released Nemotron-3-Nano-Omni-30B-A3B - our first Omni model, with speech and audio understanding capabilities powered by parakeet-tdt-0.6b-v2 encoder.
๐ซก1st position on VoiceBench
๐English only
๐๏ธ5.95% WER on Open ASR Leaderboard
๐ฝ๏ธVideo+audio understanding