We are at #ICML2026 in Seoul, presenting SPEED-Bench, the new unified benchmark for Speculative Decoding from our group @nvidia
Come say hey and chat @ COEX Hall A Poster #809, 10:30am to 12:15pm
How much of SQLite, FFmpeg, PHP compiler can LMs code from scratch? Given just an executable and no starter code or internet access.
Introducing ProgramBench: 200 rigorous, whole-repo generation tasks where models design, build, and ship a working program end to end. 🧵
ארבע שנים מלאות בניסיונות כושלים של ביבי לערער את המשטרים באזור, כולל את הדמוקרטיה הישראלית.
אם בסוף המשטר שלו ייפול, זה באמת יהיה ״מזרח תיכון חדש״, כמובטח.
Introducing SPEED-Bench: A new standard for evaluating Speculative Decoding.
If you're still using random tokens for throughput benchmarking or tiny prompt sets for draft model quality, it’s time for a major upgrade. 👇
🚀Excited to present our new paper that has been accepted to #WACV2026!
Text-to-image models often fail at simple spatial tasks, like placing a dog to the right of a teddy bear.
Our solution: Learn-to-Steer.
We learn a loss function directly from attention maps and apply it during inference.
This work was done together with @AtzmonYuval and @GalChechik
📰arXiv: https://t.co/flonjHFovr
🌐Project page: https://t.co/ReSmHuUzl3
📽️Video: https://t.co/MYscutOv8L
🧵
New paper: We compressed OpenAI’s gpt-oss-120B into a smaller, faster derivative (gpt-oss-puzzle-88B) with no accuracy loss:
⚡ Up to 1.63× higher token throughput on 8×H100
⚡ Up to 2.82× on a single H100
Launching mini-SWE-agent 2.0, the simplest coding agent. Near SoTA performance, with the agent/model/environment only ~100 lines each. Powering benchmarks and RL training at NVIDIA, Anyscale, Stanford and many more!
@ori_press Nice! Very similar to the soliloquy phenomenon that we observed in Claude 3.5 Sonnet in EnIGMA. I wonder what happens during training that makes these models hallucinate full system responses like that
🚀 Excited to share our new paper: “Fast Autoregressive Video Diffusion & World Models with Temporal Cache Compression & Sparse Attention.”
We address attention bottlenecks in auto-regressive video diffusion, enabling ×5–×10 speedup and constant memory over long rollouts.
New eval! Code duels for LMs ⚔️
Current evals test LMs on *tasks*: "fix this bug," "write a test"
But we code to achieve *goals*: maximize revenue, cut costs, win users
Meet CodeClash: LMs compete via their codebases across multi-round tournaments to achieve high-level goals
🎉 I am excited to present our new paper!
Our paper improves personalization of text-to-image models, by adding one special cleaning step on top of existing personalized models.
With just a single gradient update (~4 seconds on an NVIDIA H100 GPU) and a single image of the target concept, our method improves both text alignment and image alignment. For example, it improves LoRA by (+7% / +14%). This is achieved by adding new loss terms and taking into account the prompt and seed.
This work was done together with @dvir_samuel and @GalChechik.
🌐 Paper page: https://t.co/5KeXClcVd3
📄 arXiv paper: https://t.co/U1TXwt35EJ
More details in the comments below.