First stop in Japan: Build-a-Claw. 🇯🇵
Jensen made a surprise stop in the heart of Tokyo to meet with devs, give away some DGX Sparks, and see what they're building with open models.
The NVIDIA Vera Rubin NVL72 compute tray. 200 AI petaFLOPs. Assembled in one minute.
That's what a single-wide, third-generation NVIDIA MGX rack enables. No cables. No hoses. No fans.
It's 100% liquid cooled at 45°C and brings together Vera Rubin Superchips, ConnectX-9 SuperNICs, and BlueField-4 DPUs to deliver the lowest token cost and best performance per watt.
#NVIDIAVeraRubin
$NVDA Vera Rubin NVL72 compute tray delivers 200 AI petaFLOPs and assembles in just one minute.
The fully liquid-cooled system combines Vera Rubin Superchips, ConnectX-9 SuperNICs and BlueField-4 DPUs to lower token costs and improve performance per watt.
Vera CPUs are custom-built for agentic runtime.
We've been working together with NVIDIA on running our sandbox infrastructure that powers Perplexity Computer on Vera CPUs, with significant improvements. We will be sharing more on this soon!
The Nemotron family just passed 100M downloads!
Huge thank you to the community building with us and showing what’s possible with open models. Cheers to OSS 🍾
AI is shifting from model training to always-on token production, and that shift demands a new business model.
NVIDIA is partnering with AI clouds to deploy large‑scale, multi‑tenant AI factories through revenue-sharing and credit-support. This opens up compute access to the fast‑growing AI ecosystem of startups, model builders, enterprises, research organizations and regional AI players.
Nemotron 3 Ultra reached 35B tokens per day on OpenRouter.
Great to see the community putting open, customizable models to work. https://t.co/lg8RCghO3d
Agentic AI is changing the rules for inference. With DeepSeek V4, NVIDIA Blackwell delivered 20x lower cost per token out of the box, running a 1.6T parameter MoE model with a 1M token context on day one.
But the real story is how:
NVIDIA is the only platform co-designed end-to-end across five rack-scale systems—engineered to operate as a unified AI factory rather than a collection of discrete components.
That’s what enables:
→ Higher throughput for agentic workloads
→ Lower latency across multi-step reasoning loops
→ Sustained improvements in token economics over time
As AI factories scale, cost per token becomes the metric that matters and extreme co-design is the advantage that compounds.
📗 https://t.co/389KyfelQ0
We tried a new thing with NVIDIA to roll out Codex across a whole company and it was awesome to see it work.
Let us know if you'd like to do it at your company!
The standalone LPU’s business case rests on a core trade-off: best-in-class latency, but poor token-serving economics at scale. It can serve tokens extremely fast and win premium, latency-sensitive workloads, but it cannot scale throughput like a GPU cluster. That limitation is exactly why Nvidia wanted the IP: not to replace GPUs, but to complement them. Combined, they offer what neither could alone the throughput of GPU clusters with the latency of dedicated silicon. (4/4)
Link to the Newsletter: https://t.co/0HHjcUoFig
NEWS: NVIDIA announces the NVIDIA Nemotron 3 family of open models, data, and libraries, offering a transparent and efficient foundation for building specialized agentic AI across industries.
Nemotron 3 features a hybrid mixture-of-experts (MoE) architecture and new open Nemotron pretraining and post-training datasets, paired with NeMo Gym, an open-source reinforcement learning library that enables scalable, verifiable agent training.
Read more: https://t.co/ldf247t3Zz