@ShivAroor The dog whistle has been sounded and all of you have assembled for the witch-hunt. When are you going to pivot to real journalism for a change? What a sellout.
Wow @SemiAnalysis just dropped day-zero numbers on @deepseek_ai V4 Pro
Blackwell (B300): 8,075 tok/s/GPU
AMD MI355X: 6.99 tok/s/GPU
Hopper (H200): 186 tok/s/GPU at higher interactivity
Same interactivity. ~1,000× the throughput per GPU vs MI355X. Blackwell also keeps serving all the way out to 80 tok/s/user, more than 2× the range where AMD flatlines
That number isn’t a typo!
But the throughput isn’t the story
The story is the words “Day Zero”
DeepSeek V4 Pro is a 1.6T-parameter Chinese frontier model. It dropped Friday. Shortly after, the open source community released day-0 deployment recipes, with NVFP4 checkpoints, TensorRT-LLM + WideEP, Dynamo, optimized V4 kernels in the pipeline
That ecosystem doesn’t exist for the alternatives. That’s the real moat
If you’re a neocloud operator, an enterprise AI buyer, or a hyperscaler sizing your next infrastructure build, this is what you’re actually buying when you buy NVIDIA
You’re not buying a chip. You’re buying the guarantee that when the next frontier model drops, Chinese or American, open or closed, reasoning or multimodal, your fleet gets strong day-0 performance with continuous improvements over time
I wrote about this four weeks ago in 27× On The Same Iron. MLPerf v6.0 showed GB300 delivering 2.7× more throughput than its v5.1 debut. Same silicon. Pure software. 60%+ cost-per-token reduction. Every operator already running Blackwell got that gain for free
That was the proof point on installed-base compounding. This is the proof point on day-zero readiness for models that didn’t exist when you bought the iron
Now look at the alternative
You buy MI355X. The next frontier model drops. Now you wait. For AMD’s kernel team. For SGLang or vLLM upstream. For the open-source community to prioritize a stack that ships on a fraction of the world’s deployed compute. Weeks. Sometimes months. Sometimes never at peak
Same story for any TPU buyer who isn’t Google. Same story for anyone betting on chip startups that haven’t earned a place in the open-source optimization queue
Hopper, the last generation, is beating brand-new AMD silicon by a wide margin on a model that didn’t exist a week ago
That’s not a benchmark. That’s a procurement decision
Day zero is the moat
$NVDA $AMD $GOOGL
@LayoffAI You’re clearly double counting people over and over and artificially inflating the number for optics. Hypothetically - If a person switches employers ie transfer every year over 10 years you’re increasing the cumulative count by 10. It’s just the same 1 person 🤷
@VipulDivyanshu 8.3% util isn’t great.
The numbers don’t add up, 8.3% peak util = 0.083*19 peak fp16 ANE TFLOPs = 1.57 delivered TFLOPs. At 1.2W from your screenshot you’re more like at 1.3 TFLOPs/W. The 80x number is very exaggerated. A100 is 312 peak fp16 TFLOPs at 400W.