@jessiedong_@CoreWeave@nebiusai@LambdaAPI Gimlet is the box that changes the map. Multi-silicon inference is economic dispatch: send each slice of a model to the cheapest chip that meets the SLA, like a grid operator dispatching power plants. Once dispatch sits above the hardware, GPU owners compete on marginal cost.
The public board on 27 Sept, per GPU-hour:
On-demand median: $8.55
Spot median: $4.18
Tracked on-demand range: $4.99 to $17.80
Verda: $8.37 on-demand, $4.18 spot, $6.28 on a 24-month term.
Same silicon. The price depends on when, how and from whom you buy it.
Nebius lifts B300 on-demand to $9.50 per GPU-hour on 1 October.
Our 3-year B300 commit is $4.50.
Same chip, less than half the price. The market has split in two: capacity you rent by the hour, and capacity you commit to for years.
The result: a growing gap between when B300 capacity arrives and when buyers need it.
It's already compressing prices on 5-year+ commitments and clusters not RFS until after Q1 2027.
This may prove temporary. For now, B300 is starting to price off RFS date.
Two drivers:
1.Upgrade path. GB deployments offer a cleaner route to Vera Rubin, and that's pulling more buyers toward GB300.
2.Timing. A significant share of B300 supply comes online in 2027. Many buyers still want capacity live before year-end.
B300 is still the most requested GPU in our deal flow. Yet for the first time, there's more B300 on offer than buyers asking for it.
It's the only architecture in our top six where that's true. On GB300, B200, H200, H100 and Vera Rubin, demand still outruns supply.
The most interesting number in MLPerf v6.1 is not Rubin's 3.7x.
It is 1.6x, the gain NVIDIA got on Qwen3-VL from software alone between v6.0 and v6.1.
Your existing fleet got faster for free. Price that in before you assume you need new silicon.
Anthropic can commit $45bn to Rubin capacity 15 months before it exists. You cannot.
That asymmetry is the whole game. The buyers who get Rubin in 2028 are the ones who reserved in 2026.
There is a third door between waiting and writing a cheque that size.
$80bn of compute booked in five days. CoreWeave sitting on a $104bn backlog and 3.7GW contracted.
The Rubin capacity you plan to buy in 2028 is being reserved this quarter.
Anthropic signed $45bn with Nscale on 26 August and $35bn with Lambda on 31 August.
The Nscale contract runs on Vera Rubin and does not switch on until late 2027.
That is $45bn committed 15 months before first power. Rubin is not being sold. It is being sold forward.
NVIDIA spent Hot Chips presenting six chips. The number that matters is the one they wrapped them in: a 100MW reference factory.
11 PB of HBM4. 2 ZFLOPS of inference.
The unit of purchase is no longer a GPU. It is a building, and it is reserved years out.
The more AI โthinksโ, the more important memory becomes. The intuitive assumption is that reasoning models primarily create demand for more FLOPS. But reasoning generates more decode tokens, and decode is predominantly memory-bandwidth-bound.
You can order the GPU rack. You can't order the power on the same clock. Grid connection now runs 5 to 7 years. Transformers 3 to 5. Switchgear sold out to 2028. To run 600 kW racks in 2027, you are buying grid capacity and long-lead gear in 2026, before the GPUs exist.
Also in here: 800 VDC distribution moves 150% more power through the same copper and strips about 200kg of busbar per rack. Projected 39GW on 800 VDC by 2030.
https://t.co/gGDmmjJJd1
One NVIDIA rack:
2022 Hopper, 35 kW
2025 Blackwell, 125 kW
2026 Vera Rubin, 210 kW
2027 Kyber, 600 kW
2028 Feynman, 1 MW
29x in six years, same footprint.
NVIDIA ships a new platform annually. A greenfield data centre takes 18 to 36 months to build.
SpaceX's S-1 says Anthropic pays about $1.25bn a month for Colossus 1. 300MW, roughly 220,000 GPUs.
Do the division:
$50bn per gigawatt-year
About $7.78 per GPU-hour
Powered shell in Norway leases at $202/kW/month. Memphis implies $4,167.
The silicon is the other 95%.
The H100 hour just hit $3.08, its highest in a month.
A three-year-old chip is supposed to be depreciating fast. Instead it is at a monthly high, and the compute behind it is still in demand.
Depreciation is an assumption. This is the price at https://t.co/TYz89lPhU3
Four straight years of ~10x annual price cuts on inference. Every layer underneath posted record revenue anyway.
One layer moved the other way: power. PJM capacity cleared at $28.92 per MW-day for 2024/25. It's now $333.44.
Compute deflates. Watts don't.