GPU rental prices this week:
B300 $8.89/hr
B200 $6.42/hr
MI300X $4.99/hr
H200 $4.54/hr
H100 $3.85/hr
Datacenter GPUs cost more than they did in July. B300 +5.7%, H100 +4.6% since July 6. Consumer cards slipped.
Same open weights, many hosts. Cheapest vs most expensive, per M input tokens:
DeepSeek V4 Flash: 6.5x
GLM-5.2: 3.5x
DeepSeek V4 Pro: 2.9x
Qwen3.8 2.4T: 1.25x
Open weights built the supply side. Nothing yet gives it one price.
Every DeepSeek model, every host, one page.
After the Aug 16 increase, 16 of 17 third-party hosts undercut DeepSeek's own API on V4 Flash. On V4 Pro it depends on the hour.
Official rates, host tables, workload math:
https://t.co/Ei5BrM6lne
Oil, power and freight all have a spot price, a forward curve and a hedge.
LLM tokens have a pricing page.
Same supply structure (many sellers, one product). None of the market.
A posted price you can't lock isn't a cost. It's a forecast.
The harness routes each request to the right model and we at @mercatus_ai gives the inference itself a market.
Weโre building the forward market for inference: physically settled token contracts that let providers sell capacity before it exists, and let buyers lock supply before they need it.
DeepSeek, Aug 16: some token prices up 12x, with 10 days' notice.
Same workload, annual cost:
Before: $1,905
After: $3,132 to $6,263
Locked 12mo forward at +5%: ~$2,000
Posted prices move by announcement. Locked prices don't.
https://t.co/quGEeVC6gN
Same week:
Stripe agrees to pay $7B + for OpenRouter (the layer that prices tokens today)
The CFTC opens comment on compute derivatives (the layer that prices them next year)
One exists. One doesn't. That gap is the whole story.
https://t.co/eUNfxMhgFU
US AI inference fell 45.9% in 30 days. It still costs 5x China's.
US TPI: $5.37 per 1M tokens
Coding TPI: $2.63
CN TPI: $1.07
Blended from real usage, not list prices. Rebalanced daily.
https://t.co/FtHxmd92zH
DeepSeek just put time-of-day pricing on AI inference.
Aug 16: V4 Pro moves to peak/off-peak billing.
Input: $0.435 โ $0.66 off-peak, $1.32 peak
Output: $0.87 โ $1.98 / $3.96
Power grids have billed this way for decades.
https://t.co/nh58BGDuGv
Oil until the 1970s. Electricity until the 1990s.
Posted prices, then a spot market, then forward contracts, then a public price curve. Every commodity follows the same path.
AI tokens entered the spot stage in 2026.
US vs Chinese AI inference, per blended 1M tokens:
30 days ago: ~$8.30 vs ~$1.16 (7x gap)
Today: $4.56 vs $1.12 (4x gap)
US prices collapsed 45% in a month. Chinese prices didn't move. Coding inference got MORE expensive (+4.2%).
https://t.co/sbw9MlaEw1
GLM-5.2 trades between $0.07 and $2.10 per million input tokens across 20 providers.
Same weights, 30x apart. The cheapest price fell 95% in 90 days. The official list never moved.
The spread is infrastructure economics, not the product.
GLM-5.2 per 1M input tokens:
Official API: $1.40, unchanged since June
Cheapest host in May: $1.40
Cheapest host today: $0.07
A 95% collapse in 90 days, a 30x spread across 20 hosts. And now the first model with a traded forward price.
https://t.co/mpI0vSWNxs
The same open-weight model sells at seven different prices right now.
DeepSeek V3.2: $0.20 to $0.30+ per million input tokens depending on the provider, and $0.18 to $0.57 across the year. The spread is infrastructure economics, not the product.
The same open-weight model can cost 3x more depending on who serves it.
DeepSeek V3.2 runs $0.18 to $0.57 per million input tokens across providers. Same weights, different economics: utilization, precision, caching, margin strategy.
https://t.co/sqLzEyUwqZ