@jukan05 Bullish on the field of chip companies if there are any stipulations OpenAI has for how and where Cerebras can sell their chips being the partnership is so co-reliant.
As inference becomes the dominant workload (inference in everything from agents, robotics, world models and beyond) there will be a whole redistribution of margin dollars from the model layer to infrastructure. Labs need innovative models at the highest token efficiency to optimize their margin with higher efficiency, lower TCO which hopefully can lead to falling costs because demand is growing faster than then the price of compute can decrease. Tokens need to flow like proverbial water.
Compute is the new QB market.
Anthropic signed a SpaceX : $1.25B/month for 200k+ GPUs (~$8.50/GPU-hr).
A month later, Googleโs follow-on: $920M/month for ~110k GPUs (~$11.50/GPU-hr).
~35% higher. Once a max-level deal hits, the next QB starts from that number. The market just needs the next available QB.
@GavinSBaker Really interesting point mentioned around how if someone were to break their LTA and end up sitting on excess compute not provided to end customers then theyโll never be able to go back and ask for more supply provision
AI inference is a cost of goods sold.
Token generation is a manufacturing process.
Training converts only 30โ40% of a GPUโs theoretical compute into useful work. Production inference runs lower.
Traditional serving rack: 5โ10kW
NVIDIA GB200 rack: ~120kW-132kW
Every underused watt comes out of gross margin.
On the importance of CPU development: ARM CEO Rene Haas noted on their earnings call that a data center today runs ~30M CPU cores per gigawatt. He put the requirement at ~120M cores per GW inside the same power envelope.
As agents route and orchestrate tasks, CPU cores only gets more load-bearing. GPUs are throughput engines aiding in big, dense, predictable matrix work. Agentic workloads are the opposite: thousands of short, branching, I/O-bound steps, tool calls, retrieval, API waits, serialization, scheduling. Thatโs all CPU work, and itโs what keeps the accelerator fed. Starve it GPU/XPUs sit idle.
4x the cores. Same watts. More maestros need for the orchestra.
Interesting confirmation note on the China vs US public markets per Tony Kim on the Molly Shea Sourcery Podcast due to a lack of depth in the private capital markets compared to the United States, Chinese companies are utilizing public markets as their primary funding mechanism. Consequently, Kim projects a pipeline of 30 to 40 potential IPOs in China with the rest of 2026. This stands in stark contrast to the United States, where the volume of robotics IPOs is currently much lower.
Strong validation from Microsoft that inference cost-efficiency, not frontier capability, is the economic battleground. Amy Hood, Microsoft CFO, directly noted their AI infra as a "fungible fleet" as they and others become infrastructure agnostic so they can swap parts based on what becomes available.
(a) Azure (consistent with all cloud platforms) remains supply-constrained with demand exceeding available capacity for multiple consecutive quarters;
(b) Efficiency gains across the CPU/XPU fleet are being monetized almost immediately because of imbalances identified;
(c) Strategic pivot toward model substitutability and workload-tiered inference, where cheap first-party models handle the majority of tasks and frontier models handle the remainder;
(d) First-party silicon (Maia 200, Cobalt 200) framed openly as a margin and performance-per-dollar growth lever alongside NVIDIA and AMD hardware.
@chamath How do you think about the field of currently established AI hardware startups that are deploying to the different enterprise sectors outside the main hyperscalers/clouds? Can Cerebras really supply to everywhere and anywhere?