@OpenAI and @Cerebras have signed a multi-year agreement to deploy 750 megawatts of Cerebras wafer-scale systems to serve OpenAI customers.
This has been a decade in the making.
Deployment begins in early 2026, and when fully rolled out, it will be the largest high-speed AI inference deployment in the world.
OpenAI and Cerebras were both founded in 2015 with radically ambitious goals.
OpenAI set out to build the software that would push AI toward general intelligence.
Cerebras set out to rethink computing hardware from first principles.
Our teams met as far back as 2017. We shared ideas, early work, and a common belief:
there would come a point when model scale and hardware architecture would have to converge.
That point has arrived.
ChatGPT set the direction for the entire industry. It showed the world what AI could be.
Now we’re in the next phase - not proving capability, but delivering it at global scale.
The history of technology is clear on one thing:
speed drives adoption.
The PC industry didn’t operate at kilohertz.
The internet didn’t change the world on dial-up.
AI is no different.
As models grow more capable, speed becomes the bottleneck.
Slow systems limit what users can do, how often they engage, and whether AI becomes infrastructure or remains a novelty.
Cerebras was built for this moment.
By keeping computation and memory on a single wafer-scale processor, we eliminate the data-movement penalties that dominate GPU systems. The result is up to 15× faster inference, without sacrificing model size or accuracy.
That speed changes product design, user behavior, and ultimately productivity.
For consumers, it means AI that feels instantaneous.
For the economy, it means agents that can finally drive serious productivity growth.
For Cerebras, 2026 will be a defining year.
With this collaboration with OpenAI, Cerebras’ wafer-scale technology will reach hundreds of millions - and eventually billions - of users.
We’re proud to work alongside OpenAI to bring fast, frontier AI to people around the world.
This is what a decade of long-term thinking looks like.
A giant data center uses less water than a small restaurant.
In the entire US, data centers use less water than the almond growers in California.
By 5-7x.
These facts are not well known. They should be.
I chatted with Tom Giles on @BloombergLive about how we deliver AI inference at scale for years to come, and the biggest debates facing our industry: power, water, permitting, and the friction slowing U.S. data center deployment.
AI demand is growing quickly. We need an equally serious conversation about how we build for it.
Worth a listen if you’re interested in where AI infrastructure is headed: https://t.co/R7QZOEpcSM
GPT 5.6 Sol is 10 miles ahead of every Anthropic model when it comes to computer use.
They seemed to have resolved every problem they had with 5.4 and 5.5.
It used to compact every 7 actions, would get lost, make mistakes, and take 10-20x longer than any human.
I had some automations that would take about 90 minutes, and I could do that same task in ~5 minutes. Now it takes Codex under 10 minutes.
Not quite human speed, but pretty close.
That same workflow would have also compacted at least 10x before. It did not compact one time.
I'm telling you right now, when @cerebras inference is launched, you will understand why fast inference is the future of automation.
When an agent can not only control your computer, but do it as fast (or faster) than you can, it changes everything.
Buy your AI like you're shopping at Costco.
Not Safeway.
When @Costco first opened, people didn't know how to approach it.
So they shopped it like a Safeway - they walked down every aisle.
That's a horrible way to shop Costco:
It takes hours. And you end up with 4 things you didn't need.
Each was $22. Nearly $100 of mistakes.
But that changed quickly.
Those in the know, went straight to the back.
You get the $4.99 chicken.
18 cupcakes for your kid’s birthday party.
Bang. You get out of there.
(maybe a hotdog on the way out)
You get the value, and you get out.
Same was true for the cloud.
At first enterprises said, this is great, we can bypass our own IT orgs.
And told engineers to have at it.
Put in their credit card.
No surprise there.
Lots of waste.
Now we manage it.
Its not less productive.
Its just a more thoughtful use of resources.
Now the same movie is playing out with AI.
Engineers were told to use as many tokens as they wanted.
Just like with the cloud, and with Costco.
They are on occasion buying massive tubs of unnecessary Mayonnaise.
But like the cloud and Costco, the value unlocks when you get strategic.
Every company needs to find its $4.99 chicken.
The purchase where a few dollars unlocks unbelievable value.
A combination of closed and open source models.
The smartest, most expensive models where you need them the most.
Open source, where top tier smarts is overkill.
Some teams are enormously productive with unlimited tokens, and should get everything they need.
Others just need a few cheaper tokens to be productive.
With unconstrained resource use you get waste.
With thoughtful resource use productivity explodes.
Tokens are no different.
everyone loves writing about whether ai has made markets too expensive
but no one gives airtime to the reality that the global growth outlook without ai would be dramatically worse
ai is the only major new growth engine in sight for years now
and we’re so early still.