Previewing Ultrafast mode for @OpenAI's GPT 5.6 Sol, powered by Cerebras.
GPT-5.6 Sol Ultrafast generates responses at up to 750 tokens per second. That's the full, GPT-5.6-Sol model -- up to 14× faster than the same model on Standard processing.
It speedran Humanity’s Last Exam in 11h 11m, nearly 7× faster than Claude Fable 5 with comparable accuracy.
Today, @AMD and Cerebras introduced a powerful disaggregated inference solution, pairing the right engine to each phase of the inference pipeline.
This is what agentic AI has been waiting for: the fastest production inference at massive scale.
.@CrowdStrike has selected @cerebras to power its AI security platform, Falcon AIDR.
Speed is now a security feature.
This is a signal of where cybersecurity must go.
AI has accelerated cyber attacks.
Reconnaissance is faster.
Phishing is faster.
Malware variation is faster.
Exploit development is faster.
And AI has created an entirely new attack surface: the AI systems enterprises themselves run. Prompt injection. Jailbreaks. Agent manipulation. Attacks that unfold in seconds.
Defenders need AI that can respond just as quickly.
That is why inference speed matters.
Faster inference gives security AI time to observe live telemetry, call tools, validate evidence, and act before the window to respond closes.
Cerebras was built for this: inference up to 15x faster than leading GPU-based solutions.
Falcon AIDR on Cerebras brings that speed to one of AI’s most important applications: protecting the enterprise in real time.
Proud to partner with CrowdStrike.
.@cerebras' new data center in Toronto.
Racks and racks of wafer scale chips.
Designed and manufactured in the U.S.
Deployed across North America - and beyond.
More coming soon.
.@cerebras data centers are live across North America.
More are coming in Europe.
Here's a quick peek inside one of several new facilities under construction in Toronto, Canada.
Racks and racks and racks of CS-3s.
The fastest AI in the world, delivered across the world.
Today Cerebras announced that @awscloud will be deploying Cerebras CS-3s in their data centers.
Together, Cerebras and AWS will be delivering the fastest inference solution in the world.
It has been an extraordinary 30 days for Cerebras. In February, we announced that we would be powering OpenAI’s fast inference offering.
And within days of signing the contract, we delivered the first tranche of capacity to production.
And today, we announce a multi-year partnership with the industry’s largest cloud provider, AWS.
Together, we will be building fast inference solutions comprised of Trainium 3 doing prefill and Cerebras’ Wafer Scale Engine doing decode.
Cerebras is now powering the world's #1 AI provider, OpenAI. And will be available in the worlds #1 cloud, AWS.
Now back to work.
Just one month after announcing our partnership with @OpenAI, we’re launching our first model together: OpenAI Codex-Spark, powered by @cerebras.
Codex-Spark is built for real-time software development.
In coding, responsiveness is the product.
It is not a nice to have.
Codex-Spark is optimized for targeted code edits, logic revisions, and frontend iteration. It gives developers near-instant feedback so they can stay in flow.
Powered by the Cerebras Wafer-Scale Engine, it runs at over 1,000 tokens/s. That speed fundamentally changes the experience.
We did not build this to win a benchmark.
We built it so developers could move faster.
I’m proud of how quickly the OpenAI and Cerebras teams have brought this to life.
This is what fast execution looks like - deep engineering collaboration, rapid iteration, and shipping real products developers can use today.
We are just getting started.
When inference is fast, entirely new markets open up.
We plan to lead that shift with our partners at OpenAI.
Cerebras Systems today announced the closing of a $1 billion Series H financing at a post-money valuation of approximately $23 billion. The round was led by Tiger Global, with participation from Benchmark, Fidelity Management & Research Company, Atreides Management, Alpha Wave Global, Altimeter, AMD, Coatue, and 1789 Capital, among others.
For more information on Cerebras, visit https://t.co/INfDHvdQwu
@OpenAI and @Cerebras have signed a multi-year agreement to deploy 750 megawatts of Cerebras wafer-scale systems to serve OpenAI customers.
This has been a decade in the making.
Deployment begins in early 2026, and when fully rolled out, it will be the largest high-speed AI inference deployment in the world.
OpenAI and Cerebras were both founded in 2015 with radically ambitious goals.
OpenAI set out to build the software that would push AI toward general intelligence.
Cerebras set out to rethink computing hardware from first principles.
Our teams met as far back as 2017. We shared ideas, early work, and a common belief:
there would come a point when model scale and hardware architecture would have to converge.
That point has arrived.
ChatGPT set the direction for the entire industry. It showed the world what AI could be.
Now we’re in the next phase - not proving capability, but delivering it at global scale.
The history of technology is clear on one thing:
speed drives adoption.
The PC industry didn’t operate at kilohertz.
The internet didn’t change the world on dial-up.
AI is no different.
As models grow more capable, speed becomes the bottleneck.
Slow systems limit what users can do, how often they engage, and whether AI becomes infrastructure or remains a novelty.
Cerebras was built for this moment.
By keeping computation and memory on a single wafer-scale processor, we eliminate the data-movement penalties that dominate GPU systems. The result is up to 15× faster inference, without sacrificing model size or accuracy.
That speed changes product design, user behavior, and ultimately productivity.
For consumers, it means AI that feels instantaneous.
For the economy, it means agents that can finally drive serious productivity growth.
For Cerebras, 2026 will be a defining year.
With this collaboration with OpenAI, Cerebras’ wafer-scale technology will reach hundreds of millions - and eventually billions - of users.
We’re proud to work alongside OpenAI to bring fast, frontier AI to people around the world.
This is what a decade of long-term thinking looks like.
Extremely proud to be a part of this incredible team today and everyday.
“OpenAI is partnering with Cerebras to add 750MW of ultra low-latency AI compute to the OpenAI platform.”
Last month, I closed out a project I started in 2023: publishing a biography of my paternal grandmother, Zvart. The story traces her ancestors' stories from Trabzon (Ottoman Empire), Krasnodar (Russian Empire and USSR), then her life in Iran, France, and eventually Canada. To write it I recorded a few dozen phone conversations with her where she told her life's stories. It was by far one of the most joyful and interesting experiences I've had as I got to know someone I've known my whole life from a completely new perspective. Since finishing, I've been encouraging anyone who will listen to pursue similar projects. Each of our families' stories have as much value to us individually as the major events of history have to the world, and it's a shame to lose any node of our stories. Preserve them while you can :)
GLM-4.7 from @Zai_org is live on Cerebras!
- Frontier intelligence for coding, tool-driven agents, and multi-turn reasoning
- Record coding speed: ~1,000 tokens per second (up to 1,700 TPS for other uses)
- Strong price-performance: ~10x higher than Sonnet 4.5
Today the US Government reached an agreement with UAE's AI National Champion and @cerebras strategic partner, @G42ai , enabling the Emirati firm to deploy Cerebras solutions in the UAE.
We’re grateful for the Administration’s decision, which allows us to scale U.S.-built wafer-scale AI systems into one of the world’s fastest-growing AI hubs.
Working with partners like G42 in the UAE, we can tackle some of the world’s hardest problems with American innovation at the core.
https://t.co/G3lOSUOelv