Ten years ago, we started @cerebras around an approach many believed was impossible.
As a computer architect, it is hard for me to imagine a more exciting time. Model releases are accelerating, and hardware tapeout is compressing from multi-year roadmaps to annual launches.
Hot Chips is my favorite conference, and it’s where I launched Cerebras 7 years ago. This year’s conference was especially exciting, and so much innovation was shared. I am watching the industry recreate itself: SRAM is mainstream, DRAM is moving into the third dimension, networks are being fundamentally redesigned, and AI is helping design and program the chips themselves.
The industry has never moved faster and some of the hardest architectural questions are still wide open.
Just one month after announcing our partnership with @OpenAI, we’re launching our first model together: OpenAI Codex-Spark, powered by @cerebras.
Codex-Spark is built for real-time software development.
In coding, responsiveness is the product.
It is not a nice to have.
Codex-Spark is optimized for targeted code edits, logic revisions, and frontend iteration. It gives developers near-instant feedback so they can stay in flow.
Powered by the Cerebras Wafer-Scale Engine, it runs at over 1,000 tokens/s. That speed fundamentally changes the experience.
We did not build this to win a benchmark.
We built it so developers could move faster.
I’m proud of how quickly the OpenAI and Cerebras teams have brought this to life.
This is what fast execution looks like - deep engineering collaboration, rapid iteration, and shipping real products developers can use today.
We are just getting started.
When inference is fast, entirely new markets open up.
We plan to lead that shift with our partners at OpenAI.
There will be many winners in AI hardware.
The dominance of the graphics processing unit will recede and be replaced by multiple different architectures. This will happen first for Inference.
The graphics processing unit will not go away, it will just be substantially less dominant. Not because they are bad - they’re brilliant engineering. But because their fundamental architecture, off-chip high-capacity but low-speed memory was built, as their name suggests, for graphics.
And as the market puts AI into production the demand for speed will accelerate and drive a predictable shift in hardware.
Big customers want speed, control and cost efficiency.
This is exactly why we started Cerebras nearly 10 years ago. We believed that AI would require a fundamentally different architecture.
🚨 15–16 years & still no homes! 🚨
We, the homebuyers of Whispering Towers (Mulund) & Majestic Towers (Nahur) booked flats in HDIL projects. Possession promised in 2013. Till today, not a single flat delivered.
#JusticeForHomebuyers#HDIL@timesofindia@abpmajhatv (1/6)
In 2016, @sama and I first met. @OpenAI was a vision. @cerebras was powerpoint. Sam and the OpenAi founders became one of the early investors in @cerebras.
In the following years, the Cerebras and OpenAI frequently met to explore working together. But the timing was too early—LLMs hadn’t been invented yet.
Today, the story comes full circle. OpenAI just released its most powerful open weight reasoning model—and it runs fastest on Cerebras Systems. Not a little bit faster than the competition. It smokes the competition.
Running on our third generation Wafer Scale Engine, OpenAI gpt-oss-120B runs at up to 3,000 tokens/s – the fastest speed achieved by an OpenAI model in production. Reasoning that take minutes on @nvidia GPUs take a single second on Cerebras Systems.
Good things take time.
92% of companies plan to increase AI investments over the next 3 years.
But only 1% have fully integrated AI into their workflows.
The gap?
Real-time performance and scale.
That’s why we created Cerebras Supernova.
Join us - Space is limited: https://t.co/syzPat5Wzu
Deploying LLMs with low latency and real-time responsiveness is no small task. But with @datarobot and Cerebras Inference, you’ll be ready to customize and deploy LLMs that deliver speed, precision, and real-time responsiveness.
Dive in: https://t.co/DBwmgOM8E8
Cerebras' collaboration with Mayo Clinic has delivered the state of the art foundation model for genomics. And, the model is showing tremendous promise in predicting which Rheumatoid Arthritis drug most likely to successfully treat the disease. AI is making peoples lives better
We're proud to announce the TKO will be the official keyboard of @DreamHack PC Freeplay for Dallas, Atlanta, and beyond.
⌨️Demo the TKO
🖥Play Games
💸Score Deals
🎁Win Prizes
And to celebrate, we're giving away passes to DreamHack Dallas THIS WEEKEND.
https://t.co/6dzZlG0ekE
Eager to get under the hood of our extraordinary CS-2 system? Take a virtual tour to learn what a marvel of mechanical, thermal, electrical, and semiconductor engineering co-design the CS-2 is! https://t.co/2LQU7iaG1p #AI#deeplearning#AIHardware
We're proud to announce we've raised $250M in #SeriesF funding to accelerate our global business expansion and relentless pursuit of #AI innovation! https://t.co/4iCiEKL6J8 #deeplearning
“Training which historically took over 2 weeks to run on a large cluster of GPUs was accomplished in just over 2 days — 52hrs to be exact — on a single CS-1.” Our collaboration with @AstraZeneca is driving innovation in drug discovery using #AI solutions
https://t.co/OarkThxH4q