Previewing Ultrafast mode for @OpenAI's GPT 5.6 Sol, powered by Cerebras.
GPT-5.6 Sol Ultrafast generates responses at up to 750 tokens per second. That's the full, GPT-5.6-Sol model -- up to 14× faster than the same model on Standard processing.
It speedran Humanity’s Last Exam in 11h 11m, nearly 7× faster than Claude Fable 5 with comparable accuracy.
Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model for efficient, everyday work, and GPT-5.6 Luna, a fast and affordable model for high-volume work.
https://t.co/OoM83SyISN
🧮 What does 8×7B actually mean? It is NOT 8 experts with 7B active parameters per token.
Turns out it’s actually 13B active parameters.
But wait — where does 13B come from?
If you’ve ever tried to make sense of MoE math, our next post in the MoE 101 guide by @dmsobol (and interactive calculator) breaks it all down.
🟧 Cerebras Inference self-serve is finally here 🟧
– Pay by credit card starting at $10
– Run Qwen3 Coder, GPT OSS & more at 2,000+ TPS
– 20x the speed of GPU-based model providers
Go ahead. Melt our wafers.
https://t.co/rorHSQvWh7
Over the past year, @cerebras has delivered 20x faster inference than Nvidia GPUs, powering GPT-OSS-120, Llama, Qwen, DeepSeek & more at 2k–3k tok/sec.
We’re serving trillions of tokens/month via our cloud, on-prem, & partners like @awsmarketplace, @IBM, @huggingface, @openrouter and @vercel.
Today we announced $1.1B Series G at $8.1B valuation to accelerate our mission: building the world’s fastest AI infra for the world's fastest and best builders.
This might be the most information dense blog I've ever written. Added "show me the math" section into MoE 101 p4 episode. We believe it fully models MoE training perf on both gpu and cerebras wse devices.
https://t.co/uW6H78ZE56
🧵1/n
K2 Think is now available and runs fastest on Cerebras Inference at 2,000 TPS on Cerebras – 20x faster than GPU.
Launched by @mbzuai and @G42ai, K2 Think is a world leading open-source reasoning model that beats GPT OSS 120B and DeepSeek on math and reasoning.
After part 3 of MoE 101 series we got two main questions:
1. why is MoE forward pass slower than dense network? 2. why can't I train 64 experts on a single GPU and hit OOM?
we discuss both problems and solutions in part 4: https://t.co/uW6H78ZE56
1/n 🧵
🎂 Cerebras Inference turns 1! 🚀
Let's break it down:
- 6x faster than when we launched — From @Meta Llama to @Alibaba_Qwen 3 to @OpenAI OSS, models running on Cerebras deliver 𝟯,𝟬𝟬𝟬+ 𝘁𝗼𝗸𝗲𝗻𝘀/𝘀𝗲𝗰
- Our inference powers the best AI natives, global enterprises, and developers including Meta, @IBM , AlphaSense, @Docker, @GSK , Mayo Clinic, Core42, @vercel and more
- Largest model served: ~half a trillion parameters, 7x bigger than launch
- #1 provider of tokens on @huggingface 🤗
- Serving billions of tokens per day on @openrouter
- The leading code gen inference provider, powering AI developers with @Cline and @windsurf
We’re just getting started… but couldn’t have done Year 1 without YOU. Build on, builders. ❤️🔥
https://t.co/jREGhLHuxL
Introducing Cerebras Inference
‣ Llama3.1-70B at 450 tokens/s – 20x faster than GPUs
‣ 60c per M tokens – a fifth the price of hyperscalers
‣ Full 16-bit precision for full model accuracy
‣ Generous rate limits for devs
Try now: https://t.co/39xaLQwNfj
@joshhart Century Country Club or Quaker Ridge Country Club. And if you’re looking for a pickup bball game, stop by 4 or 5 Puritan Rd around the corner. Easily the best runs in town.