Most of the reasoning a model does isn't the answer. It's the model talking itself into the answer.
With Ember-1 we went after that overhead: about 40% fewer tokens on Kimi K3, and the quality held up on live customer A/B tests, coding workloads and external benchmarks.
I spent a lot of late nights watching models ramble. Glad to see it paid off. Try it on Fireworks.
Ember-1 is a specialized model from Fireworks Research designed to make every token go further.
Built on Kimi K3, it produces shorter reasoning traces, using roughly 40% fewer tokens while maintaining top-tier quality.
https://t.co/PLD1l5aeh4
We launched an intelligence benchmark! Real work evals are great to help shortlist which models are worth considering for your applications. Kudos to the @FireworksAI_HQ team and all our amazing launch partners!
Today we're launching the Specialized Intelligence Index (SII): one destination for real-work benchmarks across industries, built by the teams that use them every day.
Hear from Fireworks co-founder @the_bunny_chen on the importance of specialized benchmarks:
Many assume that agent spend goes toward output tokens.
When we ran DeepSWE on Astra vs. DeepSeek V4.1-Flash, input tokens outnumbered output 174 to 1. 99.6% were cache hits. Those hits are 60% of the bill.
Net result? Same quality. $0.43/task vs $6.52. https://t.co/dVUPe5EzWP
Reinforcement learning at @cognition's scale is a hard infrastructure problem. We are proud to be part of the stack behind it.
Congrats to the team on SWE-2!
Read more about how we think about RL at Fireworks: https://t.co/b5w5LXtGCl
Congrats to the Cognition team @carlobaronio@silasalberti and others for training another frontier coding model at a fraction of the cost! amazing to partner with this team as always
Introducing SWE-2, our closest model yet to the frontier.
On leading evals, it scores on par with recent frontier models – at up to 70% lower cost.
We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
Congrats to @genspark_ai on being able to match Claude Opus at a fraction of the cost - kudos to the amazing research team at @FireworksAI_HQ that partnered in training this model!
Meet Gen-1 Slides, Genspark’s first model designed for knowledge work.
Trained with @FireworksAI_HQ from an open-weight base, Gen-1 Slides now powers Standard mode in Genspark AI Slides, delivering frontier quality decks at roughly 1/17th of Opus 5’s price.
Fewer midnight fixes. Better decks by default.
Read the full blog: https://t.co/cn5WgCtfW2
We invite you to Forge.
At Forge, we will challenge you to make your own frontier.
You'll spend the day with the builders, researchers, and technical leaders pushing AI forward.
Lin Qiao. Jensen Huang. Jay Parikh. And more.
Apply to attend: https://t.co/9TTDskVI73
The unexpected benefit of open source is that it allows model training companies to train on using your software (and optimizing it) for free.
Blender may have just won as a 3D asset creation software because the models will be better at using it than any proprietary ones
Everyone knows that labs can benchmaxx: optimize the model to look good on the leaderboard.
Sometimes forgotten is that benchmarks can benchmaxx-maxx: optimize the benchmark to make its leaderboard look good.
same model. same list price. 5.7x apart in real cost.
glm-5.2 across 6 hosts:
@FireworksAI_HQ 18% of list
@sferenceai 30%
@tensorx_ai 43%
@nebiustf 100%, caches nothing
the price sheet tells you almost nothing. cache hit rate is the price.
GLM-5.3-Flash is live on Fireworks on day… 2
Why? Because we take quality very seriously. We found a benchmark discrepancy we couldn’t explain, so we delayed the launch to investigate.
Day 0 (Wed): we saw 2x longer thinking on reasoning-heavy benchmarks (AIME & GPQA) for open source engines compared with @Zai_org API. Same scores, worse token efficiency. Agentic benchmarks looked good.
We decided to investigate further, as overthinking might become a quality problem if max_tokens are reached
Day 1 (Thu): as other non-official providers launched, their APIs had thinking in the range of open-source engines: longer than https://t.co/uYD5faZSKs.
We launched a private preview endpoint with disclaimers to a few customers and worked with them to assess quality
Day 2 (Fri): the official https://t.co/uYD5faZSKs API updates. We rerun benchmarks: reasoning is now similarly long, consistent with vllm/sglang. Rest of the benchmarks, both public and internal, check out too.
We launched GLM-5.3-Flash publicly: https://t.co/zvS4ugxbT6
More details below
Wait!
Ox Alpha.
OX Alpha.
Ox is a male cattle.
Saand is the Hindi word for male cattle.
Alphas are usually the biggest in the group.
Ox Alpha is Bada Saand!
Created by Sarvam, confirmed!
Harvey is a great example of how American companies are building world-class specialized models: they took an open-source base (Kimi K3), post-trained it on legal data, and delivered state-of-the-art performance on legal benchmarks at a fraction of the cost of frontier models. Restrictions that kneecap open models would do nothing to stop Chinese labs from shipping the next Kimi. They would, however, cripple the ability of startups like Harvey to create high-performance, low-cost vertical models. Of course some of the closed labs would love this — it eliminates their competition.
Specialized intelligence is hot in case you didn’t notice:
@harvey’s Tenet
Cognition’s SWE
Cursor’s Composer
GenSpark’s DeepResearch
… the list goes on
Join the movement with @FireworksAI_HQ training platform
Congratulations to @harvey on Tenet, a new SOTA model for legal work! Proud of the work by the @FireworksAI_HQ team - we partnered closely with Harvey to build this model and push the frontier of specialized intelligence!