Most of the reasoning a model does isn't the answer. It's the model talking itself into the answer.
With Ember-1 we went after that overhead: about 40% fewer tokens on Kimi K3, and the quality held up on live customer A/B tests, coding workloads and external benchmarks.
I spent a lot of late nights watching models ramble. Glad to see it paid off. Try it on Fireworks.
Most of the reasoning a model does isn't the answer. It's the model talking itself into the answer.
With Ember-1 we went after that overhead: about 40% fewer tokens on Kimi K3, and the quality held up on live customer A/B tests, coding workloads and external benchmarks.
I spent a lot of late nights watching models ramble. Glad to see it paid off. Try it on Fireworks.
Ember-1 is a specialized model from Fireworks Research designed to make every token go further.
Built on Kimi K3, it produces shorter reasoning traces, using roughly 40% fewer tokens while maintaining top-tier quality.
https://t.co/PLD1l5aeh4
Today we're launching the Specialized Intelligence Index (SII): one destination for real-work benchmarks across industries, built by the teams that use them every day.
Hear from Fireworks co-founder @the_bunny_chen on the importance of specialized benchmarks:
Today we're launching the Specialized Intelligence Index (SII): one destination for real-work benchmarks across industries, built by the teams that use them every day.
Hear from Fireworks co-founder @the_bunny_chen on the importance of specialized benchmarks:
Jev is all the rage this weekend, and honestly, fair.
But every LLM already ends in a softmax over its vocab. Point it at "yes"/"no" and congrats, calibrated classifier.
Wrote this up a few months ago. $2, plus a proof you never renormalize.
The underrated part here isn’t "LLMs can classify" - it’s that you can get well-calibrated class probabilities without adding a new head or touching the architecture.
Once you map labels → tokens and fine-tune with standard LM loss, the optimization naturally pushes probability mass onto the label tokens. In the limit, renormalizing over the label set becomes almost redundant - the raw next-token distribution already behaves like a calibrated classifier.
This is what makes the approach so practical in production: you reuse the same model, the same serving stack, and the same training loop, and still get reliable confidence scores for routing, moderation, intent detection, etc. The "classifier" is just a different way of reading the base model’s distribution.
Divergent numerics are a common cause of RL training collapse. That's why we spend so much time at @FireworksAI_HQ aligning the numerics between trainer and inference engine. Check out our latest blog post for the details on how we achieve zero numerics drift.
The hard part of reinforcement learning on a frontier model is the infrastructure that keeps training and inference numerically identical: zero KLD, end to end.
We've solved this challenge, and are now offering it as a managed service, starting with GLM 5.2.
The underrated part here isn’t "LLMs can classify" - it’s that you can get well-calibrated class probabilities without adding a new head or touching the architecture.
Once you map labels → tokens and fine-tune with standard LM loss, the optimization naturally pushes probability mass onto the label tokens. In the limit, renormalizing over the label set becomes almost redundant - the raw next-token distribution already behaves like a calibrated classifier.
This is what makes the approach so practical in production: you reuse the same model, the same serving stack, and the same training loop, and still get reliable confidence scores for routing, moderation, intent detection, etc. The "classifier" is just a different way of reading the base model’s distribution.
Most people still think of LLMs as free form text generators. In practice, a huge number of real production workloads are actually classification problems. Things like intent detection, routing, moderation, or picking from a small set of labels.
The interesting part is that you don’t need a separate classification head or a custom model to do this. You can turn any generative model into a classifier by mapping each class to a token and letting the model’s next token distribution act as the class probability distribution.
It stays fully compatible with standard fine tuning. It keeps your inference stack simple. And the best part is that you can get well calibrated class probabilities for as little as a couple of dollars in compute. We used Qwen3 4B, trained with LoRA on Fireworks, and the total cost was around 2$.
Our latest blog breaks down the full workflow. How token based classes work. How calibration naturally emerges during training. And how you can serve high quality classifiers without changing model architecture at all.
If your team is building routing systems, moderation pipelines, intent detection, or any production workflow that depends on confidence scores, this is worth a read: https://t.co/MN5vQSu0k5
We’ve spent the last several months working closely with customers building real agents - research agents, coding copilots, multi-step workflows with tools. The pattern was clear: traditional RL setups don’t match how production agents actually behave.
Today we’re introducing Fireworks RFT, a reinforcement fine-tuning system built around real agent trajectories: multi-turn, tool calls, retries, everything. Instead of asking teams to redesign their agents for RL, we plug directly into the workflow they already have.
The early results have been impressive - agents that outperform closed models, use tools more effectively, and run at dramatically lower cost.
If you’re working on agents for research, coding, structured reasoning, or domain workflows, this is a meaningful step forward.
🚀 Fireworks Reinforcement Fine-Tuning (RFT) launched!
After many months of iteration with real world use cases, we are excited to launch Fireworks RFT public preview. It’s a managed RL service that turns open frontier models (e.g. DeepSeek V3, Kimi K2) into custom agents for tools, code & reasoning.
Our design principle is to bring RL to integrate directly with your agents in production, instead of you simulating or redesigning your agents to do RL. Fireworks RFT trains on full agent trajectories (multi-step, tools, retries), so the model learns your workflow in production.
Here are real-world results:
👉 Genspark Deep research agent → +10% quality vs SOTA closed model, +33% tool calls, ~50% lower cost
👉 Vercel V0 coding agent → +33% better than closed models, 93–94% error-free, 40× faster.
🔥 Fireworks RFT public preview is open now - start training today for free 👇
https://t.co/cRhz9IXuMO
We partnered with @vercel to post-train a model using Reinforcement Learning, achieving performance that surpasses a frontier model on a code editing task.
A particularly interesting aspect of this use case is the length of the generations. Because code completions are long, generation length has a major impact on RL iteration speed - the auto-regressive sampling during inference, not the weight updates, becomes the bottleneck.
That’s why having the fastest inference platform in the world is critical for building an efficient large-scale RL training system.
If you’re interested in exploring RL at scale, reach out to us at @FireworksAI_HQ
Fireworks teamed up with @Vercel to set a new standard for developer AI.
Lightning-fast and super-accurate AI code generation with Vercel's v0 tool achieved an incredible 93% error-free rate and 40x faster speeds!
Dive into the full story and see how they did it: https://t.co/oOjtwda9Hb
#AI #CodeGeneration #Vercel #FireworksAI #DevTools #MachineLearning
Check out our latest work applying Reinforcement Learning (RL) at scale to outperform a frontier model on the Deep Research use case. Deep Research isn’t a niche scenario - it’s a core capability that frontier labs actively optimize for - which makes this result especially notable.
If you’re interested in exploring post-training an open model with RL, reach out to us at @FireworksAI_HQ .
And while you’re at it, give @genspark_ai a try - their product is fantastic.
@genspark_ai has transformed its Deep Research Agent by leveraging Fireworks’ Reinforcement Fine Tuning (RFT) solution to train an open-source model in just four weeks.
This powerful solution delivered a 12% quality improvement and 33% more tool calls compared to a frontier closed model, resulting in a significant 50% cost reduction and a superior quality result.
This collaboration highlights the power of open models and reinforcement learning in delivering better AI solutions.
Check out our case study to learn more:
https://t.co/89rLVpSNHf
I'm very excited to share that @FireworksAI_HQ has raised $250M in Series C funding co-led by @lightspeedvp and @IndexVentures , and participation from @sequoia Capital and @EvanticCapital, bringing up valuation to $4 billion. In total, we raised $327M from prior rounds led by @benchmark and @sequoia, with participation of strategic investors including @nvidia , @AMD , @databricks, @MongoDB , and many angel investors.
When we founded Fireworks in 2022, the vision was simple, but the problem was complex: give builders the speed, cost, and control they need to win. Our mission is to reach Artificial Autonomous Intelligence – automated product and model co-development to reach maximum quality, speed and cost-efficiency using generative AI. This round propels our execution to expand our current product of tuning and inference platform towards Artificial Autonomous Intelligence.
↗️ We have onboarded hundreds of thousands of developers to customize the latest models to create unique applications and impactful user experiences. 10x from Series B.
↗️ 10K+ organizations are running on Fireworks. Companies including Notion, Shopify, Uber, GenSpark, Vercel have built and scaled on Fireworks.
↗️ We process more than 10 trillion tokens daily. More than 20x from Series B.
We will use the new funding in the following investment:
📖 Deepen our research in post-training and inference alignment to maximize quality, speed and cost efficiency of computation, including R&D in system research and algorithmic research.
🛠️ Expand our product towards a comprehensive tool-chain for the new user experience creation lifecycle, centered around product and model co-design and automation, from model evaluation, reinforcement learning to ultra fast inference engine.
🌎 Grow our computation footprints 3x higher in the next one year, and continue R&D to minimize computational cost and maximize system utilization.
This round is not just evidence that our bet three years ago is driving a big impact, it's concrete proof of the trust and commitment we've built with our customers, partners, and our community.
Join our mission to build Artificial Autonomous Intelligence https://t.co/esv1l4HrCo
At @FireworksAI_HQ, we think of post-training not as a one-size-fits-all solution, but as a decision framework that balances data availability, cost, and desired quality.
Where do you see the most significant gaps today - better datasets, training algorithms, base models, reward design, or evaluation? (4/x)
Choosing the Right Post Training Technique 🧠
Training doesn’t end when a foundation model is released. The real challenge is post-training - aligning a model with specific behaviors, domains, and quality expectations.
Let's break down the main techniques shaping the landscape today. 👇 (1/x)
So how do you choose? 🤔 A quick guide:
➡️ If you have lots of labeled ground-truth data → SFT is the right place to start.
➡️ If you want higher quality with preference data while keeping complexity in check → DPO is the pragmatic step forward.
➡️ If your task is verifiable and rule-driven → RLVR can unlock scalability.
➡️ RLHF → Powerful but rarely practical outside frontier research. (3/x)
11/x These advancements, driven by GSPO's robust training approach, have contributed to the exceptional performance improvements in the latest Qwen3 models. You can try out these models on the @FireworksAI_HQ platform to experience the capabilities powered by this new approach.
3/x The instability in GRPO stems from a fundamental issue: its misapplication of importance sampling weights at the token level. Importance sampling relies on averaging over multiple samples to correct for distribution mismatches. But GRPO's token-level weights, based on a single token sample, introduce high-variance noise that accumulates over long sequences and is amplified by clipping, often causing model collapse.
10/x While GSPO primarily focuses on sequence-level optimization, the paper also introduces GSPO-token, a variant designed for scenarios like multi-turn RL where finer-grained advantage adjustment at the token level might be desired. It achieves this by allowing token-wise advantage customization.