Step 5 Preview pushes our intelligence–cost Pareto frontier outward.
It comes in at 44 on the Artificial Analysis Intelligence Index, $0.71 per task.
Thanks @ArtificialAnlys for putting Step 5 Preview through the full evaluation.
More to come on Oct 15.
We’ve been quiet for too long. It’s time for a change.
Over the past few months, we’ve been pushing toward the frontier, step by step. And now, things are finally coming together.
Welcome to StepFun. Come build better open models and more general agents with us!
Introducing Step 5 Preview: Advancing the Pareto Frontier.
Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance.
- 600B total / 27B active MoE, with 1M context + Vision
- Substantially lower task cost at comparable intelligence
- Broad software engineering capabilities with sustained execution over long horizons
Try Step 5 Preview: https://t.co/fC7HHlWKFn
Model page: https://t.co/4caJR2YGD3
Open weights on Oct 15.
Step 5 Preview is built for professional knowledge work — from large-scale research to structured analysis and interactive reporting.
It coordinates research at scale and turns evidence into finished, auditable deliverables. In one agent action, it coordinated 950 web fetches and assembled 300,000 monthly records across 1,000 locations over 25 years; in another, it produced a 17-sheet analytical workbook with source reconciliation, formulas, and trend models.
Its outputs span technical engineering, creative production, analytical reporting, and public communication.
Step 5 Preview works across software environments and sustains execution over long horizons.
Its capabilities extend from software engineering and web applications to 3D workflows and programmable hardware.
Over longer horizons, Step 5 Preview keeps track of prior results, uses execution feedback to decide what to try next, and continues iterating. We test this behavior in runs lasting up to 24 hours, including tasks involving GPU kernel optimization and automated post-training.
Stepfun just announced their new SOTA model
step 5 preview is a 600b model with only 27b active params, and early numbers are looking insane:
> 67.7 on deepswe
> beats gemini 3.8 flash
> 1m context window + image and video input
> $1 input and $2.70 output
> beats most of chinese models
it’s already available in api and chat, so you can play around if you want
already got access and running some tests now, will post results soon
which model should i test it against?