On September 18, 2026, Beijing-based AI lab StepFun released Step 5 Preview, a 600-billion-parameter reasoning model that posted one of the strongest price-to-intelligence ratios seen in AI model benchmarks so far. The model scored 44 on the Artificial Analysis Intelligence Index v4.3.2, matching xAI’s Grok 4.6 and Moonshot AI’s Kimi K3, while averaging better scores than DeepSeek V4.1 Flash and Gemini 3.8 Flash. It did all of that at an Artificial Analysis measured cost of just $0.71 per completed task on the same benchmark suite.
For anyone tracking the global AI race, the headline is straightforward: another Chinese lab has reached frontier-level intelligence, and it is charging almost nothing for it. This article breaks down what Step 5 Preview is, how it scored, what it costs, and whether a 44 on the Artificial Analysis Intelligence Index genuinely counts as frontier performance.
What Is Step 5 Preview, and Who Is StepFun?
StepFun is a Beijing-based artificial intelligence company that has drawn far less international attention than Chinese peers like DeepSeek, Moonshot AI (the developer of Kimi), Zhipu AI (the developer of GLM), or Alibaba’s Qwen team. Step 5 Preview is its flagship entry into the frontier model race: a proprietary reasoning model reported at roughly 600 billion parameters, available through StepFun’s API.
Artificial Analysis published its Step 5 Preview evaluation results on the model’s official page within a day of release, placing the newcomer directly into comparisons with models from labs many times better known. Key details of the launch include:
- Release date: September 18, 2026, currently in preview status
- Parameters: approximately 600 billion, proprietary with closed weights
- API pricing: $1.00 per 1 million input tokens, $0.05 per 1 million cached input tokens, and $2.70 per 1 million output tokens
- Output speed: about 100 tokens per second, faster than the reasoning-model average of roughly 65
- Time to first token: approximately 2.96 seconds via StepFun’s API
One detail stands out from the full evaluation run: Step 5 Preview generated about 160 million output tokens while completing the Intelligence Index suite, well above the 90 million median for comparable models. It is a verbose model, and that verbosity affects the final bill, a point we return to in the cost section below.
Step 5 Preview Benchmarks: 44 on the Artificial Analysis Intelligence Index
The Artificial Analysis Intelligence Index is one of the most closely watched independent AI benchmarks in 2026. Version 4.3 of the index combines ten separate evaluations, including AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. Together they test reasoning, knowledge, mathematics, coding, and real-world agentic task completion.
Against that backdrop, Step 5 Preview’s score of 44 is remarkable. The median score among reasoning models in its price tier is just 25, meaning the model performs far above average for what it charges. Its average benchmark scores beat DeepSeek V4.1 Flash, which reached 40 at maximum reasoning effort, and Gemini 3.8 Flash, which scored around 40 to 41 at medium and high effort settings. At the same time, Step 5 Preview matches Grok 4.6 and Kimi K3, both of which sit at 44 on the same index version.
| Benchmark | Step 5 Preview | Grok 4.6 | Kimi K3 | DeepSeek V4.1 Flash |
|---|---|---|---|---|
| Intelligence Index | 44 | 44 | 44 | 40 |
| Finance and Accounting Index | 45 | 49 | 47 | 45 |
| AA-Briefcase v1.1 | 47% | 51% | 51% | 47% |
| GDPval-AA v2.1 | 53% | 55% | 51% | 55% |
| AutomationBench-AA | 51% | 67% | 58% | 69% |
| Terminal-Bench 4.0 | 33% | 21% | 13% | 27% |
| SciCode | 59% | 56% | 59% | 52% |
| Humanity’s Last Exam | 46% | 43% | 47% | 39% |
| GDP.pdf | 15% | 17% | 22% | 13% |
| CritPt | 21% | 17% | 23% | 14% |
| AA-Omniscience Accuracy | 42% | 48% | 48% | 46% |
| AA-Omniscience Non-Hallucination Rate | 57% | 66% | 47% | 4% |
| AA-LCR v1.1 | 88% | 80% | 89% | 84% |
| MMMU-Pro | 76% | Not published | 81% | 77% |
Step 5 Preview’s standout showings are Terminal-Bench 4.0, where its 33% tops all three rivals, and long-context reasoning with 88% on AA-LCR v1.1. It also leads Grok 4.6 on Humanity’s Last Exam and CritPt, and it avoids DeepSeek V4.1 Flash’s weakest spot, a 4% non-hallucination rate on AA-Omniscience. Its weaker areas are AutomationBench-AA at 51% and Omniscience accuracy at 42%, although those keep it level with rivals on the composite index.
Scores on this index shift with reasoning effort settings, which is why comparisons require care. Models like Gemini 3.8 Flash and Claude Opus 5 can be run at multiple effort levels, and their scores move accordingly. For deeper context on how Chinese flash-class models compare at similar price points, see our earlier DeepSeek V4.1 Flash vs. GLM 5.3 Flash benchmark comparison.
What Step 5 Preview Costs: $0.71 Per Intelligence Index Task
Benchmark scores only tell half the story in 2026. The other half is what those scores cost to produce. Artificial Analysis measures a “cost per task” figure: the real dollars a model spends completing standardized tasks across its evaluation suite, including every input and output token generated along the way. This captures an important reality that list prices miss, since verbose models burn more tokens and cost more to run.
On that measure, Step 5 Preview lands at $0.71 per task. The comparison landscape looks like this:
- Step 5 Preview (StepFun): $0.71 per task at an Intelligence Index of 44
- Grok 4.6 (xAI): $1.86 per task at an Intelligence Index of 44
- Kimi K3 (Moonshot AI): $2.00 per task at an Intelligence Index of 44
- DeepSeek V4.1 Flash (DeepSeek): $0.27 per task at an Intelligence Index of 40 at max effort
- DeepSeek V4 Pro comes in slightly cheaper at $0.67 per task
- GPT-5.6 Luna is the only model from a US company with a lower per-task cost
- Muse Spark 1.3 and Claude Opus 5 score one point higher at 45, but sit at much higher price points

The result is a clear positioning. There are cheaper models with lower benchmark scores, and there are more expensive models with better benchmark scores. For its level of intelligence, Step 5 Preview appears to be the cheapest option available. It charges $1.00 per 1 million input tokens, $0.05 per 1 million cached input tokens, and $2.70 per 1 million output tokens, far below the output rates of Grok 4.6 ($6.00) and Kimi K3 ($15.00). Anyone comparing API options across providers can explore the full landscape in our LLM API pricing comparison of every major model.
| Model | Input per 1M tokens | Cached input per 1M | Output per 1M | Cost per Index task |
|---|---|---|---|---|
| Step 5 Preview | $1.00 | $0.05 | $2.70 | $0.71 |
| DeepSeek V4 Pro | Not published | Not published | Not published | $0.67 |
| DeepSeek V4.1 Flash | $0.15 | $0.003 | $0.60 | $0.27 |
| GPT-5.6 Luna | Not published | Not published | Not published | Below $0.71 |
| Grok 4.6 | $2.00 | $0.50 | $6.00 | $1.86 |
| Kimi K3 | $3.00 | $0.30 | $15.00 | $2.00 |
Is Step 5 Preview Really a Frontier Model?
DeepSeek V4.1 Flash’s API prices are off-peak figures. Peak-time prices are double: peak hours run 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays, with all other hours, including weekends and Chinese public holidays in full, billed at off-peak rates.
This is the question the numbers invite. If frontier-level intelligence on the Artificial Analysis Intelligence Index is defined by the models hovering around the top of the current table, the frontier line in September 2026 sits at roughly 45. Meta’s Muse Spark 1.3 scores 45 at xhigh reasoning effort (its maximum setting reaches 48), and Anthropic’s Claude Opus 5 scores 45 at medium effort.
Step 5 Preview sits exactly one point below that line, at 44, matching two models widely treated as frontier-adjacent: Grok 4.6 and Kimi K3. Whether it is “at” the frontier or “approaching” it depends on definitions:
- By absolute score, it is one point behind the current frontier-line models at comparable effort settings
- By cost efficiency, it sits on the intelligence-versus-cost frontier, delivering near-frontier scores at a fraction of frontier pricing
- By trajectory, preview models typically improve before general release, so the final Step 5 could close the one-point gap

Measured per dollar, no other model currently combines a score of 44 with a per-task cost anywhere near $0.71. On the intelligence-versus-cost chart that Artificial Analysis publishes, that combination is what defines the Pareto frontier: the set of models that cannot be beaten on both dimensions at once. Whatever label it carries, Step 5 Preview belongs in that set.
The Bigger Picture: China’s Latest Entry at Frontier Level
Step 5 Preview did not appear in a vacuum. It is the latest in a sequence of Chinese models that have climbed into frontier-adjacent benchmark territory at strikingly low prices:
- DeepSeek kicked off the pattern, with V4 Pro and the cheaper V4.1 Flash extending its price-performance reputation
- Moonshot AI’s Kimi K3 reached a score of 44 on the Intelligence Index
- Zhipu AI’s GLM models pushed open-weight performance upward while undercutting US pricing
- StepFun’s Step 5 Preview has now joined them at 44, at the lowest per-task cost of the group

The economics are worth pausing on. As we explored in our analysis of why Chinese AI models are so much cheaper than OpenAI and Anthropic, Chinese labs consistently undercut US rivals on token pricing while closing the intelligence gap. Step 5 Preview sharpens that story: it is not merely cheaper, it is cheaper while matching US models that cost several times more per task.
Some caveats temper the enthusiasm. Step 5 Preview is a preview model, and scores can shift between preview and general release. The model is proprietary, unlike the open-weight releases DeepSeek and Zhipu are known for. Its verbosity, 160 million tokens during evaluation versus a 90 million median, means real-world costs depend heavily on workload. And benchmark scores, however carefully constructed, are proxies for intelligence rather than proof of it.
Even with those caveats, the strategic picture is clear. US labs now face a competitor not just on price, but on the price-to-intelligence ratio that increasingly drives API purchasing decisions.
Frequently Asked Questions
What is Step 5 Preview?
Step 5 Preview is a proprietary AI reasoning model from Chinese lab StepFun, released on September 18, 2026. Reported at around 600 billion parameters, it scored 44 on the Artificial Analysis Intelligence Index v4.3.2 at a measured cost of $0.71 per completed task.
Who is StepFun?
StepFun is a Beijing-based artificial intelligence company that develops large language models. Less internationally known than DeepSeek, Moonshot AI, or Alibaba, it entered the frontier conversation with the Step 5 Preview release and its strong benchmark results.
What is the Artificial Analysis Intelligence Index?
The Artificial Analysis Intelligence Index is a composite AI benchmark published by the independent evaluation platform Artificial Analysis. Version 4.3 combines ten evaluations spanning reasoning, knowledge, mathematics, coding, and agentic task completion. It also publishes cost-per-task figures that measure what each model actually spends to complete the suite.
How does Step 5 Preview compare with DeepSeek V4.1 Flash?
Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index, while DeepSeek V4.1 Flash scored 40 at maximum reasoning effort. At $0.27 per task off-peak, DeepSeek V4.1 Flash costs less per completed task, but Step 5 Preview delivers meaningfully higher average benchmark scores, including clear leads on Terminal-Bench 4.0 (33% versus 27%), SciCode (59% versus 52%), and Humanity’s Last Exam (46% versus 39%).
What does “cost per task” mean for AI models?
Cost per task measures the real dollars a model spends completing a standardized task, including every input and output token generated. It differs from list token pricing because it accounts for how much a model actually writes, making verbose models more expensive in practice than their per-token prices suggest.
Is Step 5 Preview an open-weights model?
No. Step 5 Preview is proprietary, meaning its weights are not publicly available. It contrasts with Chinese peers like DeepSeek and Zhipu AI, which have released open-weight models alongside their commercial offerings.
Which frontier AI model is currently the cheapest per task?
For models scoring around 44 on the Artificial Analysis Intelligence Index, Step 5 Preview is the cheapest at $0.71 per task. Among US labs, only GPT-5.6 Luna has a lower per-task cost, while DeepSeek V4.1 Flash is cheaper still but scores lower on average benchmarks.
Conclusion: The Cheapest Seat at the Frontier Table
Step 5 Preview delivers a score of 44 on the Artificial Analysis Intelligence Index, matching Grok 4.6 and Kimi K3 while beating DeepSeek V4.1 Flash and Gemini 3.8 Flash on average benchmark scores, all at $0.71 per task versus $1.86 for Grok 4.6 and $2.00 for Kimi K3. By almost any reading of the numbers, it offers the lowest cost per unit of intelligence currently available at its capability level.
The model joins a growing roster of Chinese labs, DeepSeek, Moonshot AI, Zhipu AI, and now StepFun, operating at or near the AI frontier while charging a fraction of US pricing. The open questions are whether the full Step 5 release can close the final one-point gap to the frontier line, and how US labs will respond on price as the intelligence-per-dollar bar keeps rising.
