DeepSeek plans to release DeepSeek V4.1 Flash around September 10, 2026, Beijing time, alongside a Flash price cut and a temporary routing change that sends V4 Pro API requests to the new Flash model. Independent benchmark sites have not yet published scores for the new model, so this article covers only what can be confirmed right now. A fuller benchmarks update can follow once test results appear.
What Is Launching Around September 10
DeepSeek notified API users that V4.1 Flash is scheduled for release around September 10, 2026, Beijing time. Chinese tech outlets including IT Home, Sina Finance, and Ifeng Tech reported the same notice on September 9, describing V4.1 Flash as an interim model between the current V4 Flash and V4 Pro lines.
DeepSeek said the model was tested internally and externally, and claimed it surpassed the current V4 Pro on performance, cost, speed, and total task completion time. Those comparisons remain company claims until independent evaluations publish results for coding, reasoning, agent, and multimodal workloads.
No release date or pricing has been announced for the follow-up V4.1 Pro model. Until that model arrives, the new Flash model will carry traffic from both tiers.
New Flash Pricing From September 10
The new Flash price schedule takes effect at 4:00 UTC on September 10, which is noon in Beijing, midnight Eastern time, and 9 p.m. Pacific time on September 9. Peak hours are 1:00 to 4:00 a.m. UTC and 6:00 to 10:00 a.m. UTC, Monday through Friday, with all other hours billed at off-peak rates.
In U.S. dollars, per 1 million tokens:
| Billing item | Off-peak | Peak |
|---|---|---|
| Input, cache hit | $0.003 | $0.006 |
| Input, cache miss | $0.15 | $0.30 |
| Output | $0.60 | $1.20 |
Chinese outlets reported the same DeepSeek V4.1 Flash schedule in yuan, per 1 million tokens:
| Billing item | Off-peak | Peak |
|---|---|---|
| Input, cache hit | 0.02 yuan | 0.04 yuan |
| Input, cache miss | 1 yuan | 2 yuan |
| Output | 4 yuan | 8 yuan |
That is a cut from the previous Flash rates. In yuan, cached input falls from 0.05 to 0.02 (down 60%), uncached input from 1.5 to 1 (down about 33%), and output from 4.5 to 4 (down about 11%). In dollar terms, the listed V4 Flash off-peak rates of $0.007, $0.22, and $0.66 fall to $0.003, $0.15, and $0.60, reductions of about 57%, 32%, and 9%.

For background on how DeepSeek prices cache hits and why they matter for long agent runs, see our earlier breakdown of DeepSeek V4 rates and cache economics.
Pro Requests Will Be Served by V4.1 Flash
After V4.1 Flash launches and before V4.1 Pro is released, DeepSeek will route all requests sent to the Pro endpoint to V4.1 Flash and bill them at Flash prices. The endpoint name stays the same, but the underlying model changes.
The saving is largest for workloads currently pointed at V4 Pro. Against listed V4 Pro off-peak rates of $0.022 for cached input, $0.66 for uncached input, and $1.98 for output, the Flash schedule represents reductions of roughly 86%, 77%, and 70%.
Developers should treat this as a forced migration rather than a simple discount. Applications tuned around V4 Pro output style, tool use, or reasoning behavior should be retested against V4.1 Flash, even though no API configuration change is required. For context on how Flash-class models have compared with Pro-class models on coding work, see Qwen3.8-Flash matching DeepSeek V4 Pro on coding benchmarks.
What the Two-Day Beta Confirmed
On September 8, DeepSeek opened an interim V4.1 Flash build for limited testing under the model ID deepseek-v4.1-flash-expires-on-0910, which stops working on September 10. Confirmed details from the beta:
- New model architecture with native multimodal input for text and images.
- Billed at existing V4 Flash rates during the test period.
- Capped at 20 concurrent requests per account, so it was a test endpoint, not a production endpoint.
- Called through the existing API address by changing only the model name.
- DeepSeek asked testers whether the interim build could fully replace production V4 Pro, a question the routing plan has now answered at the product level.
The native image input continues a direction DeepSeek explored with the experimental V4 Flash Vision model, which added image input to the Flash line in August.
One Early Design-Only Signal, With Caveats
A third-party design arena, the Open Design LLM arena for design tasks, displayed an early V4.1 Flash result of 81.2 out of 100, behind one rival at 82.7 and ahead of another at 77.6, with sub-scores of 28.4 out of 30 for requirement fulfillment and 52.8 out of 70 for design quality, plus a 57.7% delivery rate, 5.3-minute completion time, and $0.023 reported cost per run.
This is a narrow design-task comparison, not a broad test of coding, reasoning, or agent workloads. Open Design has not published a methodology page explaining how it computes the headline Quality rating, so treat the number as an early signal rather than a verified benchmark. The Hugging Face page for DeepSeek V4 Flash shows how much broader a full benchmark table looks for the prior model, spanning knowledge, reasoning, code, math, and long-context tests.
What Is Still Unconfirmed
As of September 9, major independent benchmark boards including Artificial Analysis and Vals.ai list no scores for V4.1 Flash. That means the following remain unconfirmed:
- Standard coding and reasoning benchmarks for the release build.
- Agent benchmarks comparable to the Terminal Bench 2.1 and DeepSWE figures published for V4 Flash.
- Speed claims from anecdotal reports, which have no published test setup.
- Rankings on Chinese-language boards, where machine translation makes the test scope unclear.
- Any comparison with specific rival Flash-class models at equal effort settings.
The sensible position is to report the launch, pricing, and routing as confirmed product facts, and to wait for independent results before judging whether a cheaper Flash model can carry premium-tier work. The official DeepSeek API pricing documentation is the authoritative source for the rates once the change takes effect.
Frequently Asked Questions
When does DeepSeek V4.1 Flash launch?
Around September 10, 2026, Beijing time. The new pricing takes effect at 4:00 UTC on September 10.
What happens to V4 Pro API requests?
Until V4.1 Pro is released, requests sent to the Pro endpoint are served by V4.1 Flash and billed at Flash prices.
What are the new Flash prices?
Off-peak: $0.003 per million cached input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens. Peak rates are double, during 1:00 to 4:00 a.m. and 6:00 to 10:00 a.m. UTC on weekdays.
What was the expires-on-0910 model?
A limited beta ID for an interim V4.1 Flash build, live from September 8 until September 10, capped at 20 concurrent requests per account and billed at V4 Flash rates.
Are there confirmed benchmarks for V4.1 Flash?
No broad independent benchmarks have been published yet. One design arena shows an early 81.2 score, but that covers only design tasks with an undisclosed methodology.
Conclusion
V4.1 Flash is a product move first and a benchmark story second. The confirmed facts are the release window, the lower Flash price schedule, and the temporary routing of Pro traffic to the cheaper model. Performance judgment should wait until independent boards publish scores for the release build, at which point this article can be updated with a fuller comparison.
