DeepSeek to launch V4.1 Flash on Sept 10, routing all V4 Pro API traffic to it
DeepSeek will officially release its V4.1 Flash model around September 10, 2026 Beijing Time, and until V4.1 Pro ships, every API request sent to the Pro model will be routed to V4.1 Flash and billed at Flash's price. The company announced the plan in a notice that reached the front page of Hacker News, where the thread drew 399 points and 209 comments.
DeepSeek says V4.1 Flash has "comprehensively surpassed V4 Pro across all key metrics, including performance, cost, speed, and task completion time," following what it describes as extensive internal and external testing. The routing decision is framed as a commitment to user responsibility, with the company asking users to send feedback if they hit problems while comparing V4 Pro against V4.1 Flash.
The model has been in users' hands for days. An internal beta of what was described as an intermediate version of V4.1 Flash opened on September 8, 2026, according to a translation of the announcement circulated on r/LocalLLaMA. That beta runs under the model name deepseek-v4.1-flash-expires-on-0910, keeps the existing base_url unchanged, and carries pricing identical to deepseek-v4-flash with a rate limit of 20 concurrent requests per account. The beta announcement described a new model architecture with native multimodal support, stronger capabilities, faster speeds, and lower costs.
Pricing changes land with the launch. Effective 12:00 Beijing Time on September 10, 2026, Flash series off-peak rates are $0.003 per unit for input cache hits, $0.15 for input cache misses, and $0.6 for output, with peak-hour prices double the off-peak rates. DeepSeek told users to plan usage accordingly.
The forced migration drew pushback in the Hacker News thread. One commenter argued that users who validated a workflow on V4 Pro may not want it suddenly running in production on V4.1 Flash, and suggested keeping V4 Pro available but deprecated for a defined period before removal. The same commenter noted that because these are open-weights models, providers such as Together.ai or OpenRouter can keep serving V4 Pro as long as they continue to host it.
Others in the thread disputed how much that matters. One reply said model behavior is non-deterministic, so a validated workflow is not a guarantee in the first place. A counter-reply from someone working for an education department running a student chatbot said model changes go through painstaking content safety reviews, and that every prior model upgrade measurably changed how well the system prompt's restrictions on discussing sex, drugs, and mental health with children were followed. A further reply argued that API providers offer no guarantees about serving behavior, including resource allocation and versioning, and that in-house or raw-hardware hosting is the only way to control it.
What remains unconfirmed is the exact release timing and whether the routing of Pro traffic is temporary or open-ended. DeepSeek's notice ties the arrangement to the period before V4.1 Pro is released, but gives no date for that release. The company has also not published the benchmark numbers behind its claim that Flash beats Pro on every metric, so that comparison currently rests on DeepSeek's own statement rather than independent evaluation.
DeepSeek is forcing every V4 Pro API user onto a newer, cheaper model without an opt-out, testing how much control customers have over which model version serves their production traffic.