WEDNESDAY 9 SEPTEMBER 2026 latent·wire 62 PIECES ON FILE
← AI NewsAI News

DeepSeek V4 Flash 0731 lists on OpenRouter at a fraction of official pricing

A Reddit user on r/LocalLLaMA flagged that DeepSeek V4 Flash 0731 is listed on OpenRouter at prices roughly four times below DeepSeek's own official rates, raising questions about how the discount is possible. According to the post, OpenRouter lists non-peak input and output prices of $0.05 and $0.16 per 1 million tokens for the model, while official DeepSeek pricing sits at $0.22 and $0.66 per 1 million tokens. The user, who says they are setting up GPU nodes and plan to self-host models under 300B parameters, called the OpenRouter figure "too cheap" for a model they consider overkill for simple summarization and email-sifting tasks but capable of handling complex workloads that smaller models would struggle with.

The post also notes a discrepancy in cached-read pricing. DeepSeek direct charges $0.007 per 1 million tokens for cached reads, while the OpenRouter listing, surfaced via openinference, shows $0.013 for the same. The user did not offer an explanation for either gap, and the thread's framing is a question rather than a conclusion.

The pricing comparison matters because the user argues the flash model undercuts smaller, ostensibly cheaper alternatives on the same marketplace. They cite phi-4, which they describe as having a small context window, plus qwen3.6 35B-A3B and qwen3.5-9B, as models that are similar in cost or more expensive per token despite offering less capability. Their self-hosting targets, qwen3.8 27B and qwen3.8-flash-next, are, in their words, "BY FAR more expensive to acquire over API" than the DeepSeek flash listing, which is part of why they still see value in running their own hardware: true privacy.

The post is a single community observation and does not resolve why the OpenRouter price diverges from DeepSeek's official rate. Possible explanations, such as a promotional discount, a pricing error, or a difference in how the two providers meter tokens, are not addressed in the excerpt. The user's message is cut off mid-sentence, so their full reasoning and any follow-up discussion in the thread are not available.

No official statement from DeepSeek or OpenRouter about the pricing gap is cited, and there is no other recent coverage of the same subject to corroborate or contradict the figures. The numbers in the post are presented as observed listings rather than confirmed contract terms, and marketplace prices can change without notice.

For anyone comparing API costs, the practical takeaway from the post is that per-token price alone does not settle the choice between a frontier flash model and a smaller dedicated one, especially when privacy and self-hosting are on the table. The user's plan is to keep their home datacenter for workloads where data control matters, even if the API price for DeepSeek V4 Flash 0731 looks cheaper on paper. What remains unknown is whether the OpenRouter listing holds, whether DeepSeek adjusts its own pricing to match, and whether the gap reflects a real cost advantage or a temporary anomaly.

Why it matters

A community-reported fourfold gap between OpenRouter and official DeepSeek pricing for V4 Flash 0731 raises unresolved questions about how API prices for frontier models are set and whether the discount is sustainable.