DeepSeek announced on its platform banner (September 9) that V4.1 Flash releases around September 10 Beijing time: a new-architecture, native-multimodal Flash-class model the lab claims has “comprehensively surpassed V4 Pro across all key metrics” — performance, cost, speed, and task-completion time — with independent numbers not yet published. The Flash series reprices in the same breath: off-peak $0.003 per million input tokens on cache hits, $0.15 on a miss, $0.60 on output (peak double), against the current card’s $0.007 / $0.22 / $0.66. The structural sentence is the routing one: from V4.1 Flash’s launch until V4.1 Pro ships, “all requests to the Pro model will be routed to V4.1 Flash and billed at Flash’s price.”
For an operator that is your lane being swapped, not repriced. Requests you send against deepseek-v4-pro will silently be served by a different, cheaper, multimodal model at Flash rates — roughly a 75% cut on the Pro input card, off-peak — with no opt-out until V4.1 Pro exists, so anything holding V4 Pro assumptions (evals, reasoning-effort behavior, routing pins, cost models) needs re-validating before the swap lands; the test ID deepseek-v4.1-flash-expires-on-0910 had already surfaced in the API in the days before. The take is the sharpest version yet of the story that ran yesterday: the day after the FBI/NSA/CISA advisory named DeepSeek an industrial-scale distiller, the same lab shipped a model it claims beats its own flagship while undercutting it — the cheap-lane cost figure keeps getting cheaper and faster regardless of how the training-data ledger is accounted for.
Source: news.ycombinator.com (DeepSeek platform banner, verbatim), panews.io (Cailianshe relay), api-docs.deepseek.com (current price card), forums.developer.nvidia.com