DeepSeek shipped it overnight. V4.1 Flash is officially released (September 10, Beijing time, per the API notice and the account’s announcement thread that tops Hacker News this morning) — and this changes the shape of this morning’s earlier note, which correctly flagged it as API-first with no weights in the open. The weights are now out: deepseek-ai/DeepSeek-V4.1-Flash on Hugging Face, MIT, FP8, 48 safetensors shards and a full tech report. The architecture is a real departure, not a retrain: a 40-layer causal encoder-decoder (20/20) with 552B backbone + 196B Engram parameters, activating only 8B per token in prefill and 16B in decode, a global KV cache down to 890 bytes/token (about a quarter of V4-Flash — the ~1GB/1M-context regime), native multimodal input, and a continuous 1–100 reasoning-effort dial. Claims of topping its own V4 Pro remain the lab’s word until independent numbers land, but the open artifact is the tell: the distiller the NSA/CISA/FBI named this week shipped cheaper-and-better open weights the same week.
For an operator, the date is the headline. DeepSeek’s notice to API users postpones the V4-Pro retirement to 12:00 Beijing time on September 14 — 21:00 PDT on the 13th — at which point every deepseek-v4-pro request routes to V4.1 Flash and bills at Flash’s price. Anyone holding Pro pins, evals, or capacity models gets four days to re-validate against Flash-rate economics, and there’s no opt-out. On the local lane, note the footprint: V4-Flash (284B/13B active) fit a 2×DGX-Spark rig; V4.1 Flash’s 552B+196B at FP8 is community-sized around 510GB across ~4 machines — the cheap lane is still cheap, it just changed machine, so don’t size the Spark deploy against the old card. The position: DeepSeek retiring its own flagship into a cheaper model is them admitting the Pro lane’s economics no longer justify themselves — the open flash tier is the product, and the closed frontier has to answer on price.
Source: huggingface.co (weights + tech report), news.ycombinator.com (API notice verbatim), news.ycombinator.com (tech-report breakdown)