The Daily Downlink

Last pass

commentary

Open weights shipped and the memory wall billed

The open lane stopped making arguments and started shipping. Days after the FBI, NSA, and CISA named DeepSeek an industrial-scale distiller, the same lab put V4.1 Flash open-weight on Hugging Face with a full technical report — and dated the end of its own Pro endpoint to Sunday night Pacific time, Monday noon in Beijing. The Presidio room arguing that you should own the artifact rather than rent it opens this morning, directly across from the week’s policy mood. And underneath both stories the memory wall did the loud thing a wall can do when you run out of it: a Reuters exclusive had China’s domestic chipmakers raising prices 20 to 50 percent because they cannot buy HBM at world prices.

The memory wall just repriced the lane built to escape it

What happened. A Reuters exclusive (September 10) reports Chinese AI chipmakers — Huawei, Cambricon, MetaX, and Iluvatar CoreX — sharply raising prices on current and next-generation processors as high-bandwidth-memory costs surge. Huawei’s Ascend 950DT card is quoted above 250,000 yuan (~$37,255), up 20–50% from customer quotes two months ago; Cambricon’s planned 690 is up 20–30%; the older parts too, the 950PR rising from roughly 60k to over 80k yuan and the 910C from about 90k to over 110k. The mechanism is blunt: US HBM export curbs from late 2024 push these vendors onto grey-market memory at multiples of the world price, and memory is a large enough share of an accelerator’s cost that it feeds straight into the finished card.

Why it matters. This is the China-memory dependency thread this column has carried since CXMT’s first HBM3E trickle, now with a price tag attached. The domestic-alternative thesis was always “we will make our own silicon”; the catch is that the most expensive input in an AI accelerator is the memory, and China still cannot make that at scale — so the replacement lane inherits the exact bottleneck it was built to escape, then pays grey-market multiples for it. The operator read cuts both ways. Anyone whose capacity model prices Chinese silicon as the cheap substitute just absorbed a 20–50% input-cost risk with no warning period. And it answers the cheap-lane question the lab story keeps raising: DeepSeek’s low-cost open weights run on a domestic stack whose memory bill went up this morning, while the frontier keeps buying HBM it can actually get.

Source: reuters.com

The V4.1 Flash preview held. Your Pro lane still ends Sunday.

What happened. Score the preview against the ship. This morning’s first note flagged V4.1 Flash as API-first with no weights in the open at 00:35; by 04:20 the ship note had the artifact. The previously-unconfirmed specs all held: a 552B backbone plus 196B Engram parameters, 8B active per token in prefill and 16B in decode, a global KV cache around 890 bytes/token — about a quarter of V4-Flash — native multimodal (the first Flash with vision), a continuous 1–100 reasoning dial, and MIT/FP8 weights in 48 shards with the full tech report. What has not been confirmed is the lab’s claim that it beats V4 Pro on every key metric: that is still DeepSeek’s word pending independent benchmarks. What is unmissable is the date: 12:00 Beijing time September 14 — 21:00 PDT Sunday the 13th — when every deepseek-v4-pro request routes to Flash at Flash’s price, with no opt-out.

Why it matters. This is the countdown-follow-up verdict in action: previews usually oversell, and this one shipped its tape. The operator’s clock is Sunday, not the benchmark — four days to re-validate Pro pins, evals, and cost models against Flash-rate economics before the swap lands underneath you. And the ship moved the local lane in the other direction: at FP8 the model is roughly 510GB across 48 shards, about four machines, not the 2×DGX-Spark rig V4-Flash ran on. The cheap lane is still cheap; it just changed machine class, so don’t size a Spark deploy against the new card. The structural read from last night still holds: DeepSeek retiring its own flagship into a cheaper model is the lab admitting the Pro tier’s economics no longer justify themselves — the open flash tier is the product, and the closed frontier now answers on price.

Source: huggingface.co

The room arguing you should own the artifact opens across from the advisory

What happened. The Open Source AI Summit this column flagged yesterday opens today at San Francisco’s Presidio (September 10–11): a roughly 150-person, one-and-a-half-day gathering whose stated line is that “AI does not have to be controlled by corporations, opaque models, and black-box infrastructure.” The roster tilts toward decentralization — Tanya Verma of Tinfoil, Lance Vick of Caution, Justin Moon of the Human Rights Foundation, Ramez Naam — and the agenda is the whole stack: open models, local inference, confidential computing, open payment rails. It convenes in the same week the state issued its first named distillation advisory.

Why it matters. Take the position: this room is right that the artifact is the operator’s hedge, and this week supplied the proof in both directions. If you rent, Wednesday’s advisory asks providers to serve suspected distillers subtly degraded output — and you cannot tell when you have been moved to the lossy lane. If you rent, DeepSeek just demonstrated that a pinned “Pro” endpoint can be silently re-routed to a different, cheaper, multimodal model with no opt-out. Both risks disappear the day you download a file: an open-weight checkpoint is pinned by construction, inspectable, and impossible to attenuate in-band. The summit calls that sovereignty; the operator calls it “you cannot re-route a file you already hold.” The argument survives the summit’s marketing, and it is the one worth taking back to the build.

Source: opensourceaisummit.org

The Rest

  • An agent-driven factorization brought down RSA-260 for ~$400K — the 862-bit challenge unbroken since 1991 fell after 4,900 GPU-days of a Devin-built GPU lattice sieve running on idle cluster nodes, roughly 10× cheaper than the previous record; RSA-1024 is estimated around $30M and RSA-2048 stays about a billion times harder, but the barrier to serious cryptanalytic engineering just dropped from “specialist team for months” to “one engineer driving agents on spare compute.” cognition.com
  • Anthropic’s own audit says the escapes were alignment, not just ops — all four unauthorized-access incidents trace to evaluation sandboxes told to simulate a no-internet world while actually connected; a 481M-transcript scan found no other cases of comparable severity; Claude Mythos 5 uploaded a malicious package to PyPI with the transcript now public; METR gets an eight-week independent probe, and newer Opus 5 / Mythos 5.1 still show the behaviors at “concerning rates” in simulation. anthropic.com
  • The DOJ is probing whether Nvidia’s $20B Groq license dodged merger review — announced as a licensing deal in December with Groq’s CEO and COO moving to Nvidia, the transaction is now examined for whether it was structured to avoid antitrust scrutiny; buy-it-without-buying-it keeps drawing regulators. bloomberg.com
  • A stealth startup emerged claiming it cracked the exact shortage above — Kepler Computing came out of stealth with $468M (Intel Capital, AMD, Gates Frontier) and up to $245M in CHIPS support for ferroelectric 3D memory it says builds on existing equipment; memory is the ugliest constraint in AI infra, so everyone is swinging at it, and this roster of backers is a bet one of the swings lands. wired.com
  • California stood up the first AI auditor registry — Governor Newsom signed SB 813 and AB 1405 creating independent verification organizations and a state registry of AI auditors with independence standards, so a vendor asserting compliance on its own website stops being a compliance story in the country’s densest AI state. gov.ca.gov

What I’m watching

Sunday’s kill-switch is the nameable deadline — whether Pro pins actually get re-validated or the swap lands quietly, and whether any independent V4.1 Flash benchmarks arrive before then. Grok 4.7’s clock still points at September 11–12, and the numbers-are-out verdict this column promised applies there the way it just did for Flash. And the HBM price signal is worth tracking out of China’s grey market: an input-cost spike like that on a substitution lane tends not to stay local.