The Daily Downlink

Last pass

commentary

Open models got a balance sheet — and the distribution point went portable

Openness stopped being a licensing argument this week and became a profit-and-loss statement. Z.ai — the lab behind the GLM family, the open models this column actually runs — reported first-half revenue of $141.95 million, up nearly 400 percent year over year, with its open platform and API business up 28x to $122.8 million, 86 percent of the total. In the same window, the UAE’s MBZUAI-backed IFM released the largest fully-open model fleet in history, and Nvidia confirmed it is taking DGX-Spark-class compute mobile in October. The open-vs-closed fight is no longer settling on press releases; it is settling in filings, license files, and purchase orders.

The model you run now has receipts

What happened. Z.ai’s first interim report since its January Hong Kong IPO: H1 revenue of $141.95 million, up nearly 400 percent year over year; net loss narrowed to $308.33 million from $350.87 million; R&D spending up 33.6 percent to $317.14 million. The open platform and API cloud business did $122.8 million, up 28x and 86 percent of revenue, against an enterprise on-premises business that fell 54.6 percent to $9.98 million. The company reports 7.4 million enterprise and developer users on its model-as-a-service platform, inference over a cluster of more than 100,000 domestic Chinese chips, and unit token cost down 80 percent from the start of the year.

Why it matters. This is the first real balance sheet for the open tier, and the number to read is not the headline growth — it is that 86 percent of the money comes from an API business, not from server sales or consulting. That settles a question this column has been pricing since Ox Alpha turned out to be GLM-5.3-Flash: open weights can carry stated pricing power, as a business, not a perpetual subsidy. The texture that matters for operators is how Z.ai now sells it — long-horizon task success rate and unit intelligence cost rather than single-turn quality or raw token volume, which is the harness-and-price-of-a-task framing moving into a securities filing. And the 100,000-chip domestic cluster with costs down 80 percent is the computational-independence story: the cheapest open models now run on a stack that owes nothing to the landlord, even as Meta answered the open tier by shipping its cheapest frontier model closed.

Source: constellationr.com

The Gulf out-opened everyone, including the training data

What happened. MBZUAI’s Institute of Foundation Models released K2 Horizon on September 3: six models from 0.9 billion to 375 billion parameters, under Apache 2.0, with weights, code, training data and full methodology — a fleet that ships the receipts, not just the artifact. Available day one on Hugging Face, vLLM and SGLang, with API access through Compass, Cerebras and Nebius. IFM claims state-of-the-art-in-class results for the 0.9B, 3.7B and 7B models, a 36B-A4B mixture-of-value-attention architecture, and a diffusion-distillation technique that roughly triples generation speed.

Why it matters. Watch what the counter-program is aimed at: Meta shipped its cheapest frontier model closed this week with its open-weights promise still un-landed, and IFM’s founder answered in kind — “open source is much more than open weights.” Sovereign Gulf funding just produced the most transparent frontier artifact in existence, down the same capital lane that is bankrolling Together’s capacity. For an operator the realistic read is the small end of the fleet: a 7B claiming best-in-class under 10B and fitting on a phone is the local-first tier, and the 0.9B is aimed at a wrist. What makes those claims checkable rather than vendor noise is the license-file discipline that full data plus Apache-2.0 actually buys — reproduce and verify, not take-it-on-faith.

Source: ifm.ai, wired.me

The landlord carried the local layer into a laptop

What happened. At IFA 2026, Nvidia confirmed RTX Spark Windows PCs ship in October, with Lenovo’s Yoga Pro 9n and 9n 2-in-1 plus an Acer desktop joining the six OEMs already announced. The onboard spec is DGX-Spark-class: a 1-petaflop RTX Blackwell GPU with up to 128GB of unified memory and a 20-core Grace CPU in a thin 3.6-pound laptop. Alongside it, the company shipped PAIR, a free tool that routes inference across the PCs on a home network, and promised a Windows Agent framework for “agents that run safely in the background” under OS-level control; one-click local-agent setup arrives for Hermes and OpenClaw on Windows, while Perplexity’s Portable Computer is already local on Linux RTX boxes.

Why it matters. Read this the way you’d read the Hugging Face acquisition that closed the day before: the distribution point of the open layer is consolidating on one vendor’s rails, and it just went mobile. The 128GB unified-memory figure is the operator-reading number — it decides whether a 27B-to-30B dense model runs entirely on-device with a real context window, or sends anything serious to the cloud — while “1 petaflop” is a marketing unit that says nothing about sustained tokens per second out of a thin chassis. Treat the laptop as a developer terminal, not a production box, and PAIR as the first credible attempt to make a fleet of local boxes behave like one GPU — the same pooling model a DGX-Spark owner already runs. What the Windows Agent framework actually gates is the thing to watch.

Source: blogs.nvidia.com, thurrott.com

The Rest

  • The Ban Artificial Superintelligence Act — Sanders and Casar introduced legislation to permanently ban superintelligent AI and pause advanced development until a federal regulator sets safety rules, citing the July agent-escape incident; a proposal, not a law, but the first text that turns “rogue agents” into statutory language. sanders.senate.gov
  • Crusoe raised ~$3B at ~$30B — the neocloud closed a round co-led by Atreides and Valor with Mubadala in, alongside a five-year $13 billion GPU-cloud contract with Jane Street; the capacity lane keeps printing money for anyone who can build a shed. reuters.com
  • CXMT is making HBM3E in small quantities — China’s memory maker is reportedly producing the AI memory at low volume with a 2027 expansion plan, as Alibaba’s T-Head and Cambricon test it; years behind SK Hynix and Samsung on HBM4, but the dependency-diversification story gets its first goods. reuters.com
  • San Jose coupled approvals to neighborhood payback — a Council committee voted to fold community benefits into the data-center standards and to study reinvesting a share of data-center tax revenue in the surrounding area; the engagement window just got a harder edge. kqed.org

What I’m watching

Four clocks. Grok 4.7 is still unshipped with Musk’s mid-September window the only countdown, and docs.x.ai still tops at grok-4.6. Meta’s Spark open weights are still “coming soon” — a founder promise against an artifact on the meta-models org. Whether Z.ai’s unit-intelligence-cost framing, as much marketing as measurement, survives the next open flash-model price war. And whether RTX Spark’s October launch delivers sustained local throughput or just a laptop-shaped benchmark card. Open Source AI Summit hits the Presidio on September 10-11.