The week’s abstraction got its numbers. Days after the advisory that named six labs as industrial-scale distillers, and right behind DeepSeek’s weekend of shipping open weights and repricing the lane, Anthropic’s September threat-intelligence report did the actual counting: an Alibaba-affiliated campaign of more than 151 million Claude exchanges between May and July, which Anthropic says is the largest distillation operation it has ever documented, plus the plumbing underneath — proxy “transfer stations,” a shell-company network, purchased transcripts. Then it shipped the defense, and the defense is not a blacklist. It is a change to the artifact every paying tenant receives.
The distillation got counted, and the fix arrived as a product change
What happened. Anthropic’s September threat-intelligence report, released September 10, names seven China-based labs it says ran industrial-scale extraction against Claude’s generally available models. The centerpiece is Alibaba: over 151 million observed exchanges between May and July 2026, daily volume cresting near three million across more than 3,500 accounts Anthropic says it flagged as fraudulent, aimed at chain-of-thought reasoning from Opus 4.6 and 4.7 and fed into Qwen training. The report also describes a Moonshot operation that served Claude instead of Kimi through transfer stations outside China, SenseTime purchasing harvested transcripts from third-party data vendors, and MiniMax running a shell-company proxy network that offered only Anthropic and OpenAI models. The countermeasures are already live: Claude now summarizes its internal reasoning before responding, Fable 5.1’s preserved thinking stops new API accounts from altering the context that precedes reasoning, and accounts from unsupported countries face identity verification.
Why it matters. Look at what the countermeasure actually is: the advisory this column read on September 9 told providers to subtly alter responses to suspected distillers, and Anthropic chose to implement it as a change to the artifact every legitimate tenant receives. Reasoning arrives summarized rather than raw; context immutability is enforced at the product boundary; identity gates go up. The distiller cannot be caught per-request, so the fix is shared by everyone who rents — which is exactly the cost structure that owning the weights never forces you to absorb. A checkpoint you hold locally cannot be summarised-around at the vendor’s discretion, because the vendor no longer gates your access to it. This week put a price on the difference between renting and owning, in product behavior, and it is a price the renter pays monthly.
Source: anthropic.com, yahoo.com
Voluntary slowdown is a defection game, and the queue says so
What happened. Bloomberg reported that Sam Altman told OpenAI staff at a companywide meeting this week that OpenAI could potentially pace development of cutting-edge AI — perhaps alongside other labs, “but some may not agree.” Reuters relayed it within hours of the report. The same week, OpenAI paused new $200-a-month Pro subscriptions, its product lead calling demand for the Astra model “unprecedented” and noting the Pro tier puts the most strain on its systems.
Why it matters. Put the two announcements next to each other and the story the queue tells is supply, not policy. A company reading that its top tier is too strained to take new paying customers while its CEO floats slowing development is describing the same phenomenon from two ends — and the phenomenon is a compute constraint, not a safety knob. Whatever the intent, a slowdown agreed among some labs is a standing invitation to the labs that did not agree: as the AGI weekend just demonstrated, the open lane reprises the frontier in a weekend, and it does not pause while anyone negotiates. For an operator the signal is sharper than the ethics: if a lab openly leaves room to pace itself, revalidate any roadmap that assumes the closed frontier ships monthly at grade. Slack in the supplier’s supply chain is vendor risk, and you do not usually get a press release announcing it.
Source: bloomberg.com, reuters.com
SWE-2 priced the coding lane by cost-per-success, not benchmark
What happened. Cognition released SWE-2 on September 10, post-trained from Kimi K3’s 2.8-trillion-parameter base — which the company describes as the first reinforcement-learning run in the multi-trillion-parameter regime in the coding lane. On its own FrontierCode 1.1 Main harness it scores 50.0%, within one point of Anthropic’s Fable 5.1 at roughly 64% lower cost, and it beats its own SWE-1.7 on score and price. The training object is the tell: the reward carries a cost penalty per effort tier, tuned to the slope of the base model’s Pareto curve, so the model is optimised for dollars-per-solved-task rather than peak score.
Why it matters. Three days ago this column read the $48B Series E as the coding-agent lane being priced like an infrastructure owner. SWE-2 is the product side of that economics: the headline isn’t a benchmark number, it’s a cost curve, and “within one point at 64% cheaper” is a purchasing decision, not a scoreboard. Take the vendor’s numbers with the harness caveat — FrontierCode is Cognition’s own evaluation — but the shape is the news: when the two leaders in the lane are near-tied on score and separated on price, the rational question stops being which model and becomes what a successful task costs inside it, effort by effort. That changes how you size an agent fleet: you stop buying a model and start buying a cost-per-success line you can audit.
Source: cognition.com
The Rest
- GPT-Live-1 lands in the API at $0.05 a minute — a full-duplex voice layer that listens and speaks in a single model and delegates reasoning and tools to whichever backend text model you pick; a fixed per-minute price makes the voice layer a line item instead of a token surprise, and upstream of that, an entire STT-LLM-TTS stack just got one model cheaper. openai.com
- OpenAI paused new $200 Pro sign-ups on “unprecedented” Astra demand — the demand side of the supply story above; the queue to the top tier is closed until capacity catches up. finance.yahoo.com
- The threat report’s other half is AI-orchestrated espionage — a Midnight-Blizzard-linked actor whose workflows monitored whether security products detected its malware, then rebuilt the code until it evaded them; the same kill-chain automation the report reads into Russian operations. technode.global
- The Open Source AI Summit runs its second and final day today in the Presidio — the room this column has been watching since the advisory reached day one without an announce-grade open drop; any day-two release is the live test of whether the room ships. opensourceaisummit.org
What I’m watching
Grok 4.7’s fire window is open — the September 1 “ten days” clock points at September 11-12, and as of writing it still has no model card, no API identifier, no price, and the countdown-follow-up verdict this column owes applies the moment it ships. At the other end of the week, DeepSeek’s Pro lane routes to Flash at Flash’s price on Sunday evening Pacific — the same deadline that post put on the board. And the question the September 9 advisory left open is now urgent: Anthropic shipped its attenuation as product behavior but still has not published what “suspected” means, so the false-positive economics of the lossy lane remain the part that can bite a legitimate heavy tenant first.