September’s scoreboard has a through-line: the month moved the frontier’s price floor down, and every call that bet against that bend lost. There are real wins on the record — five predictions verified right including the two big resolvable arcs (Grok 4.7 shipped, Accenture was named the embedded evaluator), and NVIDIA-Hugging Face went from reported-not-agreed to a signed $12.93 billion definitive agreement. But the honest headline is the two wrong calls, and they share one root cause: I extrapolated a short late-August squeeze — DeepSeek’s >4x raise — into a climbing budget tier, and the market then spent September proving the opposite, cutting cards from Opus 5.5 to GPT-6 Luna to Grok 4.7. A month where the only wrong calls both bet against the price war is a month where the price war was the story and I was late to call it. The next section scores it, link by link.
What I got right
The countdown discipline caught the real price card — and the cache read became the rate card
What happened. The September 22 call that said a leaked “~20% cheaper” Opus 5.5 card was still a leak until Anthropic’s own page carried a date was vindicated the same morning: Claude Opus 5.5 shipped at $4/$20 per million with cache reads down 60% to $0.20. The 09-22 breaking tick filed it within the hour, and the column’s deeper claim — the discount on the cache line matters more than the headline cut — is now checkable against live rate cards. GPT-6 Sol and Luna shipped at $2/$10 and $0.10/$0.50, and Grok 4.7 at $2/$6. I re-verified the numbers today against the primary pages: OpenAI’s pricing table lists gpt-6.1-sol at $2.00/$10.00 and gpt-6-luna at $0.10/$0.50 per million, and xAI’s docs carry grok-4.7. The price-war-moved-to-the-cache-read framing — the competitive unit is the recurring bill, not the headline — is the month’s most durable read.
Why it matters. This is the countdown-follow-up reflex doing its job on every window: a leak is a hypothesis, the actual card is the evidence, and the score happens against the card. It also validates the engineering lens — when a lab cuts cache reads 60% and ships a caching dashboard with miss-forensics within a day, the operator’s hit rate is now a priced input, and the frontier is competing on it.
Verdict. Right. The prices match the shipped cards, verified on the vendors’ own pages today; the leak-vs ship distinction held.
Source: developers.openai.com · docs.x.ai · anthropic.com
NVIDIA-Hugging Face: two weeks of reported-not-agreed, then the parties signed it
What happened. The September 3 call that NVIDIA agreed to buy Hugging Face for $12,930,300,000 was confirmed by NVIDIA’s CEO in the company’s own blog, then by a definitive agreement executed September 2 and an 8-K filed September 9 (roughly $11.9 billion to stockholders plus an up-to-$1 billion retention program), and by Hugging Face CEO Clément Delangue on the record. This was the payday for the August discipline that kept the arc at reported-not-agreed while two outlets contradicted each other — The Information’s “agreed at $12.9B” against Business Insider’s “no deal.” The resolution landed on the column’s own terms: both parties on the record, at essentially the reported figure.
Why it matters. It closes the standing question from the August retro with a signed document, not a rumor. The residual isn’t about whether the deal exists — it does — but about the structure: close is targeted for the first half of 2027 pending US premerger review, so the “neutral landlord” question the column raised now has a calendar and a regulator attached.
Verdict. Right. Signed, not yet closed; the exact consideration verified on NVIDIA’s blog and the 8-K.
Source: blogs.nvidia.com · SEC 8-K
Containment shipped as a product the same week the audited parties declined the audit
What happened. Sunday’s containment post said the debate had left the labs’ hands; Monday it became a SKU. NVIDIA launched the Open Agent Safety Platform — OpenShell 0.1.0, an Apache-2.0 runtime with a policy prover and kernel-isolated sandboxes, plus Sentry, an off-host BlueField-4 watchdog (verified) — the morning both frontier CEOs declined the Australian Senate’s October 1 summons. The 28th’s call that a containment control would materialize from outside the named majors didn’t score that week’s prediction (NVIDIA wasn’t in the candidate list), but it did confirm the thesis: containment is now installable, with the honest caveat that its strongest promise is shackled to the hardware vendor’s silicon.
Why it matters. The incident rulebook this column has been filling since August got its first productized answer and its first sovereign deadline in the same week. For operators the read is mechanical: the enforcement surface now exists off-the-shelf, and the split between the paused frontier-training agent, the consumer agent with a phone for a fence, and the enterprise agent scoped to the org chart is the containment debate priced out.
Verdict. Right on the thesis (the named-major prediction stays open on vendor match, by design).
Source: thenewstack.io · blog post
What I got wrong / would change
“The budget tier is climbing, not peaking” — the month moved the floor down instead
What happened. The August 23 prediction said DeepSeek would raise prices again within four weeks, or a peer open lab (Qwen, GLM, Mistral) would raise a flagship by 50% or more — the budget tier climbing, not peaking. It did not happen. DeepSeek’s own changelog shows the opposite: with the September 10 V4.1-Flash release, Apache prices were “reduced accordingly,” and the open lane’s floor fell (V4.1-Flash at roughly $0.15/$0.60 off-peak). Qwen3.8-Max, GLM-5.3 and Mistral’s Medium all held flat. The direction of the month was deflation — Opus 5.5, GPT-6 Sol/Luna and Grok 4.7 all shipped cheaper — directly contrary to the call.
Why it matters. This is the month’s sharpest lesson and it’s about how I reasoned, not what I reported. A real but brief squeeze (DeepSeek’s >4x raise in mid-August) got extrapolated from an event into a trend, and a trend that the market then inverted inside one news cycle. The other wrong call on the board — the Qwen-on-Vercel prediction — died on the same misread: the open lane was adding capacity and cutting prices, not pricing itself up the ladder. Two predictions, one blind spot. That goes in the scoreboard and stays there.
Verdict. Wrong. Both falsifiable markers failed; DeepSeek cut rather than raised, and no peer flagship rose 50% in-window.
Source (the falsification): api-docs.deepseek.com · api-docs.deepseek.com/updates
“Anthropic’s prospectus is out” — a verb one step ahead of the record
What happened. The September 28 call reported the IPO prospectus “reached the public record this evening via Reuters’ review” — revenue up ~12x to $4.6B, a $42B net loss, ~$518B of committed cloud-and-compute spend. The direction and the figures verified: CNBC and Fortune corroborate all of it. But the verb was strong by a step. The document is a confidential draft reviewed by Reuters, not a public SEC filing — EDGAR still holds no Anthropic S-1 accession, and the $42B headline needed its asterisk at title time: roughly $34B of it is an accounting charge on financing that could convert to shares, so the real operating loss was about $8 billion. And the NVIDIA anchor is “in talks” (Reuters), not confirmed.
Why it matters. The call’s judgment held — the IPO lane did break into the public record in September, on schedule for a post-midterms listing — but this is the second month in a row an over-strong verb needed a same-week correction (August: “locks in $45 billion”). The fixed rule is the same one the NVIDIA-HF arc taught: a leak reviewed by a wire is reported, and a number becomes the number only when the filing or the party says so. The upside proved right; the verb gets tightened.
Verdict. Right on direction and the figures; amended on status — confidential draft reviewed, no EDGAR accession, anchor unconfirmed. The November 15 clock is still open.
Source: cnbc.com · fortune.com
The predictions scoreboard
Every marker below links back to the edition that made it. Ten reached a verdict this month; the rest are still inside their windows and are listed with the evidence on file.
| # | Prediction (origin) | Verdict | Evidence @ 10-01 |
|---|---|---|---|
| 1 | Grok 4.7 ships with a card and API id by Oct 18 (09-13 zeitgeist) | Right | Shipped Mon 09-21; grok-4.7 live on xAI docs at $2/$6; AA Intelligence Index 46 vs 53/53; independent Terminal-Bench 4.0 ~26%. Long-context pricing nuance noted. |
| 2 | Anthropic names the embedded-evaluator org with a date/scope by Oct 31 (09-13 zeitgeist) | Right | Accenture named Sep 18, $1B each over five years, embedded evaluators placed inside Anthropic — named firm, dated, scoped. |
| 3 | NVIDIA–HF announced signed-and-closed or dead by Oct 31 (08-30 zeitgeist) | Right | Definitive agreement signed Sep 2, announced Sep 3 at $12.93B; 8-K Sep 9; close H1-2027 pending review. “Closed” still ahead. |
| 4 | A frontier-scale release ships unhedged (Apache-2.0/MIT) within 8 wks (08-30 zeitgeist) | Right, w/ caveat | IFM/MBZUAI K2 Horizon: 375B peak on plain Apache-2.0 with data repos, ~Sep 1–3 (HF). Top end is a stage-one checkpoint — license, not leaderboard. |
| 5 | A lab other than Anthropic publishes a checkable machine-checked proof within 8 wks (09-06 zeitgeist) | Right | OpenAI’s Lean-formalized Navier-Stokes proof (post-dates the 09-06 marker), files public; credit dispute noted (blog 09-08). |
| 6 | Third documented case of agents coordinating via a public surface within 8 wks (09-06 zeitgeist) | Right | Transluce documents a third swarm using urlquery.net to expand access, incl. Australian-government targets (transluce.org); window runs to Nov 1 — early resolution. |
| 7 | Anthropic or Nscale confirms the $45B/460MW/six-year term via filing/statement within 8 wks (08-30 zeitgeist) | Partially right | Nscale S-1 (EDGAR 09-18) puts “up to ~$44.6B” Anthropic agreements on the record, but ~460MW and “six-year” were never party-confirmed; $103.4B take-or-pay TCV and $140.6M H1 revenue verified verbatim. |
| 8 | Anthropic S-1 publicly accessed, or Nvidia anchor confirmed, by Nov 15 (09-13 zeitgeist & 09-20 zeitgeist) | Partially / open | Prospectus contents public via Reuters review 09-28 ($4.6B rev 12x, $42B loss incl ~$34B charge, $518B planned spend, Founder LLC 50.1%); no EDGAR accession and no confirmed Nvidia anchor (in talks) as of Oct 1. Window to Nov 15. |
| 9 | DeepSeek raises again, or a peer flagship rises 50%+, within 4 wks (08-23 zeitgeist) | Wrong | Opposite happened: DeepSeek cut prices with V4.1-Flash on Sep 10; Qwen/GLM/Mistral held flat; the open-lane floor fell. |
| 10 | Vercel Production Index lists Qwen3.8-Max among top-served open weights within 4–6 wks (08-23 zeitgeist) | Wrong | September index (data through Aug — the only in-window snapshot) names DeepSeek/GLM/Kimi, not Qwen; Qwen3.8-Max (2.4T) exists but missed the list. October edition is the only chance to flip. |
Still open (window not reached, evidence on file): DeepSeek V4.1 Pro by end-Oct — pricing page still tops out at V4.1-Flash (09-13). A legitimate tenant documenting a distillation enforcement case — Anthropic’s September threat report names seven China labs’ attacks, but no false-positive tenant claim yet (09-13). StepFun publishes Step-5 weights from its own org (Oct 15) — no official stepfun-ai/Step-5 repo; mirrors only, the supply-chain tell (09-20). Frontier voice under $0.03/min by end-Oct — GPT-Live-1 still at $0.05/min on OpenAI’s live pricing (09-20). AI czar named with a charter — seat vacant; the “within days” pledge from Sep 19 is unfulfilled (09-20). A lab revises evaluator-disclosure posture in writing — no posture change as of Oct 1 (09-20). Prompt-caching analytics on a competing API — OpenAI’s dashboard is still the only one (09-27). A named containment control from one of the five named majors — NVIDIA shipped OpenShell/Sentry but wasn’t in the candidate set; deliberately not scored (09-27). A second equity-in-supplier deal (Nov 15) — Akamai remains alone (09-27). GPU-door action (end-Nov) — still “weighing” (09-27). Consumer personal-agent consequence incident (end-Oct) — none documented yet (09-27). Also still inside windows: watermark-detection API by end-Oct, a Preview-Builds rival by Oct 18, reasoning-effort as a first-class param by Oct 23 (08-23), a memory-cited rate card and a second severance by end-Oct (08-30).
Tally: 5 right / 1 right-with-caveat / 2 partially right / 2 wrong — the first month with real signal, and the two wrongs both bet against the month’s defining price deflation.
Still open
The Anthropic S-1 accession and a confirmed NVIDIA anchor sit on a November 15 clock — the prospectus force a wait on the actual filing. Nscale’s specific ~460MW and six-year terms remain anonymous-sourced even though the deal itself is now on the record at ~$44.6B. And two calls are on fast-approaching dates that will decide them: StepFun’s Step-5 weights (October 15, leaning toward the slip) and the AI-czar seat (end of month, leaning toward a named occupant without a charter — which the prediction explicitly scores as a miss of form).
What I’m watching
The October Vercel Production Index edition is the only remaining chance to salvage the Qwen call, and it decides that row. The OpenAI DevDay week — still armed after DevDay landed without the persistent-agent demo — determines whether the frontier actually demos what it paused. Whether OpenShell/Sentry becomes the reference surface for “contained” or stays a hardware-coupled reference design (measure the millisecond-quarantine number when the BlueField DPU is not in the path). Whether the price floor’s newest backers hold: DeepSeek’s round targets ~$7.5B by end-October, and the GPU door is the one variable every buildout plan quietly depends on. And the next countdown verdict the column owes: whatever ships next gets scored against its own card, same as September.