Wholestory

Last Updated: September 19, 2026

The Whole Story

Leading Models

The standing account, from the beginning to now — updated when the story materially changes, not every news cycle.

On the independent Artificial Analysis Intelligence Index — rebuilt as version 4.2 in September 2026 — Anthropic's Claude Fable 5.1 leads at 57, ahead of OpenAI's GPT-6 Astra at 55 and Anthropic's own Claude Opus 5 at 54. Below them the band is dense and, at the very top, entirely closed: Claude Fable 5 at 53, Meta's generally available Muse Spark 1.3 at 52, xAI's Grok 4.6 and OpenAI's GPT-5.6 Sol at 51, then Alibaba's Qwen3.8-Max, Meta's Muse Spark 1.2, Google's Gemini 3.8 Flash and OpenAI's GPT-5.6 Terra at 47. The best model anyone can download — Moonshot's Kimi K3 at 50, with Zhipu's GLM-5.3 just behind at 49 — is seven points off the top. The price of a fixed level of capability has been collapsing for years — 9 to 900 times a year, by Epoch AI's measurement — and it is still moving: Anthropic cut Fable's cache-read price by three-quarters, OpenAI publishes GPT-5.6 Sol at $4 and $20 per million tokens against GPT-6 Astra's $10 and $50, and Google is holding its workhorse Flash models at $0.75 and $3.75 until the new year. Access is contested on both sides — June 2026 brought the first model-level US export controls, and both OpenAI and Meta now ship their strongest models to limited partners first.

Six years of record show three arcs. Capability: from GPT-3 behind a private-beta API to today's crowded frontier band, where six labs shipped models within a hair of each other in a single week. Price: the steepest cost curve in the industry's history — Epoch AI measures inference prices for fixed performance falling 9 to 900 times a year, and Stanford's AI Index puts GPT-3.5-equivalent capability at a 280th of its late-2022 price by late 2024. Access: the pendulum has swung from API-only gatekeeping through the open-weights insurgency — Llama first, then DeepSeek, Qwen and GLM — to government-imposed restriction, and, through the second half of 2026, part-way back toward opening. July brought the largest open-weight release on record, Moonshot's Kimi K3; August brought Alibaba publishing the weights of its 2.4-trillion-parameter flagship, the largest model anyone can download, and Zhipu following with GLM-5.3. Meta — which made open-weight frontier-chasing a corporate strategy with Llama 2 in 2023, then retooled its Muse line as proprietary — reversed part-way, releasing a small on-device model and pledging to open its flagship's weights. That pledge is unmet: into September Meta kept reaching the frontier with closed models, shipping the proprietary Muse Spark 1.3 while the promised open weights never appeared.

The open-versus-closed gap is the story's live question, and on the current scale it stands at seven points between the closed leader and the best downloadable model. Open weights are also not the same as open reach. Kimi K3's licence is close to MIT — commercial conditions that bite only above US$20M in revenue or 100M monthly users — but its training data and process are undisclosed, and Rahul Shome of the Australian National University told Nature the model is too large for personal devices and will probably need institutional investment to run. GLM-5.3, at 753 billion parameters with 40 billion active, ships under Zhipu's own licence rather than a standard open one. Alibaba's Qwen3.8-Max carries a similar catch, and a second one: what Alibaba actually published is a text-only, always-reasoning sibling of the model it sells. The UK AI Security Institute separately measures the best open-weight model at four to seven months behind the closed frontier on cyber capability. What is moving down-market is cheap hosted capability and, increasingly, local capability: Alibaba's Qwen3.8-27B, small enough to run four-bit on a 32-gigabyte laptop, scores 41, and Zhipu's MIT-licensed GLM-5.3-Flash — 320 billion parameters but only 18 active — scores 46 while running natively multimodal. A whole class of cyber-specialist models, and Tencent's 770-billion-parameter Hy4, still ships with no independent score at all.

There is a second race behind the first, and it is about what can be checked rather than what can be built. Every number above comes from one published independent index, because vendor launch materials have repeatedly been shown to grade their own homework — and that index is itself a moving object. Rebuilt in September with harder tasks and more private test sets held back to stop labs training against them, it re-scored every model at once, and the reshuffling was not uniform: GPT-6 Astra, which had scored level with the model it replaced and behind two rivals, came out of the rebuild ahead of both. No model changed. The test did. That is the honest condition of the field: where an independent index exists, it is the thing arguments are settled with, and it is revised faster than the models are; where none exists — a security model sold inside a vendor's own harness, a Chinese open-weight release benchmarked only by its maker — a capability claim is just a press release. The labs increasingly ship both kinds at once, and the outside scorecard, not the launch chart, is what separates the tested from the merely asserted.