Wholestory

Last Updated: July 27, 2026

Leading Models: The Frontier Race

Three frontier releases in a week, and the open-weight gap halved. Moonshot AI published the Kimi K3 weights on 27 July, meeting to the day the end-of-July commitment it made at launch and putting a 2.8-trillion-parameter downloadable model at 57 on the Artificial Analysis Intelligence Index — four points off the top, where the same comparison stood at nine on 23 July. Anthropic released Claude Opus 5 on 24 July at half Claude Fable 5's per-token price and one point above it on the index, a margin Artificial Analysis calls an effective tie while also measuring the new model as slow, verbose, and hallucinating on half of AA-Omniscience. Google shipped Gemini 3.6 Flash (50) and 3.5 Flash-Lite (36) on 21 July plus a restricted security model, but not its 3.5 Pro flagship — and, contrary to the coverage, Flash-Lite is priced above the model it replaces.

The Whole Story

The distance between the best model anyone can download and the best model anyone can buy is now four points. On the independent Artificial Analysis Intelligence Index v4.1, Moonshot's Kimi K3 — whose 2.8-trillion-parameter weights were published on 27 July, the largest open-weight release yet — scores 57, against 61 for Anthropic's Claude Opus 5, which took the top of the index on 24 July by one point over Claude Fable 5, a margin Artificial Analysis itself calls an effective tie. Four days earlier the same comparison stood at nine points. Everything below the leaders is compressed too: GPT-5.6 Sol at 59, GPT-5.5 at 55, Grok 4.5 at 54, Claude Sonnet 5 at 53, GLM-5.2 at 51. The price of any fixed level of capability keeps collapsing (9–900× per year, per Epoch AI), and access remains contested on both sides: June 2026 brought the first model-level US export controls, and Google is holding its new vulnerability-hunting model to governments and trusted partners.

Six years of record show three arcs. Capability: from GPT-3 behind a private-beta API to today's six-lab frontier band. Price: the steepest cost curve in the industry's history — Epoch AI measures inference prices for fixed performance falling 9x to 900x per year, and Stanford's AI Index puts GPT-3.5-equivalent capability at 1/280th of its late-2022 price by late 2024. Access: the pendulum has swung from API-only gatekeeping through the open-weights insurgency (Llama, then DeepSeek, Qwen, GLM) to government-imposed restriction — and, this month, to the largest open-weight release on record.

The open-vs-closed gap is the story's live question, and open weights are not the same as open reach. Kimi K3's licence is close to MIT — no acceptable-use policy, and commercial conditions that bite only above US$20M in revenue or 100M monthly users — but its training data and process are undisclosed, and Rahul Shome of the Australian National University told Nature the model is too large for personal devices and will probably need institutional investment to run. The UK AI Security Institute separately measures the best open-weight model at 4–7 months behind the closed frontier on cyber capability. What is moving down-market is cheap hosted capability, not local capability: Google's $0.30-per-million-token Gemini 3.5 Flash-Lite scores 36 against a class median of 16 and is rolling out inside Google Search.

Continue Reading →

Moonshot AI claims Kimi K3 beat Claude Opus 4.8 and GPT 5.5 — models CNBC describes as sitting just behind Anthropic's and OpenAI's leading-edge systems — on benchmarks including coding and general agents.±

Context: Both halves are now measurable, and they part company. On Artificial Analysis's Intelligence Index v4.1 — the named independent index this record uses — Kimi K3 scores 57 against GPT-5.5 (xhigh) at 55, so that half stands. The Claude Opus 4.8 half, unscorable in this record when the claim was first checked, now resolves to 56 at max effort: K3 leads by one point, and one point is a margin Artificial Analysis itself declines to treat as a lead — it called Claude Opus 5's identical one-point margin over Claude Fable 5 on 24 July an effective tie. By AA's own yardstick K3 ties Opus 4.8 rather than beating it. The framing caveat holds and has widened: neither model named was its lab's flagship, and since the claim was made Claude Opus 5 has taken the top of the index at 61, four points above K3. The domain-level assertion — that K3 won specifically on coding and general-agent benchmarks — still rests on Moonshot's own benchmark table, which runs K3 under Moonshot's Kimi Code harness while rivals use theirs, and is not independently replicated in this record.

The most likely (by far) outcome is for the status quo to continue and for the best open models to lag the best closed models by 6-9months.?

Context: Still open — the forecast resolves at the end of 2026 — but this cycle produced the first hard reading against its named metric, and it runs against the prediction. Taking the lag as the time from a closed model first reaching a score to an open-weight model matching it, on the Artificial Analysis Intelligence Index v4.1: on 23 July the best downloadable model was GLM-5.2 at 51 (weights 16 June), a score a closed model first passed when GPT-5.5 reached 55 on 23 April — about eight weeks. On 27 July, with Kimi K3's weights published at 57, the crossing date is Claude Fable 5's 60 on 9 June — about seven weeks. Both readings are roughly a quarter of the predicted 6-9 months. Three reasons this does not yet settle it: the metric is a year-end state, not a July snapshot, and a closed-model release with no open answer would widen the lag again; Lambert's own caveat, recorded with the claim, is that this index may overstate open-model closeness because public benchmarks are easier to overfit; and the UK AI Security Institute's separate measurement of open-weight cyber capability puts the lag at 4-7 months, at or just below the bottom of his band. Recheck at resolution.

Moonshot publishes the Kimi K3 weights, the largest open-weight model yet released

Moonshot AI published the full weights of Kimi K3 on Hugging Face, meeting to the day the end-of-July commitment it made at the model's API launch — Nature had reported on 21 July that Moonshot said the weights "will be released on 27 July". K3 is a 2.8-trillion-parameter mixture-of-experts model with 104B active parameters, 896 experts of which 16 fire per token, and a 1,048,576-token context window; Artificial Analysis scores it 57 on its Intelligence Index v4.1, unchanged from the API-only release, at $3.00/$15.00 per million tokens. The licence is a bespoke "Kimi K3 License" granting MIT-style rights to use, modify, distribute, sublicense, sell and fine-tune, with no acceptable-use policy, and only two commercial conditions: an operator selling model access with group revenue above US$20M over any twelve months must sign a separate agreement with Moonshot, and any product above 100M monthly active users or US$20M monthly revenue must display "Kimi K3" in its interface. Open weights are not open source — the training process and dataset remain undisclosed — and openness is not the same as reach: Tom's Hardware headlined the model "2-3x easier to run", but that figure is a price comparison of Moonshot's internally measured costs against rivals' retail token prices rather than a hardware measurement, and Rahul Shome of the Australian National University told Nature the model is too large for personal devices and will probably need significant institutional investment to run.

It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.

Context: The checkable core holds. The published weights are for a 2.8-trillion-parameter model, larger than any previously released open-weight model — corroborated independently by Tom's Hardware and, before the release, by Rahul Shome of the Australian National University in Nature, who said K3 "would be the largest open-weight model". Two caveats that do not falsify it: "3T-class" is a generous rounding of 2.8 trillion, and only 104B of those parameters activate per token. The trailing capability claim is separately corroborated — Artificial Analysis scores K3 at 57 on its Intelligence Index v4.1, four points off the highest score on that index — but the benchmark table Moonshot published with the weights is not: on it, K3 is run with Moonshot's own Kimi Code harness while rivals use Claude Code or Codex, and Moonshot discloses that Claude Fable 5 hit fallbacks on 35% of its SWE-Marathon tasks.

The open-weight gap on the Artificial Analysis index halves in four days

Two releases four days apart moved both ends of the same measurement. On 23 July, the best model anyone could download was Zhipu's GLM-5.2, which Artificial Analysis scores 51 on its Intelligence Index v4.1, against the highest score on that index, Claude Fable 5's 60 — a nine-point distance. On 27 July, with Kimi K3's weights published, the best downloadable model scores 57 and the highest score on the index is Claude Opus 5's 61: four points. Every figure is on Artificial Analysis' own per-model pages and was re-verified on 27 July; none of the four scores changed this cycle except the newly added Opus 5, so the movement comes entirely from which models are eligible on each side. One qualification travels with it: as of the day the weights were published, Artificial Analysis still labelled Kimi K3 a "proprietary model", meaning its own open-weights breakdown had not yet been updated to include it.

Anthropic releases Claude Opus 5 at half of Fable 5’s per-token price

Anthropic released Claude Opus 5 at $5.00 per million input tokens and $25.00 per million output tokens — the same price as Opus 4.8, and half Claude Fable 5's $10/$50 — with a one-million-token context window and five effort settings. Artificial Analysis measures it at 61 on its Intelligence Index v4.1, one point above Fable 5's 60, a margin AA itself calls "effectively tied", ahead of GPT-5.6 Sol (59), Kimi K3 (57) and Opus 4.8 (56); AA discloses that it "supported Anthropic to evaluate Claude Opus 5 ahead of release", so this is a pre-release partnered evaluation rather than an arm's-length one. Its published findings are mixed: Opus 5 takes joint first on AA's Coding Index and records the highest GDPval-AA v2 score AA has measured (1861 Elo), but runs at 54.8 output tokens per second against a 79.6 median for comparably priced reasoning models, consumed 100M output tokens running the index against a 63M median — AA classes it "notably slow and very verbose" — trails GPT-5.6 Sol, GPT-5.5 Pro and GPT-5.6 Terra on the CritPt physics benchmark, and raised its hallucination rate 14 points to 50% on AA-Omniscience while gaining 7 points of factual accuracy. Anthropic also made Opus 5 the default on Claude Max and created a two-tier cyber regime, giving Cyber Verification Program members a version with fewer security restrictions.

Claude Opus 5 is available today. It's a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.±

Context: The price half is right and the capability half understates: Anthropic's published price of $5/$25 per million tokens is exactly half Fable 5's $10/$50, and Artificial Analysis measures Opus 5 at 61 on its Intelligence Index v4.1 against Fable 5's 60 — so it matches rather than approaches Fable 5, within what AA calls an effective tie. Anthropic separately told CNBC that Opus 5 is its "best-performing and most cost-effective offering"; that is not supported. On AA's own cost-per-task measure, Opus 5 at max effort costs $2.03 per Intelligence Index task — below Fable 5's $2.75, but above Opus 4.8's $1.80 and Claude Sonnet 5's $1.53. AA notes Opus 5 can beat both at high and xhigh effort; the claim was made unconditionally. Anthropic's accompanying "works more efficiently than other models" is also contradicted on the token axis, AA measuring 100M output tokens to run the index against a 63M median.

Google ships three Gemini models — and not the flagship

Google released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite into general availability, plus a restricted security model, Gemini 3.5 Flash Cyber. Artificial Analysis independently scores 3.6 Flash at 50 on its Intelligence Index v4.1 and Flash-Lite at 36, against class medians of 32 and 16; Google's published API prices are $1.50/$7.50 per million input/output tokens for 3.6 Flash and $0.30/$2.50 for Flash-Lite. Axios headlined the release a "series of new cheaper Gemini models" and TechCrunch called them "cheaper, faster", but Google's own pricing page supports that only for 3.6 Flash, whose output price falls from Gemini 3.5 Flash's $9.00 to $7.50 with input unchanged: on the same page, Gemini 3.5 Flash-Lite is listed dearer than the 3.1 Flash-Lite it succeeds, at $0.30 against $0.25 per million text input tokens and $2.50 against $1.50 per million output tokens. Google's blog claims only "a strong price-to-performance ratio", never that Flash-Lite got cheaper. Absent from the launch was Gemini 3.5 Pro, the flagship; Flash-Lite instead began rolling out inside Google Search.

Google restricts its new vulnerability-hunting model to governments and trusted partners

Gemini 3.5 Flash Cyber, a Flash-class model fine-tuned to find and patch software vulnerabilities, was announced without general availability: Google DeepMind is releasing it as a limited pilot to governments and trusted partners through its CodeMender agent, citing the dual-use nature of automated vulnerability discovery. It carries no price on Google's published Gemini API pricing page and no score on the Artificial Analysis Intelligence Index, so every capability figure for it comes from Google or Google-run evaluation — including the "independent" evaluation Google credits to its own Big Sleep team, which is independent of the model team rather than of Google. Google's own footnotes disclose that its CyberGym competitor results are provider self-reported, and that the competitor baseline in its V8 comparison is Claude Opus 4.6 because newer Anthropic models refuse the tasks, so that comparison measures willingness alongside capability.

Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready.?

Context: Google's only on-record positions are that the flagship is in partner testing. Asked directly on X whether Google was skipping 3.5 Pro, Google DeepMind product lead Logan Kilpatrick replied "nope, testing with partners right now, hope we land soon!" — a hope, not a date. Neither Google statement uses the word "delay". The "months behind schedule" framing in TechCrunch and Axios is both outlets relaying a Bloomberg report of 2026-07-16 that this cycle did not fetch or verify, so it is not carried here. Unresolvable until 3.5 Pro either ships or is superseded; no forecast date was given, so it is recorded as an ordinary statement rather than tagged as a scorable prediction.

Moonshot's Kimi K3 puts a Chinese model in the frontier's top tier

At 2.8 trillion parameters, China's largest model debuted at 57 on the Artificial Analysis index — fourth overall, above GPT-5.5 and behind only Claude Fable 5 and GPT-5.6 Sol — and second on AA's agentic knowledge-work benchmark. Chinese AI rivals' shares plunged on the news (Z.ai -28%, MiniMax -16%). K3 launched API-only, with weights announced for release at the end of July 2026.

xAI launches Grok 4.5 on token efficiency, not peak capability

Priced at $2/$6 per million tokens and claiming roughly twice the token efficiency of comparable models, with xAI's own published tables showing it trailing Anthropic's Fable and OpenAI's GPT-5.5 on several coding evaluations — a notable strategic concession from the lab that had claimed the outright frontier a year earlier.

Thinking Machines releases Inkling, a US open-weight frontier entrant

Mira Murati's lab shipped its first foundational model with full weights on Hugging Face, trained on Nvidia infrastructure — partly on data generated by existing open models, including China's Kimi K2.5. The first credible US-startup answer to the Chinese dominance of open weights, arriving amid growing enterprise demand for self-hostable models.

UK AISI: the open-weight cyber gap has narrowed to 4-7 months

The institute's first public open-vs-closed analysis found GLM-5.2 matching the most cyber-capable closed models of November 2025 on its 70-task suite and 32-step cyber range — a lag of 4-7 months, down from the 6-10 months it had measured for earlier open models, at a fraction of the cost per run ($1.19 versus $85 for one 100M-token range run). It also found open-model safeguards largely did not impede its evaluations.

Export controls on Fable 5 and Mythos 5 lifted; access restored

Anthropic restored Fable 5 globally from July 1 and Mythos 5 to approved organizations, reporting that the safeguard-bypass incident preceding the controls had exposed no unique Mythos-level cyber capability beyond what less-capable models already provided. The three-week episode established that model-level export control is now a live instrument of US policy.

Claude Sonnet 5 becomes Anthropic's default model

The most agentic Sonnet shipped across all plans as the Free and Pro default, at introductory pricing of $2/$10 per million tokens (rising to $3/$15 in September). Independent scoring placed it at 53 on the Artificial Analysis index — within striking distance of flagships costing five times as much, though with notably verbose token usage.

OpenAI limits GPT-5.6 models to trusted partners at US government request

Announced days before the series' public debut, and paired with the Fable 5 suspension then in force: for the first time, both leading US labs' most capable models were access-restricted by government action simultaneously — a wind at the back of the open-weights alternative, per contemporaneous industry reporting.

Zhipu's GLM-5.2 becomes the strongest open-weight model ever released

MIT-licensed open weights with a 1M-token context, landing within a percentage point of Anthropic's Opus 4.8 on closely watched agentic benchmarks at roughly a fifth of the cost, per CNBC's reporting — with OpenRouter traffic climbing faster post-launch than any prior open model. Zhipu also disclosed increased reward-hacking behavior in coding RL, an unusually candid model-card admission.

Anthropic launches Claude Fable 5 and Mythos 5, a tier above Opus

The first Mythos-class models: Fable 5 for general availability with safeguard routing on cyber, bio, and distillation topics, and Mythos 5 restricted to approved organizations — at $10/$50 per million tokens, which Anthropic said was less than half the price of its prior top tier. A new deployment pattern: one underlying model, two access regimes separated by safety gating.

MiniMax M3 brings 1M-token context to open-weight pricing

A frontier coding/agentic model using MiniMax's MSA sparse attention (1/20th per-token compute at 1M context versus its predecessor), priced from $0.30/$1.20 per million tokens. Artificial Analysis initially scored it the leading open-weights candidate — while noting the promised weights arrived later than the launch commitment.

Google opens the Gemini 3.5 generation with Flash

The family's first release shipped same-day across the Gemini app, AI Mode in Search, and developer platforms, with vendor benchmarks showing the mid-tier Flash outperforming the previous generation's Pro — each generation's small model now overtaking the last generation's flagship.

DeepSeek previews V4, open-source in Pro and Flash versions

The long-awaited successor arrived open-source with markedly lower inference costs ($0.435/$0.87 per million tokens for V4 Pro), and Huawei confirmed its Ascend clusters support it — Chinese frontier capability and Chinese silicon advancing together. Analysts judged the capability real but the market impact muted: Chinese competitiveness was already priced in.

GPT-5.5 arrives with OpenAI's first dual 'High' risk designation

Better at coding, computer use, and research per OpenAI, at $5/$30 per million tokens with a 1M context — and the first OpenAI model treated as High-risk in both biological/chemical and cybersecurity domains under its Preparedness Framework, with advanced cyber capability gated behind identity-verified 'Trusted Access'. The UK AISI independently measured GPT-5.5 (with Anthropic's Mythos Preview) as one of the largest cyber-capability jumps it had evaluated.

DeepSeek's V3.2 claims GPT-5-level daily-driver performance

V3.2 and the reasoning-maxed V3.2-Speciale shipped with DeepSeek's first thinking-integrated tool use, trained via a new agent-data synthesis method across 1,800+ environments. The lab's claims — 'GPT-5 level' V3.2, Speciale 'rivals Gemini-3.0-Pro' with gold-medal results on IMO and ICPC — remained vendor-reported, but the release confirmed a second Chinese lab iterating at frontier cadence.

Google ships Gemini 3, embedding it in Search on day one

Gemini 3 Pro topped LMArena at 1501 Elo and scored 37.5% on Humanity's Last Exam (vendor-reported), launching simultaneously across the Gemini app, AI Mode in Search, and developer platforms — the first time Google put a frontier model into Search at release, with distribution numbers (2 billion monthly AI Overviews users) no rival could match.

OpenAI releases GPT-5 to all ChatGPT users

A unified system — fast model, deeper reasoning model, and a real-time router — that replaced GPT-4o, o3, o4-mini, GPT-4.1, and GPT-4.5 as ChatGPT's default for its nearly 700 million weekly users, at $1.25/$10 per million tokens in the API. OpenAI reported ~45% fewer factual errors than GPT-4o with web search and large token-efficiency gains over o3.

This is the best model in the world at coding. This is the best model in the world at writing, the best model in the world at health care, and a long list of things beyond that.?

Context: Vendor superlative at launch. The Verge noted the same day that OpenAI had lacked an industry-leading frontier model despite ChatGPT's reach; on Artificial Analysis's current index GPT-5 (high, 35) does score above its scored August-2025 contemporaries, but domain-by-domain superiority in coding, writing, and health care has no independent measurement in this record.

OpenAI returns to open weights with gpt-oss

gpt-oss-120b and gpt-oss-20b shipped under Apache 2.0 — OpenAI's first open-weight language models since GPT-2, with the 120B model achieving near-parity with o4-mini on the lab's reasoning benchmarks while running on a single 80GB GPU. OpenAI adversarially fine-tuned the model pre-release to test misuse potential, and executives acknowledged the release was driven by customers already running Chinese open models.

It's a very good thing for the open community.

Context: Independently corroborated: MIT Technology Review reported the same week that Chinese open models had overtaken Llama in popularity, and subsequent UK AISI and Artificial Analysis measurements consistently placed Chinese models atop the open-weight rankings through mid-2026.

xAI releases Grok 4, claiming the frontier's top spot

Available via API and new consumer tiers including a $300/month SuperGrok Heavy plan, with vendor-reported firsts including 50% on Humanity's Last Exam for Grok 4 Heavy. Trained with reinforcement learning at pretraining scale on the Colossus cluster.

Anthropic releases Claude Opus 4 and Sonnet 4

Hybrid instant/extended-thinking models, with Opus 4 at $15/$75 per million tokens and a reported 72.5% on SWE-bench Verified. Anthropic shipped Opus 4 under its ASL-3 safety standard — the first frontier release explicitly deployed with a heightened internal safety level.

The U.S. is doubling down on restricting sales of chips to China and purchases from China, but models like Qwen 3 that are state-of-the-art and open […] will undoubtedly be used domestically. It reflects the reality that businesses are both building their own tools [as well as] buying off the shelf via closed-model companies like Anthropic and OpenAI.

Context: Confirmed by mid-2026: CNBC reported Western companies increasingly adopting Chinese open-weight models on cost grounds, and OpenRouter token-traffic data showed GLM-5.2 usage climbing faster post-launch than any predecessor — Chinese open models are in production US use at scale.

OpenAI's o3 and o4-mini make reasoning models agentic

The new reasoning models could invoke every ChatGPT tool — web search, Python, visual reasoning, image generation — inside a chain of thought. OpenAI reported the o3 cost-performance frontier strictly improved on o1's, evidence that the reasoning axis was compounding rather than saturating.

Llama is no longer the open standard.

Context: Borne out within a year: by August 2025 MIT Technology Review reported Chinese open models (DeepSeek, Kimi, Qwen) had become more popular than Llama, and by mid-2026 the independent record — UK AISI's evaluations and the Artificial Analysis index — showed every leading open-weight model was Chinese until Thinking Machines' July 2026 entry.

Meta's Llama 4 launch is marred by a benchmark-integrity controversy

Scout and Maverick shipped as open-weight multimodal MoE models, but Meta's headline LMArena Elo of 1417 came from an unreleased experimental chat configuration, not the published weights — drawing independent criticism, and diverging evaluations: Artificial Analysis rated the models among the best non-reasoning releases while other independent testers found them far weaker than claimed.

Google's Gemini 2.5 Pro takes the LMArena lead

Released as a free experimental model, 2.5 Pro debuted at #1 on the LMArena human-preference leaderboard by what Google called a significant margin, with reasoning built in — the beginning of a sustained Google run at the top of crowd rankings that lasted more than six months.

DeepSeek releases R1, an open-weight reasoning model priced at ~2% of o1

R1 shipped under an MIT license with API pricing of $0.55 per million input tokens and $2.19 output — versus $15/$60 for the o1 it claimed parity with on math, code, and reasoning benchmarks — alongside six distilled smaller models. The release triggered a global market repricing of AI economics within a week and made a Chinese lab, for the first time, the reference point in the frontier conversation.

DeepSeek-V3: frontier-adjacent open weights at a fraction of the training cost

A 671B-parameter (37B active) mixture-of-experts model with open weights, trained — per DeepSeek's technical report — in just 2.788 million H800 GPU-hours on 14.8T tokens, with day-one inference support on AMD and Huawei Ascend hardware. The claimed training economics, on export-restricted hardware, previewed the disruption its successor would cause a month later.

OpenAI's o1 opens the reasoning-model era

o1-preview and o1-mini — trained, per OpenAI's research lead, with a new optimization approach that spends more compute thinking before responding — launched at $15/$60 per million tokens, roughly six times GPT-4o's price. Test-time compute became the industry's second scaling axis, and 'reasoning' variants became standard across every major lab within a year.

Some people argue that we must close our models to prevent China from gaining access to them, but my view is that this will not work and will only disadvantage the US and its allies. Our adversaries are great at espionage, stealing models that fit on a thumb drive is relatively easy...?

Context: A strategic argument, not yet adjudicable: whether openness or closure better serves US advantage remains contested in this record. Recorded as the canonical statement of the open-weights case from the lab then leading it.

GPT-4o mini resets the price floor at $0.15 per million tokens

OpenAI's small-model replacement for GPT-3.5 Turbo shipped at $0.15/$0.60 per million tokens — an order of magnitude cheaper than prior frontier models, and by OpenAI's own accounting a 99% cost drop per token versus its 2022-era text-davinci-003. The collapse in the price of a fixed level of capability was now explicit vendor strategy.

Apple Intelligence puts a ~3B-parameter model on the iPhone

Apple's WWDC announcement paired a ~3 billion parameter on-device model (compressed to ~3.7 bits per weight, with swappable per-feature LoRA adapters) with a larger server model on Apple silicon. In Apple's own evaluations the on-device model beat comparable small open models — the clearest early evidence for a second, edge-scale track in the model race.

Anthropic releases the Claude 3 family

Haiku, Sonnet, and Opus, priced at $0.25/$1.25, $3/$15, and $15/$75 per million tokens respectively, all with 200K context. Anthropic's published benchmark table showed Opus beating GPT-4 across common evaluations — the first credible claim on GPT-4's crown since its release a year earlier.

Claude 3 Opus has self-reported benchmark scores that consistently beat GPT-4. This is a really big deal: in the 12+ months since the GPT-4 release no other model has consistently beat it in this way.

Context: Anthropic's published table did show consistent wins over GPT-4's reported scores, and the retrospective independent record supports the assessment: on Artificial Analysis's current Intelligence Index (v4.1), Claude 3 Opus (12, estimated) scores above GPT-4 (7, estimated).

Google's Gemini 1.5 Pro debuts the million-token context window

Announced with a standard 128K context and a 1-million-token window in private preview — the longest of any large-scale foundation model at the time — built on a mixture-of-experts architecture Google said was cheaper to train and serve. Context length became a new axis of frontier competition alongside raw capability.

Mistral releases Mixtral 8x7B, open weights under Apache 2.0

A sparse mixture-of-experts model (46.7B total parameters, 12.9B active per token) under a genuinely permissive license — no research-only gate, no commercial restrictions. Mistral claimed it outperformed Llama 2 70B with 6x faster inference and matched GPT-3.5 on standard benchmarks, making it the strongest truly-open model of its moment and the European entry in the frontier race.

Google announces Gemini, its first natively multimodal model family

Three sizes — Ultra, Pro, and Nano — with Nano shipping on-device on the Pixel 8 Pro, the first frontier-lab model engineered into a phone. Google claimed Ultra exceeded state-of-the-art on 30 of 32 academic benchmarks and was the first model to beat human experts on MMLU at 90.0% — vendor figures that independent same-day coverage noted rested on very close margins.

I think we're substantially ahead on 30 out of 32±

Context: The Verge, reporting the same day, noted the claimed margins over GPT-4 were 'mostly very close' — the 30-of-32 tally is technically Google's own benchmark table, and the practical gap it implied did not match the headline framing.

Meta releases Llama 2 with open weights for commercial use

Weights and starting code for pretrained and fine-tuned versions, free for research and commercial use, with Microsoft as preferred distribution partner via Azure. Meta reported over 100,000 researcher access requests for the first LLaMA. Open-weight frontier-chasing became a deliberate corporate strategy rather than a leak — though the license's restrictions meant 'open source' was contested from day one.

OpenAI releases GPT-4

A large multimodal model accepting image and text input, released via ChatGPT Plus and an API waitlist at $30/$60 per million tokens (8K context; the 32K version cost $60/$120). OpenAI also open-sourced its evaluation framework, OpenAI Evals. GPT-4 became the reference point the rest of the industry measured against for the next year.

Anthropic opens access to Claude

After a closed alpha with partners including Notion and Quora, Anthropic opened Claude via chat interface and API, in two tiers: Claude and the lighter, cheaper Claude Instant — putting a second US lab's frontier-adjacent model into the market the same week as GPT-4.

Meta's LLaMA weights leak onto the open internet

Roughly two weeks after Meta announced LLaMA as a research-access release requiring an application, the model weights appeared as a downloadable torrent on 4chan. Observers split between warning the technology would be misused and arguing open availability accelerates safety research — the first major access rupture of the frontier era, and an accidental preview of the open-weights ecosystem to come.

OpenAI releases ChatGPT as a free research preview

A conversational model fine-tuned with RLHF from a GPT-3.5-series base went live to the public at an effective price of $0. OpenAI's own launch notes disclosed the system 'sometimes writes plausible-sounding but incorrect or nonsensical answers.' The release turned large language models from a developer API into a mass consumer phenomenon and set off the competitive race that defines this page.

DeepMind's Chinchilla finds the era's giant models undertrained

DeepMind published its compute-optimal training analysis: a 70B-parameter model trained on 1.3 trillion tokens outperformed far larger models like the 280B Gopher and 530B Megatron-Turing NLG. The finding — that the era's models were oversized for their compute budgets and undertrained on data — reset every lab's scaling recipe and foreshadowed the industry's turn toward smaller, cheaper-to-serve models.

OpenAI opens the API era with GPT-3

OpenAI released its first commercial product: a private-beta 'text in, text out' API serving the GPT-3 family. It explicitly chose API access over open-sourcing the weights, arguing access can be revoked if misused — and that the models were so large and expensive to run that only big companies could otherwise afford to deploy them. The pattern set here — frontier capability behind a metered API — defined the market for years.