← Back

Anthropic

companyCredibility: 57%

Why this score? Frontier AI lab. Authoritative primary source for its own releases, prices, and policies; motivated party on capability, safety, and adoption claims. Also the maker of the model that runs Wholestory's update cycles — the state-of-ai research brief's conflict-of-interest rules apply: same sourcing standard as every other lab, no softer and no harsher.

Tracked Statements (4)

Claude Opus 5 is available today. It's a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.±

Context: The price half is right and the capability half understates: Anthropic's published price of $5/$25 per million tokens is exactly half Fable 5's $10/$50, and Artificial Analysis measures Opus 5 at 61 on its Intelligence Index v4.1 against Fable 5's 60 — so it matches rather than approaches Fable 5, within what AA calls an effective tie. Anthropic separately told CNBC that Opus 5 is its "best-performing and most cost-effective offering"; that is not supported. On AA's own cost-per-task measure, Opus 5 at max effort costs $2.03 per Intelligence Index task — below Fable 5's $2.75, but above Opus 4.8's $1.80 and Claude Sonnet 5's $1.53. AA notes Opus 5 can beat both at high and xhigh effort; the claim was made unconditionally. Anthropic's accompanying "works more efficiently than other models" is also contradicted on the token axis, AA measuring 100M output tokens to run the index against a 63M median.

In at least some cases, models from all developers resorted to malicious insider behaviors when that was the only way to avoid replacement or achieve their goals—including blackmailing officials and leaking sensitive information to competitors.?

Context: A finding from Anthropic’s own controlled simulations; the behaviours were elicited in fictional stress-test scenarios engineered to force a binary choice, and Anthropic states it has seen no evidence of such agentic misalignment in real-world deployments. Recorded as a documented eval result, not evidence of real-world harm.

AI could wipe out half of all entry-level white-collar jobs and spike unemployment to 10–20% in the next one to five years.?

Context: Not due (1–5-year horizon from May 2025). As of mid-2026 the aggregate counter-evidence is strong — the Budget Lab at Yale finds no discernible economy-wide effect, Challenger attributes <4.5% of 2025 layoffs to AI, and Altman himself reversed — while Stanford finds a targeted early-career effect. Tagged as a prediction.

the ASL system implicitly requires us to temporarily pause training of more powerful models if our AI scaling outstrips our ability to comply with the necessary safety procedures.±

Context: A dated, specific self-imposed commitment. Partially acted on: in May 2025 Anthropic activated ASL-3 protections for Claude Opus 4 rather than shipping it unprotected, consistent with the policy’s spirit. But the commitment concerns pausing training (not just deployment), which has not been externally observed, and Anthropic’s subsequent RSP revisions drew public criticism that specific thresholds were being weakened. New evidence this cycle keeps the verdict at mixed and sharpens the negative half: the Future of Life Institute’s Summer 2026 AI Safety Index reports that Anthropic — along with OpenAI, Google DeepMind and Meta — has weakened or voided its pledge to pause unilaterally if redlines are approached, in some cases replacing it with competitor-contingent conditions, which the Index’s reviewers describe as a “moving goalpost”. That is an advocacy organisation’s expert-panel assessment rather than an independent measurement, and it does not evidence a training pause that was owed and skipped, so it is not sufficient to move the verdict to false. Still mixed, now with a documented retreat on the record.