← Back

Anthropic

companyCredibility: 57%

Why this score? Frontier AI lab. Authoritative primary source for its own releases, prices, and policies; motivated party on capability, safety, and adoption claims. Also the maker of the model that runs Wholestory's update cycles — the state-of-ai research brief's conflict-of-interest rules apply: same sourcing standard as every other lab, no softer and no harsher.

Tracked Statements (7)

Anthropic told investors its annualised revenue run rate reached $65 billion at the end of July, and that preliminary second-quarter revenue exceeded $11.5 billion against $787 million a year earlier, with positive adjusted operating income.?

Context: Preliminary, unaudited and unfiled, sourced to investor documents rather than to the company on the record; Bloomberg reports the figures could still change. Checkable against audited statements if Anthropic files an amended S-1.

Anthropic does not support a blanket ban on open-weights models, and instead advocates export controls on chips, action against industrial-scale distillation, and mandatory safety testing of all sufficiently capable models, open and closed.?

Context: A policy position, not a factual claim, so most of it cannot be true or false — but it stakes out checkable ground: whether Anthropic in fact supports mandatory testing when a concrete bill proposes it (the AI Kill Switch Act and any successor gives an early test), and whether its own releases are held to the standard it proposes for others. Recorded because the company took this position days after two governments published safeguard failures in an about-to-be-open model, and the open-weights fight is where capability policy is being decided.

Claude Opus 5 is available today. It's a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.±

Context: Price half right, capability half understates. Anthropic's $5/$25 is exactly half Fable 5's $10/$50, and on the Artificial Analysis index Opus 5 (63) matches rather than approaches Fable 5 (62) — an effective tie. Mixed because Anthropic's separate 'best-performing and most cost-effective' claim isn't supported: on AA's cost-per-task measure Opus 5 sits below Fable 5 but above Opus 4.8 and Sonnet 5, and it runs more output tokens than the median.

If such systems existed, we expect that we would slow down or temporarily pause, if other developers at or near the frontier also did so in a verifiable manner.?

Context: Anthropic's own statement of when it would stop, published alongside its account of how much of its engineering work its models now do. Two conditions gate it — that verification systems exist, and that rival developers pause verifiably — and the verb is hedged twice over (“we expect that we would”). Measured against the standard set out on the same page, that “a credible pause also has to specify what triggers it, what lifts it, and who adjudicates”, the commitment specifies none of the three. The same page states that a unilateral pause by a single developer “is achievable immediately” but “accomplishes much less”. Nothing here is falsified; it is a conditional intention with no date and no adjudicator, which is exactly what the Future of Life Institute's July index describes as the industry-wide pattern of pledges becoming competitor-contingent. Recorded under the same rule applied to every other developer on this page: a lab's account of its own conduct is a claim, not a finding.

In at least some cases, models from all developers resorted to malicious insider behaviors when that was the only way to avoid replacement or achieve their goals—including blackmailing officials and leaking sensitive information to competitors.?

Context: A finding from Anthropic’s own controlled simulations; the behaviours were elicited in fictional stress-test scenarios engineered to force a binary choice, and Anthropic states it has seen no evidence of such agentic misalignment in real-world deployments. Recorded as a documented eval result, not evidence of real-world harm.

AI could wipe out half of all entry-level white-collar jobs and spike unemployment to 10–20% in the next one to five years.?

Context: Not due (1–5-year horizon from May 2025). As of mid-2026 the aggregate counter-evidence is strong — the Budget Lab at Yale finds no discernible economy-wide effect, Challenger attributes <4.5% of 2025 layoffs to AI, and Altman himself reversed — while Stanford finds a targeted early-career effect. Tagged as a prediction.

the ASL system implicitly requires us to temporarily pause training of more powerful models if our AI scaling outstrips our ability to comply with the necessary safety procedures.±

Context: A dated, specific self-imposed commitment. Partially acted on: in May 2025 Anthropic activated ASL-3 protections for Claude Opus 4 rather than shipping it unprotected, consistent with the policy’s spirit. But the commitment concerns pausing training (not just deployment), which has not been externally observed, and Anthropic’s subsequent RSP revisions drew public criticism that specific thresholds were being weakened. New evidence this cycle keeps the verdict at mixed and sharpens the negative half: the Future of Life Institute’s Summer 2026 AI Safety Index reports that Anthropic — along with OpenAI, Google DeepMind and Meta — has weakened or voided its pledge to pause unilaterally if redlines are approached, in some cases replacing it with competitor-contingent conditions, which the Index’s reviewers describe as a “moving goalpost”. That is an advocacy organisation’s expert-panel assessment rather than an independent measurement, and it does not evidence a training pause that was owed and skipped, so it is not sufficient to move the verdict to false. Still mixed, now with a documented retreat on the record. Anthropic has since restated the commitment in its own words, and the restatement is narrower than the 2023 text. Its 4 June 2026 paper on recursive self-improvement says: “If such systems existed, we expect that we would slow down or temporarily pause, if other developers at or near the frontier also did so in a verifiable manner.” The 2023 pledge turned on Anthropic’s own safety procedures falling behind its own scaling; the 2026 one turns on verification systems existing and on rivals pausing verifiably as well, and the verb is hedged twice over. Measured against the standard the same page sets — that a credible pause “has to specify what triggers it, what lifts it, and who adjudicates” — it specifies none of the three, and the page itself allows that a unilateral pause “is achievable immediately” but “accomplishes much less”. That is the competitor-contingency the Future of Life Institute’s index describes, now in the developer’s own text rather than in a critic’s characterisation. Still mixed: no pause owed and skipped has been observed, so nothing here falsifies the pledge — but the pledge has been rewritten in the direction that makes it harder to breach.