← Back

Simon Willison

journalistCredibility: 83%

Why this score? Independent developer and widely cited LLM analyst; hands-on testing, transparent methodology, publishes corrections; not institutionally edited.

Tracked Statements (1)

Claude 3 Opus has self-reported benchmark scores that consistently beat GPT-4. This is a really big deal: in the 12+ months since the GPT-4 release no other model has consistently beat it in this way.

Context: Anthropic's published table did show consistent wins over GPT-4's reported scores, and the retrospective independent record supports the assessment: on Artificial Analysis's current Intelligence Index (v4.1), Claude 3 Opus (12, estimated) scores above GPT-4 (7, estimated).