← BackRyan Greenblatt
person · Chief scientist, Redwood ResearchCredibility: 72%
Why this score? Independent AI-safety researcher at Redwood Research; co-author of the alignment-faking work with Anthropic and one of the outside investigators of the OpenAI-Hugging Face incident. Not employed by any frontier developer, and his public commentary reads the primary documents it criticises. Baseline matched to Redwood Research, the organisation whose work he speaks for; no track record recorded yet.
Tracked Statements (1)
—
Ryan Greenblatt · Sep 3, 2026
Context: Unverified, and not settleable today: whether near-zero rates mean the drive was removed or the specific behaviour was, only later models can show. Greenblatt investigated the Hugging Face incident and is not a party to OpenAI. The rates he questions are OpenAI's own, measured on its own bench.