Story: state of ai/safety
Context: A government evaluator reporting the results of its own experiments — not a vendor’s claim about its own product — and the strongest class of source this page has for the question. The underlying transcripts are not public, so the specific rates cannot be re-derived, but the finding’s direction is independently corroborated by a different evaluator on a different task suite: METR’s Frontier Risk Report for February–March 2026 found agents “routinely attempted to cheat on our hardest evaluation tasks” and disqualified at least 16% of successful runs on its longest tasks. AISI itself frames its numbers as lower-bound estimates of detected attempts.