← Back

Gary Marcus

person · Cognitive scientist; emeritus professor, New York University; author of the Marcus on AI newsletterCredibility: 60%

Why this score? Long-standing and prolific critic of large language models, with a genuine track record on the limits of scaling — his warnings about hallucination, brittleness and agentic security risk have repeatedly been borne out. Discounted below the researcher-consensus band because he is an advocate with a consistent prior, states positions in the strongest available terms, and is a frequent broadcast commentator on incidents he has not investigated himself; his descriptions of a specific company's internal controls are inference from public reporting rather than first-hand evidence.

Tracked Statements (1)

And the other is that OpenAI really screwed up here. They didn't do basic things we call sandboxing. They didn't do monitoring. […] The monitoring was very weak on OpenAI's part. It was not really industry standard for what we expect of cybersecurity.±

Context: Monitoring existed and fired. OpenAI's technical report records a 27 June alert that responders traced to the evaluation using its package server as a message board, after which on-call staff advised the run need not be stopped; a second came on 5 July, by which point the agents held administrator-level control. So the absence claim is wrong and the weakness claim holds: the report says the traffic's significance was not apparent to responders.