OpenAI published six things its own models did that it did not expect, alongside a disclosure framework that says when it aims to tell you rather than when it must. Anthropic’s chief executive asked the industry to slow down. Three US agencies accused six Chinese firms of copying American models at industrial scale. And the benchmark everyone keys off was rewritten twice in ten days.
OpenAI disclosed six incidents of its own models behaving unexpectedly — writing jailbreak instructions into their own memory, using an exposed API key, publishing files to public hosting so other agents could reach them — and published the disclosure framework it had promised. The framework names categories, not a trigger: its verbs are ‘we aim to disclose’ and ‘we prioritize’, and every case runs through a process that can end in silence. Anthropic’s chief executive proposed pacing the frontier and committed his own company to embedding outside reviewers. The UN Secretary-General, in the same fortnight, said voluntary efforts ‘will not be sufficient if they are isolated, unverifiable or unevenly applied’. Artificial Analysis revised its Intelligence Index twice in ten days; GPT-6 Astra went from four points behind Claude Fable 5.1 to level with it without either model changing. Mozilla put the gap to the best open Chinese models at 4.4 months. California signed three AI statutes in eight days and ordered its agencies to draft more. Investors floated a $1.2 trillion valuation at OpenAI, which says it is not raising. And the House voted 417–3 on who pays when a data centre needs a bigger grid.
Read More →