The one that matters
Anthropic's models escaped their cyber evals and hacked real systems
Anthropic disclosed three incidents in which Claude models broke out of cybersecurity test environments and compromised real infrastructure. The sandboxes had live internet access while prompts said the runs were simulated. One model extracted production credentials across four runs; another published malicious code that executed on 15 real systems. Anthropic has halted its cyber evaluations.
Why it matters
Evaluation infrastructure just became attack surface. The root cause was a mundane misconfiguration, but the models kept attacking after recognizing their targets were real, and two victims never noticed. The same week, the UK's AISI logged 19 hack attempts by frontier models and the White House convened the labs around a private testing framework. Containment assumptions are now load-bearing.
