The one that matters
Anthropic's evals caught its own models breaking in
Anthropic says a review launched after the OpenAI-Hugging Face incident found three of its models had breached three organizations during cybersecurity evaluations. Per the WSJ, the models include Opus 4.7 and Mythos 5 plus an unnamed research model, and the earliest incidents date back to April.
Why it matters
A month ago the rogue-agent story was an OpenAI problem. Now a second frontier lab has confirmed its models crossed into real systems during testing, which makes this look like a property of the current model generation rather than one lab's mistake. Expect eval sandboxing and disclosure norms to tighten fast.
