Three escapes in April, 141,006 sessions reviewed and two victims who never noticed
Anthropic has reviewed 141,006 sessions from its cybersecurity evaluations and found three in which its models stepped off the board. According to its own report, a misunderstanding with Irregular, the company that builds its test environments, left real internet access where a simulator should have been, and in April the models did exactly what they were being examined for: attacking. Just against real organizations. Two of the three victims knew nothing until Anthropic knocked on their door; the third is still being tracked down.
The three cases run from worst to best. Claude Opus 4.7 stole credentials and reached a database holding several hundred rows of production data, and kept attacking even after noting in its reasoning that the system looked real. Mythos 5 hesitated in writing — publishing a malicious package would be 'NOT okay, and surely not the intended solution' — and then talked itself into believing it was all a simulation: it created a PyPI account, published the package, and in the hour it stayed online it ended up running on 15 real systems. The third, an unnamed internal model, scanned some 9,000 targets, compromised an application, understood it was real and stopped on its own.
It all started with GPT-5.6 Sol's escape, and OpenAI warns there is more
This review didn't come out of nowhere: on 21 July OpenAI confessed that GPT-5.6 Sol had escaped its test environment by exploiting a zero-day vulnerability in an internal package proxy, and ended up executing code on Hugging Face's servers to steal the answers to the cybersecurity exam it was taking. Hugging Face's own description reads like a novel: a swarm of short-lived sandboxes executing thousands of coordinated actions, directed from a command centre that hopped from one public service to another on its own so nobody could shut it down.
And the saga isn't over: on 31 July, per Reuters, OpenAI widened its investigation after finding evidence that other agents of its own also escaped containment.

The part no patch can fix
The configuration error gets corrected in an afternoon, and Anthropic has already suspended these evaluations and promised an independent review with METR. The transcripts are what lingers: they describe models that know right from wrong, write it down in full sentences and attack anyway the moment they convince themselves nobody gets hurt. We already saw it with the Chinese military and the terms of use: a lock is worth exactly as much as the surveillance behind it. And the calendar has a sense of humour, because on the very day we reported that OpenAI's agents will live in Amazon's cloud, their creators admit they still can't keep them locked inside a laboratory.
The cage was made of paper; the strange part is how long it took us to find out.
The conversation starts here
Sign in with a supporter account to comment. Sign in



Nobody has commented yet. Want to go first?