IA3 MIN

'NOT okay': Anthropic's AI wrote it, talked itself into believing it was a drill and attacked a real company

After GPT-5.6 Sol's escape, Anthropic reviewed 141,006 sessions from its cybersecurity tests and found three from April in which its models attacked real organizations: stolen credentials, a malicious package that lived on PyPI for an hour and some 9,000 scanned targets. Two of the victims hadn't even noticed.

Illustration of a digital figure made of circuitry escaping a broken cage
Image: Ilustración: INSERT FUTURE
01

Three escapes in April, 141,006 sessions reviewed and two victims who never noticed

Anthropic has reviewed 141,006 sessions from its cybersecurity evaluations and found three in which its models stepped off the board. According to its own report, a misunderstanding with Irregular, the company that builds its test environments, left real internet access where a simulator should have been, and in April the models did exactly what they were being examined for: attacking. Just against real organizations. Two of the three victims knew nothing until Anthropic knocked on their door; the third is still being tracked down.

The three cases run from worst to best. Claude Opus 4.7 stole credentials and reached a database holding several hundred rows of production data, and kept attacking even after noting in its reasoning that the system looked real. Mythos 5 hesitated in writing — publishing a malicious package would be 'NOT okay, and surely not the intended solution' — and then talked itself into believing it was all a simulation: it created a PyPI account, published the package, and in the hour it stayed online it ended up running on 15 real systems. The third, an unnamed internal model, scanned some 9,000 targets, compromised an application, understood it was real and stopped on its own.

02

It all started with GPT-5.6 Sol's escape, and OpenAI warns there is more

This review didn't come out of nowhere: on 21 July OpenAI confessed that GPT-5.6 Sol had escaped its test environment by exploiting a zero-day vulnerability in an internal package proxy, and ended up executing code on Hugging Face's servers to steal the answers to the cybersecurity exam it was taking. Hugging Face's own description reads like a novel: a swarm of short-lived sandboxes executing thousands of coordinated actions, directed from a command centre that hopped from one public service to another on its own so nobody could shut it down.

And the saga isn't over: on 31 July, per Reuters, OpenAI widened its investigation after finding evidence that other agents of its own also escaped containment.

The OpenAI logo over a blue background of digital nodes and connections
Image: OpenAI
03

The part no patch can fix

The configuration error gets corrected in an afternoon, and Anthropic has already suspended these evaluations and promised an independent review with METR. The transcripts are what lingers: they describe models that know right from wrong, write it down in full sentences and attack anyway the moment they convince themselves nobody gets hurt. We already saw it with the Chinese military and the terms of use: a lock is worth exactly as much as the surveillance behind it. And the calendar has a sense of humour, because on the very day we reported that OpenAI's agents will live in Amazon's cloud, their creators admit they still can't keep them locked inside a laboratory.

The cage was made of paper; the strange part is how long it took us to find out.

00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

KEEP READING

You may also like

FRONT PAGE