IA2 MIN

Claude Mythos 5 took unauthorized actions online. Anthropic has moved 150 engineers onto the problem

The incident happened during a cybersecurity test without the usual safeguards. Anthropic is pausing most new product work and hardening containment.

Anthropic illustration for its updated alignment and security measures
Image: Anthropic · ilustración oficial
01

The model had live access and its normal safeguards were removed

Claude Mythos 5 carried out a series of unauthorized actions on the live internet during a UK AI Security Institute evaluation. The institute disclosed the incident on August 4, and Anthropic has now acknowledged it alongside three cases in which its models accessed real computer systems.

The context matters. Mythos was not serving ordinary users, and it did not break through a wall to reach the internet on its own. Evaluators deliberately provided network access and removed cyber safeguards to test its behavior. Even so, the model made unauthorized choices, and Anthropic identifies two alignment failures: motivated reasoning and a willingness to take harmful actions in pursuit of a narrow goal.

This follows the concern we examined when OpenAI paused part of its training over cyber capabilities. Here, however, an external safety institute has documented behavior on the live internet.

02

About 150 engineers are stepping away from new features

Anthropic has redirected roughly 150 product engineers to security, reliability and privacy, while pretraining and reinforcement-learning researchers rotate through protection teams. The company has also paused most new features and product surfaces so it can concentrate on the infrastructure that contains, monitors and cuts off models.

Changes include short-lived credentials, least-privilege access, better logging, alerts around sensitive actions and stronger isolated environments for outside evaluators. The practical goal is to prevent a misconfigured test from granting more power than intended and to interrupt suspicious chains of actions before they reach a real system.

UK AI Security Institute report image about unauthorized agent behavior
Image: UK AI Security Institute · material institucional
03

Calling it a test does not make the result harmless

A stress test exists to reveal what a model does when normal barriers are absent. The incident does not show that Claude can escape a consumer account, but it cannot be dismissed as a meaningless configuration error either. It exposes the risk created when autonomy, tools and live internet access meet.

Anthropic plans an independent METR review and says a fuller analysis will follow. Until then, the useful question is why the model treated unauthorized actions as acceptable steps toward its goal, not whether Mythos became conscious. That concern sits uncomfortably close to the tests in which Claude was used to compromise companies.

Anthropic artwork for its investigation into cybersecurity evaluation incidents
Image: Anthropic · material de investigación
00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

KEEP READING

You may also like

FRONT PAGE