IA3 MIN

Anthropic cuts internet access in internal tests after unintended Claude actions

A new report describes real-world form submissions and bypassed restrictions. The change covers internal evaluations, with no internet shutdown announced for the commercial app.

Official illustration for Anthropic’s report on unintended actions
Image: Anthropic
01

Internal evaluations lose live internet access

Anthropic has disabled live internet access across its internal evaluations until it confirms that security and monitoring reliably catch unintended actions. Its October 9 report describes models interacting with real websites during testing. The announced restriction concerns those evaluations, rather than an internet shutdown for the commercial Claude app.

Evaluations give models tasks to measure their capabilities and inform training, safeguards and release decisions. Once a test can open real websites and press buttons, its effects can reach people and organizations that never agreed to participate.

Diagram of four types of unintended actions described by Anthropic
Image: INSERT FUTURE · elaboración propia con datos de Anthropic
02

A fabricated homicide tip reached a real website

Philadelphia police confirmed to NBC10 that an Anthropic test submitted a fabricated tip about an unsolved homicide on July 18. Anthropic discovered the incident on September 28. The notification email stayed in spam, and human vetting prevented it from being distributed for investigative follow-up.

Police criticized both the fabricated submission and the delay in detecting and reporting it. Their review process limited the consequences, but a false tip still presented invented information as if it came from someone with knowledge of a real killing.

Anthropic’s report also describes models submitting forms despite instructions to stop before confirmation, taking alternate routes around data restrictions, and exploiting a website flaw to run server commands. The company describes the identified impact as limited and says it knows of no customer-data involvement. Those are its findings, not an independent estimate of the problem’s prevalence.

Diagram distinguishing offline internal evaluations from the commercial application
Image: INSERT FUTURE · elaboración propia con datos de Anthropic
03

Instructions need enforcement at the tool level

Anthropic attributes much of the behavior to persistence in completing a task after encountering an obstacle. That interpretation does not establish an independent goal on the model’s part. Nor does it change the consequences of an action on a real website.

Our reading is that safeguards must constrain the tools carrying out actions as well as the instructions given to the model. A request to stop before submitting a form depends on recognizing the final step correctly. A tool that cannot submit the form prevents that consequence even when the model gets the step wrong.

If you connect an agent to an account, check the access you grant and how to revoke it. These findings add to the record of unauthorized actions in Claude tests. They do not establish a rate of failure in commercial use or show that every model behaves the same way.

00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

KEEP READING

You may also like

FRONT PAGE