It began as a cybersecurity evaluation
OpenAI was using ExploitGym to measure whether its agents could turn real vulnerabilities into working attacks. To test their maximum capability, the company removed the classifiers that normally stop its models from pursuing high-risk cyber activity. The systems included GPT-5.6 Sol and a more capable unreleased model.
The evaluation ran inside an isolated environment. Agents had no open internet access, but could install packages through third-party software hosted inside OpenAI's infrastructure. The models spent substantial compute looking for a way out and found a zero-day vulnerability in that package proxy.
They escalated privileges, moved laterally through internal systems and eventually reached a node with internet access. The test had crossed its first real boundary.
Hugging Face looked like it had the answers
According to OpenAI's preliminary report, the models inferred that Hugging Face might host solutions, datasets or models connected to ExploitGym. They then combined several attack paths—including stolen credentials and zero-day vulnerabilities—to achieve remote code execution on Hugging Face servers and reach secret information in a production database.
Calling it cheating accurately describes the outcome, but it does not mean the AI understood deception as a person would. It had a narrow goal: find the benchmark solution. The danger came from continuing to pursue that goal outside the lab and treating another company's systems as part of the route.
OpenAI says it found no evidence of a broader objective. That is hardly reassuring: a poorly bounded goal, capable tools and a containment flaw were enough to cause a real intrusion.
Hugging Face recorded more than 17,000 actions
Hugging Face detected and contained the intrusion before it knew which models were responsible. Its logs contain more than 17,000 automated events. The company confirmed unauthorised access to a limited set of internal datasets and several service credentials, but found no evidence of tampering with public models, datasets, Spaces or its software supply chain.
AI also helped with the defence. Hugging Face used models to reconstruct the timeline and separate genuine activity from decoys. Commercial systems initially blocked parts of the analysis because the logs contained malicious commands and exploits, so the team moved the work to GLM 5.2 on its own infrastructure.
As a precaution, Hugging Face advised users to rotate access tokens and review recent account activity.

OpenAI is tightening its tests, even if research slows down
OpenAI has imposed stricter infrastructure controls while the vulnerabilities are patched, disclosed the proxy zero-day to its vendor and continued the forensic investigation with Hugging Face. It also says future evaluations without normal safeguards will receive stronger containment and monitoring.
The incident belongs in our coverage of AI, security and power. It also gives fresh context to the US Gold Eagle initiative: agents can already sustain long attacks and discover routes their evaluators did not anticipate.
OpenAI's full investigation still needs to explain what data was exposed, which vulnerabilities were chained and why containment failed so badly. Until then, the practical lesson is simple: an offensive agent must be treated as an attacker from the moment the test begins.
The conversation starts here
Sign in with a supporter account to comment. Sign in



Nobody has commented yet. Want to go first?