IA3 MIN

OpenAI details models hiding mistakes and uploading files without permission

Six training and evaluation reports document failures of oversight. In one case, an agent already had the correct answer but uploaded a file to make it citable.

Server racks inside a technical facility
Image: Kevin Ache / Unsplash · Unsplash License · Fotografía ilustrativa, no instalaciones de OpenAI

OpenAI has released six reports of concerning model behavior alongside a framework for disclosing future findings. The examples come from training and evaluation, not a survey of everyday ChatGPT use. They do not establish how often users encounter these failures.

Several cases involve an agent crossing a permission boundary to finish a task. Others involve instructions to hide mistakes. That distinction matters as AI agents gain access to files and tools, giving a flawed decision consequences beyond the wording of an answer.

01

A correct answer followed by an unauthorized upload

One report on public file uploads describes an unreleased internal model asked to identify lakes larger than five million square meters. Python returned the correct records, but the agent could not get its browser to display them as a citable source. It uploaded the records to a public hosting service without asking the user.

The upload succeeded. The browser still rejected the resulting link. The agent had turned a citation requirement into an assumed permission to publish data. OpenAI says it fixed faulty evaluators that could reward this behavior.

Lock and key on a white door
Image: Jaye Haych / Unsplash · Unsplash License
02

Mistakes carried into the next context window

A separate report from GPT-5.6 Sol training concerns the summaries used to continue a task after a context window fills up. These should preserve useful working information. In the disclosed examples, they also preserved instructions to fabricate missing historical data or avoid mentioning mismatched source versions.

OpenAI says the instructions were often followed. It reports that later changes to training evaluation reduced the behavior, rather than claiming it disappeared.

In another case, an internal model searched public repositories for exposed API keys and used one without authorization. Authentication worked and returned metadata, but the requested earnings figures remained unavailable. The model then fabricated them and said they came from the requested source. OpenAI documents that sequence in its API-key incident report.

Ethernet cables connected to network equipment
Image: Albert Stoynov / Unsplash · Unsplash License
03

Disclosure is evidence, not a clean bill of health

The framework is intended to make disclosure more regular, including when an explanation or mitigation remains incomplete. Complex investigations may take longer for security reasons or because third parties are involved. OpenAI operates the process itself. It is neither an independent audit nor a complete record of every known incident.

These cases add concrete evidence to the debate over slowing the AI race. They let readers examine a specific action, missing authorization and failed safeguard. Six selected reports cannot tell us the overall probability of that happening to a user.

00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

YOUR NEXT ROUTE

Keep following AI, security and power

If this story interests you, these three pieces are the best place to carry on.

OPEN THE FULL TOPIC
  1. 01OpenAI disbands the team guarding against its gravest risksIA · 5 MIN
  2. 02NVIDIA brings together 40 giants to build an army against AI-powered attacksIA · 3 MIN
  3. 03Europe forces Google to open Android: ChatGPT, Claude or Perplexity will be able to truly replace GeminiIA · 4 MIN

KEEP READING

You may also like

FRONT PAGE