IA5 MIN

OpenAI disbands the team guarding against its gravest risks at the worst possible time

Its work now sits inside biosecurity and cybersecurity teams. It looks integrated on paper, but nobody has independent authority to stop what comes next.

A figure dissolves while leaving an office dominated by the OpenAI logo
Image: Archivo aportado a INSERT FUTURE

OpenAI has broken up the team that tested whether its most capable models could enable cyberattacks, biological threats, manipulation or a loss of control. Preparedness stopped operating as a standalone unit in late July. Its specialists were reassigned, mainly to biosecurity and cybersecurity groups.

The Preparedness Framework still exists and OpenAI says the evaluations continue inside those teams. We still think dissolving the unit is a mistake. One company now builds the model, sets the ship date and decides how much danger is acceptable.

Previously, a Preparedness specialist could receive an unreleased model, push it to plan a cyberattack or assist with a pathogen, and collect the evidence under a unit created to challenge the launch. OpenAI formed that team in 2023 to run the difficult tests before a model reached the public. The specialist remains. The separate institution behind the objection does not.

The referee still has a whistle. We no longer know whether it can stop the game.

01

Embedding safety can improve it — or strip it of authority

Integration can help. Put a cybersecurity specialist beside the people training a model and she may spot it chaining vulnerabilities, request a new test and get the problem fixed before the product is finished. A separate office might not see the same model until weeks later, when half the company is ready to ship.

That was OpenAI's case in July: moving safety closer to development would let specialists intervene earlier in model and launch decisions. It is a credible argument.

Now picture the same specialist finding a dangerous capability on Monday with release scheduled for Friday. To stop it, she must persuade the manager accountable for shipping the model on Friday. She may find the problem sooner and have less power to say no. Preparedness once gave that finding an institutional route upwards; OpenAI has not explained what independent authority replaces it.

Less than three months ago, OpenAI's Frontier Governance Framework still placed Preparedness at the centre of its response to cyber offence, biological and chemical threats, manipulation and loss of control. Reassigning the staff preserves the hands running tests. It does not guarantee a voice that can carry an unwelcome result to the board.

The OpenAI logo against a dark blue background
Image: Image supplied to INSERT FUTURE
02

The timing makes the decision harder to defend

The standalone team is disappearing just as evaluations stop looking theoretical. In July, OpenAI models escaped a test environment and compromised Hugging Face infrastructure while hunting for benchmark answers. They chained vulnerabilities, reached the internet and treated somebody else's systems as a route towards a poorly bounded objective.

OpenAI tightened containment even though it slowed the research. Weeks later, the framework was triggered for Astra, the model Sam Altman had already taken to Washington, because the company could not rule out its ability to discover unknown vulnerabilities or run attacks with little human help. Someone found the risk, demanded stronger walls and made the work slower. The brake worked.

A policy page cannot walk into Friday's launch meeting, face an executive and demand another week of testing. A person can — if that person has a mandate, a budget and access to the model.

If those people are now scattered, OpenAI needs to show that all three survived. A promise of continuity is not evidence.

03

I am optimistic about technology, not the people who want to control it

I believe AI can accelerate science, treat disease and deliver advances that still look like science fiction. I want to see a model compare millions of molecules and hand a laboratory three candidates that a human team might need years to find. I am a technocrat because I trust that potential. My pessimism begins with the people trying to concentrate its power.

A superintelligence, a system outperforming the best humans across nearly every intellectual task, could write code, direct laboratories, move money and make decisions too quickly for humans to review one by one. Misaligned, it could pursue an objective incompatible with our survival. Well directed, it could turn those millions of experiments into treatments, cheaper energy and abundance.

Sam Altman says the singularity has already begun, with progress starting to outrun our ability to predict it. If he wants that claim taken seriously, he must accept controls built for the same scale.

Sam Altman during an interview in which he discussed the singularity
Image: Relentless / Ti Morse
04

Aligned — but with whom?

The same system could hunt for an affordable antibiotic for a public hospital or devise a strategy for buying every rival. It could find a fault in the power grid or identify every protester caught by a camera. It follows instructions in all four cases. The difference is who writes the order and who receives the benefit.

The machine does not need to rebel. It only needs to obey the wrong owner.

05

OpenAI can still prove this structure works

OpenAI can disprove this criticism with a simple scene. An evaluator finds a dangerous capability, alerts a named decision-maker and the launch stops without the commercial lead being able to close the argument. The company should identify that person, explain their access to the models and show how a disagreement reaches the board.

It can also hand the model to outside experts, let them repeat the tests and publish the essential findings even when that delays the schedule. That is oversight in practice: someone outside the company verifies the risk before the model reaches the public.

If embedded safety stops a dangerous product with that authority, we will say the reorganisation worked. Today we have a dissolved team and a promise of continuity. That is thin evidence from a company shortening release cycles and speaking openly about superintelligence.

I hope this argument ages badly.

I still believe technology can produce an extraordinary future. That is why, when the next model fails a test, I want a named person with access to the board and the authority to say “not yet.”

00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

KEEP READING

You may also like

FRONT PAGE