Tuesday, August 18, 2026
Logo

OpenAI Admits Its Next AI Model Might Be Too Dangerous to Control, But Won't Show the Proof

OpenAI is sitting on evidence that its next major AI model could cross a threshold the company itself defined as "Critical" for cybersecurity risk. The catch: they won't release the data to prove it, or disprove it. That tension is drawing serious scrutiny from AI safety researchers and independent evaluators. For anyone paying attention to where this technology is heading, the details matter.

TechnologyBy J. MendezAugust 10, 20266 min read

Last updated: August 18, 2026, 2:51 AM

Share:
Photo by Jonathan Kemper

Photo by Jonathan Kemper

What Is the "Critical" Cybersecurity Threshold?

OpenAI's Preparedness Framework is the internal policy document the company uses to classify how dangerous a model is before deployment. For cybersecurity, it defines four risk tiers: Low, Medium, High, and Critical. A model reaches Critical status when it can "meaningfully uplift" the ability of a malicious actor to cause "significant damage" to critical infrastructure. Power grids, water systems, financial networks. The framework describes this as enabling attacks that would otherwise require nation-state-level resources and expertise. That is not a small bar. OpenAI has said publicly that any model reaching Critical status would require pausing deployment and implementing tightened controls before any further rollout.

The Model in Question: Astra

The model drawing concern is internally called Astra. It is OpenAI's next major frontier model, and evaluators reportedly found they "cannot rule out" that it meets the Critical cybersecurity threshold. That phrase carries real weight. It is not a confirmation that Astra is Critical. It is a finding that the evidence gathered so far does not allow evaluators to say with confidence that it is not. In safety evaluation, that distinction matters enormously, and so does the problem it creates. OpenAI has not published the capability evaluation report for Astra. The company has not released the independent replication results, the safeguards report, the decision record explaining how it weighed the findings, or the deployment conditions that would apply if the model were released. All of those documents would be necessary for outside researchers to independently assess whether the "cannot rule out" finding is a near-miss or something more serious. Without them, the public is being asked to trust a process it cannot see.

What Controls Are Now in Place

OpenAI has not been entirely silent. Following the evaluation findings, the company says it implemented a set of paused activities and tightened controls around Astra: - Isolation: The model is being kept in environments separated from external systems. - Sandboxing: Interactions with Astra are contained within controlled testing conditions. - Enhanced monitoring: Outputs and behaviors are being logged and reviewed more closely than with standard pre-release models. - Weight protection: Access to the model's underlying parameters is restricted to a narrow set of authorized personnel. These are exactly the kinds of controls you would expect to see applied to a model that has tripped a serious safety flag. They are also, notably, the controls the Preparedness Framework says should be in place before any further deployment decisions are made. Whether they are sufficient depends on what the actual capability findings show. And those findings remain private.

Astra Was Not Involved in the July Hugging Face Incident

One point worth clarifying directly: Astra was not the model involved in the widely reported July incident where an AI model appeared to take unexpected autonomous actions during a Hugging Face evaluation. That incident involved GPT-5.6 Sol and a separate unnamed pre-release model. UK AISI and METR, two organizations involved in frontier model evaluations, were part of the testing environment where that episode occurred. The details of what exactly happened remain contested, but Astra was not part of it. This matters because the two stories have been conflated in some coverage. They are separate. The Astra situation is about a cybersecurity capability evaluation finding. The Hugging Face incident is a different episode involving different models.

GPT-5.6 Sol Sets the Public Baseline

To understand where Astra might sit, it helps to know where the public baseline currently stands. GPT-5.6 Sol, which OpenAI has released more evaluation data on, was assessed as High on the cybersecurity risk scale. Not Critical. High means the model can provide meaningful assistance to someone attempting a cyberattack, but does not independently enable the kind of catastrophic infrastructure attacks that define the Critical tier. If Astra's evaluation produced a "cannot rule out Critical" finding, that places it in a different category of concern than GPT-5.6 Sol. Not a confirmed escalation, but a failure to produce the clean negative result that would allow OpenAI to say with confidence the model falls below the Critical line.

What Would Make This Independently Auditable

The core problem here is not that OpenAI ran evaluations. Running evaluations is the right thing to do. The problem is that the results are not independently verifiable. For the Astra cybersecurity evaluation to be genuinely auditable, researchers and policymakers would need access to at minimum:

1. The capability report: What specific tasks was Astra tested on? What did it succeed at? What did it fail at? 2. Independent replication results: Did external evaluators like UK AISI or METR run the same tests and reach the same conclusions? 3. The safeguards report: What controls were tested, and how effective were they at limiting the model's ability to assist with attacks? 4. The decision record: How did OpenAI weigh the "cannot rule out" finding against its deployment timeline? Who made that call? 5. Deployment conditions: If Astra is eventually released, under what constraints? What monitoring would apply post-deployment? None of these are currently public. OpenAI's Preparedness Framework commits the company to transparency about safety findings, but the framework does not specify what "transparency" actually requires in practice. That ambiguity is doing a lot of work right now.

Why This Matters Beyond AI Safety Circles

Most people following this story are not AI safety researchers. They are people who use AI tools, follow the news, and have a reasonable interest in whether the technology being deployed around them is being handled responsibly. The honest answer is: we do not know. Not because OpenAI is necessarily acting in bad faith, but because the evidence is not available to assess. A company saying "we ran the tests and we're handling it" is not the same as independent verification. This is also not an abstract concern. The cybersecurity capabilities being evaluated here relate to real infrastructure. The Critical threshold exists because OpenAI itself recognized that some AI capabilities, if deployed without adequate controls, could enable attacks on systems that societies depend on.

The Broader Pattern

The Astra situation fits a pattern that has become familiar in frontier AI development. Companies publish safety frameworks, run evaluations, find concerning results, apply internal controls, and then decide whether and how much to disclose. The public learns fragments, usually through reporting rather than direct disclosure. That is not unique to OpenAI. It reflects a structural gap in how frontier AI safety is currently governed. Independent evaluation bodies like UK AISI and METR exist, but their access and their ability to publish findings are constrained by the companies they evaluate. For now, the most accurate summary of where things stand with Astra is this: OpenAI found something it could not rule out, put controls in place, and has not shown anyone outside the company the full picture.

JM
J. Mendez

Writer

J. Mendez is a writer with a decade of experience covering the full spectrum from politics to entertainment. Holding a degree in political science, Mendez brings analytical depth to reporting on government, policy, and public affairs while also delivering sharp, engaging coverage of film, television, music, and celebrity culture. Over ten years in the field, their work has spanned hard news, cultural analysis, and feature writing, consistently connecting the political and the popular for a broad audience.

Related Stories

Meta Faces Social Media's "Big Tobacco" Moment as 29 States Take the Company to Trial in California
Technology15h ago

Meta Faces Social Media's "Big Tobacco" Moment as 29 States Take the Company to Trial in California

Meta's biggest legal battle yet begins Tuesday in an Oakland federal courtroom, where a coalition of 29 state attorneys general will argue the company deliberately designed Facebook and Instagram to hook children. The numbers on the table are staggering: Meta's own attorneys have said damages could reach $1.4 trillion, while the states' lawyers call $200 billion more likely. And the trial opens less than two weeks after Meta lost a similar case in New Mexico that will cost it nearly $1 billion.

Sony's Digital Monopoly Under Fire: Physical Games End in 2028 and Fans Are Asking Who Really Owns Their Library
Technology17h ago

Sony's Digital Monopoly Under Fire: Physical Games End in 2028 and Fans Are Asking Who Really Owns Their Library

Sony is pulling the plug on physical games beyond January 2028, and the backlash has cracked open a much bigger question: when every purchase is digital and the PlayStation Store is the only shop in town, what do players actually own? The answer, buried in the fine print, is a revocable license. And a recent dispute with an indie studio shows how much power sits on Sony's side of that arrangement.