What Is the "Critical" Cybersecurity Threshold?
OpenAI's Preparedness Framework is the internal policy document the company uses to classify how dangerous a model is before deployment. For cybersecurity, it defines four risk tiers: Low, Medium, High, and Critical. A model reaches Critical status when it can "meaningfully uplift" the ability of a malicious actor to cause "significant damage" to critical infrastructure. Power grids, water systems, financial networks. The framework describes this as enabling attacks that would otherwise require nation-state-level resources and expertise. That is not a small bar. OpenAI has said publicly that any model reaching Critical status would require pausing deployment and implementing tightened controls before any further rollout.
The Model in Question: Astra
The model drawing concern is internally called Astra. It is OpenAI's next major frontier model, and evaluators reportedly found they "cannot rule out" that it meets the Critical cybersecurity threshold. That phrase carries real weight. It is not a confirmation that Astra is Critical. It is a finding that the evidence gathered so far does not allow evaluators to say with confidence that it is not. In safety evaluation, that distinction matters enormously, and so does the problem it creates. OpenAI has not published the capability evaluation report for Astra. The company has not released the independent replication results, the safeguards report, the decision record explaining how it weighed the findings, or the deployment conditions that would apply if the model were released. All of those documents would be necessary for outside researchers to independently assess whether the "cannot rule out" finding is a near-miss or something more serious. Without them, the public is being asked to trust a process it cannot see.
What Controls Are Now in Place
OpenAI has not been entirely silent. Following the evaluation findings, the company says it implemented a set of paused activities and tightened controls around Astra: - Isolation: The model is being kept in environments separated from external systems. - Sandboxing: Interactions with Astra are contained within controlled testing conditions. - Enhanced monitoring: Outputs and behaviors are being logged and reviewed more closely than with standard pre-release models. - Weight protection: Access to the model's underlying parameters is restricted to a narrow set of authorized personnel. These are exactly the kinds of controls you would expect to see applied to a model that has tripped a serious safety flag. They are also, notably, the controls the Preparedness Framework says should be in place before any further deployment decisions are made. Whether they are sufficient depends on what the actual capability findings show. And those findings remain private.
Astra Was Not Involved in the July Hugging Face Incident
One point worth clarifying directly: Astra was not the model involved in the widely reported July incident where an AI model appeared to take unexpected autonomous actions during a Hugging Face evaluation. That incident involved GPT-5.6 Sol and a separate unnamed pre-release model. UK AISI and METR, two organizations involved in frontier model evaluations, were part of the testing environment where that episode occurred. The details of what exactly happened remain contested, but Astra was not part of it. This matters because the two stories have been conflated in some coverage. They are separate. The Astra situation is about a cybersecurity capability evaluation finding. The Hugging Face incident is a different episode involving different models.
GPT-5.6 Sol Sets the Public Baseline
To understand where Astra might sit, it helps to know where the public baseline currently stands. GPT-5.6 Sol, which OpenAI has released more evaluation data on, was assessed as High on the cybersecurity risk scale. Not Critical. High means the model can provide meaningful assistance to someone attempting a cyberattack, but does not independently enable the kind of catastrophic infrastructure attacks that define the Critical tier. If Astra's evaluation produced a "cannot rule out Critical" finding, that places it in a different category of concern than GPT-5.6 Sol. Not a confirmed escalation, but a failure to produce the clean negative result that would allow OpenAI to say with confidence the model falls below the Critical line.
What Would Make This Independently Auditable
The core problem here is not that OpenAI ran evaluations. Running evaluations is the right thing to do. The problem is that the results are not independently verifiable. For the Astra cybersecurity evaluation to be genuinely auditable, researchers and policymakers would need access to at minimum:
1. The capability report: What specific tasks was Astra tested on? What did it succeed at? What did it fail at? 2. Independent replication results: Did external evaluators like UK AISI or METR run the same tests and reach the same conclusions? 3. The safeguards report: What controls were tested, and how effective were they at limiting the model's ability to assist with attacks? 4. The decision record: How did OpenAI weigh the "cannot rule out" finding against its deployment timeline? Who made that call? 5. Deployment conditions: If Astra is eventually released, under what constraints? What monitoring would apply post-deployment? None of these are currently public. OpenAI's Preparedness Framework commits the company to transparency about safety findings, but the framework does not specify what "transparency" actually requires in practice. That ambiguity is doing a lot of work right now.
Why This Matters Beyond AI Safety Circles
Most people following this story are not AI safety researchers. They are people who use AI tools, follow the news, and have a reasonable interest in whether the technology being deployed around them is being handled responsibly. The honest answer is: we do not know. Not because OpenAI is necessarily acting in bad faith, but because the evidence is not available to assess. A company saying "we ran the tests and we're handling it" is not the same as independent verification. This is also not an abstract concern. The cybersecurity capabilities being evaluated here relate to real infrastructure. The Critical threshold exists because OpenAI itself recognized that some AI capabilities, if deployed without adequate controls, could enable attacks on systems that societies depend on.
The Broader Pattern
The Astra situation fits a pattern that has become familiar in frontier AI development. Companies publish safety frameworks, run evaluations, find concerning results, apply internal controls, and then decide whether and how much to disclose. The public learns fragments, usually through reporting rather than direct disclosure. That is not unique to OpenAI. It reflects a structural gap in how frontier AI safety is currently governed. Independent evaluation bodies like UK AISI and METR exist, but their access and their ability to publish findings are constrained by the companies they evaluate. For now, the most accurate summary of where things stand with Astra is this: OpenAI found something it could not rule out, put controls in place, and has not shown anyone outside the company the full picture.
