OpenAI "cannot rule out" that its upcoming model Astra has "critical" cyber capabilities, a designation that has prompted the company to expand safety testing and pause internal activities that do not meet stricter security requirements, OpenAI told Axios.
The company told Axios that internal evaluations of Astra left it unable to rule out "critical cyber capabilities" in the system, a designation sitting at the upper end of the risk tiers set out in OpenAI's preparedness framework, first published in 2023 to govern how dangerous capabilities are assessed before deployment.
As a result, OpenAI has widened its security testing regime and halted internal work that does not meet tighter safeguard requirements.
The firm has also signalled it intends to slow Astra's development until adequate protections are in place, a step that could push back any eventual public release.
It is to be noted that Astra was not connected to a recent set of exploits affecting Hugging Face, the machine learning platform, the company clarified.
The caution around Astra follows closely on the heels of a very different kind of announcement. On 1 August, OpenAI published a 249-page collection of results showing Astra had solved ten mathematics problems that had stood open for at least a decade, several for far longer, spanning group theory, high-dimensional geometry, coding theory, quantum complexity, lattice cryptography and extremal combinatorics.
OpenAI put the total API cost of producing the successful runs at roughly $2,000.
Unlike typical AI benchmark claims, every result was accompanied by a machine-checkable certificate formalised in Lean, a proof assistant that verifies each logical step and rejects any argument that does not follow.
The certificates were published openly on GitHub, meaning outside mathematicians could verify the work themselves rather than take OpenAI's word for it. Thomas Bloom, who maintains the catalogue of open problems left behind by Paul Erdős, three of which featured among Astra's results, called the release "big news".
The achievement built on an earlier result in May, when the same model family reportedly disproved the 80-year-old Erdős unit distance conjecture, a proof that Fields Medalist Tim Gowers said he would have recommended for publication without hesitation.
A White House official told Axios that OpenAI had proactively briefed the administration, noting the company had “informed the administration of their plans to delay the release.”
The acknowledgement comes as the Trump administration works to formalise a review process for powerful AI models before public release, though questions including review length and access to underlying systems remain unresolved.
The episode revives scrutiny of a commitment Anthropic made previously, pledging to pause training of its most powerful models should their capabilities outstrip its ability to control them.
That pledge was scaled back in a February update to Anthropic's Responsible Scaling Policy, which now argues unilateral pauses could backfire, warning that if one developer halted work while rivals kept shipping systems without comparable safeguards, the result could leave the wider AI landscape less safe rather than more.
Anthropic has since taken a more calibrated approach with its own high-capability systems, releasing a more heavily safeguarded version of its most cyber-capable model, Mythos, in June.
Dianne Penn, Anthropic's head of product management, research and labs, told Axios the company had been "deliberately more conservative" with that release. Anthropic separately warned in a June blog post about models capable of improving themselves, coupling it with a call for a pause across the industry.
The disclosure follows remarks earlier in the week at the Black Hat cybersecurity conference, where OpenAI technical staff said the company was already easing testing while overhauling its security practices.
In a blog post published the same day, OpenAI said it had begun rolling out stricter controls, including isolated testing environments and monitoring systems tracking agentic behaviour across Astra's applications.
Catch all the Business News, Market News, Breaking News Events and Latest News Updates on Live Mint. Download The Mint News App to get Daily Market Updates.
Oops! Looks like you have exceeded the limit to bookmark the image. Remove some to bookmark this image.
Facts Only
OpenAI cannot rule out that its Astra model possesses critical cyber capabilities.
OpenAI has expanded safety testing and paused internal activities not meeting new security requirements.
Development of Astra is being slowed, potentially delaying public release.
OpenAI briefed the White House administration on plans to delay the release.
On August 1, OpenAI published results showing Astra solved ten open mathematics problems.
These mathematical results included machine-checkable certificates formalized in Lean.
The API cost for the successful mathematical runs was approximately $2,000.
In May, a model from the Astra family reportedly disproved the Erdős unit distance conjecture.
Anthropic released a safeguarded version of its model, Mythos, in June.
Anthropic updated its Responsible Scaling Policy in February to scale back a pledge to pause training.
OpenAI has implemented isolated testing environments and monitoring for agentic behavior.
Astra was not connected to recent exploits affecting Hugging Face.
Executive Summary
OpenAI is delaying the public release of its Astra model after internal evaluations failed to rule out "critical cyber capabilities." This designation, situated at the highest risk tier of the company's 2023 preparedness framework, has triggered a widening of security testing and a halt to internal work that does not meet stricter safeguards. The company has proactively briefed the White House on these delays, coinciding with efforts by the Trump administration to formalize review processes for powerful AI models.
Conversely, Astra has demonstrated significant breakthroughs in mathematics, solving ten long-standing problems across fields like lattice cryptography and quantum complexity. These results are verified by Lean certificates on GitHub, providing transparent, machine-checkable proof of the model's reasoning. This tension between high-level intellectual capability and potential cyber risk mirrors a broader industry struggle, as seen with Anthropic’s calibrated release of its Mythos model and the subsequent debates over whether unilateral pauses in development increase or decrease overall systemic safety.
Full Take
The strongest version of this narrative is that of a responsible corporate actor identifying a dual-use dilemma—where a system capable of solving profound mathematical proofs is also capable of automating cyber-attacks—and choosing safety over speed. This portrays a maturing industry adopting a "defense-in-depth" strategy before deployment.
The pattern here is the juxtaposition of "transparency" in academic achievement (GitHub certificates) against "opacity" in security risk (internal "critical" designations). By highlighting the verifiable brilliance of the model in mathematics, the narrative builds a case for the model's power, which then justifies the necessity of the secret, high-level security protocols. This creates a loop where the model's proven intelligence makes its theoretical danger more believable, even without public evidence of the latter.
Patterns detected: none
The driving paradigm is the "Capabilities vs. Alignment" race. The unstated assumption is that high-level reasoning in mathematics naturally translates to high-level capability in cyber-exploitation. This echoes the historical pattern of nuclear proliferation: the same physics that enables energy production enables weaponry.
The implication is a shift toward "closed-door" governance. If the most capable models are deemed too dangerous for public release without government-vetted safeguards, the locus of control shifts from open research to a tight circle of corporate and state actors. This risks creating a "security priesthood" where only a few have access to the most powerful cognitive tools.
Bridge Questions:
1. Does a model's ability to solve formal proofs in Lean necessarily imply an ability to find zero-day vulnerabilities in non-formalized code?
2. How does the "competition" argument used by Anthropic change the ethical calculus of a safety pause?
3. What objective, third-party metrics could replace internal "risk tiers" to verify safety without compromising security?
Counterstrike Scan: An influence campaign would likely exaggerate the "critical cyber" threat to trigger regulatory capture or panic, while simultaneously hyping the "math breakthroughs" to maintain investor confidence. The current content remains a straightforward report of corporate actions and public releases; it does not match the structural intensity of a coordinated campaign.
Sentinel — Human
The article functions as a synthesis of several related developments in AI safety and capability disclosure, skillfully linking technical achievements with institutional policy shifts.
