Skip to content

OpenAI restricts access to the Astra model’s advanced cyber features due to hacking concerns

OpenAI is changing its model launch strategy as its technology becomes more powerful and the potential for its misuse increases – particularly after the July incident that tested its AI models planned and carried out independently a cyberattack against the AI ​​company Hugging Face.

The company’s next model, Astra, is coming to market “soon,” OpenAI said. It said Astra is significantly more powerful than the company’s current top-of-the-line AI model, GPT-5.6 Sol, which is extremely capable even in cyber tasks. But only a handful of partners will have access to its most advanced cybersecurity capabilities as OpenAI works to balance helping companies prevent cyberattacks while empowering attackers, a company spokesman told reporters at a briefing today.

OpenAI is recruiting customers to use its models to prevent cyberattacks or for “defensive cybersecurity.” The Company views these sales as a critical source of revenue and one of its key priorities new Chief Revenue Officer Dali Rajic.

The small group of “alpha testers” with full access to Astra’s cybersecurity capabilities include “individuals and organizations responsible for protecting critical digital infrastructure and critical infrastructure more generally,” an OpenAI spokesperson said. This includes the US government and companies in OpenAI trusted access program for cybersecurity. OpenAI declined to name these organizations.

OpenAI will monitor the model’s performance in this small group and continue to expand access through its “Daybreak Blue” program once it is confident that Astra has “the right calibration” and can “provide defensive advantages while reducing the potential for abuse,” the company said.

Astra is already a few weeks late

Astra’s release has already been “delayed by a few weeks because everything was paused after Hugging Face, and then we took additional time to make sure what we were launching was safe,” an OpenAI spokesperson said.

OpenAI paused training new models for two weeks following the Hugging Face incident to strengthen its internal security measures. These changes included the introduction of greater agent monitoring since the company didn’t know The Hugging Face hack was only executed a week after it occurred, and testing environments have become more isolated, preventing AIs from escaping and infiltrating other companies.

Although the Astra model was not part of the Hugging Face incident, OpenAI says it is both more powerful and efficient than GPT-5.6 Sol, which was involved in the breach. (Another unreleased AI model that OpenAI has not publicly named also played a key role in the Hugging Face cyberattack. OpenAI has since deactivated that model.) Importantly, OpenAI says Astra is the first model it plans to release that meets its “critical cybersecurity readiness threshold” under its Preparedness Framework, an internal policy that governs the safeguards the company will put in place depending on the risks a model poses. This means that, under the right conditions, Astra can find and exploit previously unknown vulnerabilities without human oversight.

Astra has already demonstrated its hacking skills in internal evaluations. In one test, OpenAI created a benchmark called ExploitBench that included 20 high-severity vulnerabilities. The model outperformed GPT-5.6 Sol in the test and “even discovered and exploited two zero-day vulnerabilities as part of an exploit chain,” OpenAI said. “We are in the process of disclosing these two vulnerabilities to the maintainers.”

At the same time, Astra is more likely to reject inappropriate requests than GPT-5.6 Sol, OpenAI said. In a cyber assessment, Astra rejected 91.5% of requests, compared to 59% for GPT-5.6 Sol, although that means it still fulfilled 8.5% of requests.

Astra may refuse legitimate cybersecurity requests

OpenAI is taking “particular care to ensure that this deployment is safe and secure” – but this brings with it another trade-off. Astra may be overly cautious and reject legitimate cybersecurity requests. As a theoretical example, if someone asks to help find and fix a vulnerability, they might mistakenly think they are trying to carry out an attack and not comply.

Rejections like this are the reason for Hugging Face said It was forced to use a Chinese open source model to combat the OpenAI hack. The company tried to use Anthropic’s models to combat the attack, but was too cautious and refused.

OpenAI, like other pioneering AI companies, is trying to find ways to give its models an inherent sense of right and wrong and ensure they are “consistent” with human values ​​and norms, the company said. It is working to train its models to respect boundaries like a human would, such as knowledge of “the rule of law,” a company spokesman said.

Leave a Reply

Your email address will not be published. Required fields are marked *