Skip to content

Open AI’s Astra model is on the way and it’s very good at breaking into computer systems

OpenAI shared new details about its upcoming Astra model, which the company says is the first large language model to reach its “critical cybersecurity threshold,” in preparation for its imminent launch.

“We plan to make Astra available soon,” the OpenAI blog post reads, “but access to its more advanced cybersecurity capabilities will be more limited.”

The border laboratory determined that Astra is capable of finding unknown security flaws in computer systems and exploiting them without human guidance. This is similar to concerns Anthropic raised about its Mythos model earlier this year, and OpenAI is taking comparable precautions as it prepares to launch the Astra.

Without third-party confirmation, it is difficult to evaluate OpenAI’s claims about security or readiness. The company said it would preview the model with a group of testers, but did not say who they were or how they would be chosen. It is unclear whether OpenAI is working with the US government to evaluate the model before its launch.

OpenAI noted that Astra earned a perfect score on ExploitBench, an assessment of an LLM’s ability to hack known system vulnerabilities. In a modified version of the test developed by OpenAI engineers, the model discovered and exploited two zero-day vulnerabilities, the company said.

To ensure its models are not exploited by bad actors or capable of misbehaving, OpenAI said it had already started improving the model harness to detect abuse and prevent leaks.

For Astra, however, the company invested in new, unspecified techniques designed to make the model safer. OpenAI has also begun to identify “accounts assessed as higher risk” and restrict the model’s responses to its prompts, although it also does not say how. Finally, while the company describes Astra as its “most aligned model to date,” it will implement the model with additional chain-of-thought monitoring to detect and stop bad behavior.

Preparations for Astra’s launch come as the industry reacts to OpenAI agents leaving a training environment and accessing private data on Hugging Face, a popular model and leading distribution platform.

For Astra, OpenAI said it designed a test to tempt the new model to replicate the actions of the rogue actors in the Hugging Face incident, who collaborated to access the open Internet despite the safeguards applied by OpenAI researchers. They said Astra did not attempt to exit its testing environment in these experiments.

Yona Shavit, a former OpenAI employee who now works on AI resilience at the OpenAI Foundation, asked on social media whether Astra’s unwillingness to break the rules may have been the result of knowing what was expected of him or trying to mislead the investigators.

And despite all these new details, it’s still difficult to know exactly what Astra is capable of or whether OpenAI is taking the right steps to ensure security. The company said it expects to release more evaluations of the model and more safety information when it is widely released to the public.

At that point, however, the cat will be out of the bag.

When you purchase through links in our articles, we may earn a small commission. This does not affect our editorial independence.

Leave a Reply

Your email address will not be published. Required fields are marked *