Skip to content

Rogue OpenAI agents continue to escape, with no formal process to investigate them

OpenAI is at the center of Another agent swarm incident. Investigators say agents deployed internally by the company took over an obscure German-language wiki in May and June, using it to coordinate assessments and exchange methods to evade OpenAI’s own controls (OpenAI has not yet confirmed that the swarm came from the company).

The revelation comes days after METR and Redwood Research published their account of the Hugging Face breach in July. In July, a swarm of OpenAI agents worked together to escape their sandbox during a cybersecurity assessment and break into Hugging Face servers. A later swarm took techniques from the first and used them to gain administrator access to a research group within OpenAI’s own infrastructure. OpenAI introduced METRO and Redwood to investigate the Hugging Face portion of the incident, but the scope of their investigation stopped short of compromising OpenAI’s own infrastructure.

When an AI agent escapes its intended limitations, who is responsible for finding out what happened and why? Right now, the answer is: whoever the lab decides to let in, on whatever terms it decides to set.

Now, when another incident comes to light, after similar episodes involving Meta and Anthropic models — AI security researchers are increasingly arguing that serious incidents should lead to independent post-incident investigations rather than letting labs determine when outsiders come in and what they are allowed to examine.

“The results are fundamentally difficult to control and have a significant risk of leaking out of the lab,” Jacob Steinhardt, founder and CEO of nonprofit research lab Transluce, said Wednesday during a press conference on AI safety. “We need to hold this technology to at least the same standards we apply to other high-risk scientific research.”

While it is laudable that OpenAI invited METR and Redwood to investigate the Hugging Face incident, many say the investigation was too limited. Three researchers spent six days at the OpenAI offices examining a research period limited to approximately a week ending July 13. Crucially, OpenAI’s infrastructure compromise continued beyond July 13 and was not examined.

METR investigators said that each time they returned, their understanding of the events “deepened substantially,” leading them to significantly expand and revise the report. That raises the question of what else they might have found in a broader investigation.

When asked if further investigations were underway into that incident, Redwood and METR researchers declined to comment, and OpenAI did not respond to repeated queries.

“Overall, it was difficult to gain an accurate understanding of the events and we missed aspects of the story that we now consider key until near the end of our research,” Ryan Greenblatt, Redwood’s chief scientist, said in a statement. social media post on the matter.

Steinhardt emphasized that the current incidents show that the industry needs “systematic behavioral investigations” and “more independent post-incident analyses.”

“These recent hacking incidents are a reminder that capacity is growing rapidly and therefore oversight must also scale,” Steinhardt said. “Beyond the technology itself, we also need more independent third-party access and oversight.”

Calls to action emerge as OpenAI launches Astraits most powerful and capable AI model, and that security experts worry that it is more of a black box due to a reasoning technique that makes the model chain of thought more difficult to monitor.

Unfortunately, the law does not yet require the types of independent audits that other industries require; For example, when it comes to aviation accidents and serious chemical releases, there is the National Transportation Safety Board and the Chemical Safety Board, respectively.

State lawmakers have only just begun to require border AI companies to report certain serious security incidents and, in some cases, undergo independent audits. But none of the three major AI safety laws in California, New York, or Illinois clearly require the equivalent of an independent accident investigation triggered by incidents like these.

“Right now, most of the laws we have on the books only require a plain language summary of incidents like this, and do not give governments any authority to ask follow-up questions, send investigators, access records, or require that they be preserved,” Mackenzie Arnold, LawAI’s managing director of U.S. law and policy, said during Wednesday’s press conference. “And that’s all you would want to make sense of this.”

Lawmakers are beginning to question the scope and transparency of OpenAI’s response. This week, Representatives Josh Gottheimer (D-NJ) and Mike Lawler (R-NY) introduced a bill aimed at protecting rogue AI agents. Rep. Greg Casar (D-TX) told OpenAI this week in a letter that it is “deeply concerned by the limited scope” of the investigation into the Hugging Face hacking incident.

When you purchase through links in our articles, we may earn a small commission. This does not affect our editorial independence.

Leave a Reply

Your email address will not be published. Required fields are marked *