Skip to content

How AI security barriers are impeding the work of offensive cybersecurity researchers

For months, AI giants have devised special vetted programs and strict barriers to limit the use of their models by malicious hackers. But these limits are now hindering the work of legitimate network defenders, as well as offensive cybersecurity researchers.

In June, the US government imposed export control restrictions about the much-hyped AI models from Anthropic, Mythos and Fable. The move was motivated, at least in part, by a report that claimed it was possible to circumvent security barriers in the models designed to prevent users from using them to create and execute malicious cyberattacks.

Regardless of whether the incident was actually motivated by fears of a leakThe fact is that Anthropo has repeatedly marketed Myths like a kind of cyber doomsday machine that can only be given to carefully vetted users, and even then with strict security barriers. (Export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1; Mythos 5 has been reintroduced only to US organizations vetted as part of the government review process.)

That type of control is not unique to Mythos. Both Anthropic, with its other models, and OpenAI offer cybersecurity researchers programs that they can apply to be vetted and, if approved, access models with fewer cybersecurity restrictions: OpenAI Cyber ​​Trusted Access Program and anthropic Cyber ​​verification program.

These barriers have been widely criticized, particularly by researchers whose job is to find unknown vulnerabilities in systems and devise ways to exploit them before criminals do.

During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, saying that “I’m not really comfortable that these big random companies are making arbitrary decisions about what’s safe and what’s not.”

Dowd has spent decades find and sell “zero days” (previously unknown software flaws and the exploits that take advantage of them) to Western governments, instead of reporting them to software manufacturers for patching. Governments pay a premium for vulnerabilities precisely because they remain open, which is useful for intelligence operations.

Dowd admitted his job may make him biased, but he’s not alone. Several people who work in offensive cybersecurity (proactively probing systems for weaknesses) described to TechCrunch how they use AI tools and manage their security barriers.

Chris Anley, chief scientist at security consulting giant NCC Group, said that asking an AI model to try to exploit a bug is a key step in confirming that it is a real vulnerability worth fixing. But if a guardrail prompts the model to refuse to answer the question directly, the guardrail hurts defenders, he said.

“This is where the whole offensive versus defensive part and security barriers come into play, because ‘fix this code’ as a message is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,” Anley said. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two really can’t be thrown away.”

It’s “like a hammer,” he continued. “You can’t build a house without a hammer. It’s definitely a tool, but it’s also irreducibly a weapon.”

When he and his colleagues hit such a roadblock, they sometimes turn to open-source AI models that don’t have any guardrails.

Paolo Stagno, chief technology officer at Crowdfense, a well-known company that develops, acquires and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying that AI companies “essentially treat customers like children who need to be taken care of” with their vetted programs and guardrails.

Stagno said he and his colleagues use frontier models, but only to reverse engineer them. They avoid using AI to help find vulnerabilities or create exploits, he said, because putting that work into a cloud-based model risks leaking sensitive vulnerability data or having it absorbed into future training runs. For that step, he said, they use open source models run locally, since they do not depend on sharing data outside the model.

Giuseppe Cali, a security researcher who finds zero days and develops exploits, said security barriers do not impede his work. That’s because it doesn’t use AI for offensive work; instead, you use it for initial reverse engineering, to understand the code you are analyzing, and to create supporting tools. For that, he said, AI tools can speed up the process and allow you to focus on discovering vulnerabilities.

“I still want to be responsible for the discovery of errors and the use of weapons and that would not change if all the barriers were lifted tomorrow,” Cali said. “I’m jealous of my bugs and I like this game too much to let the models play it for me.”

A researcher at a smartphone component maker, who spoke on condition of anonymity because he is not authorized to speak to the press, said his employer is not part of Anthropic’s CVP program and, as a result, its tools are barely useful in finding vulnerabilities because the security barriers are too tight.

“If it finds out we’re doing something security-related, it just stops and can’t be used,” the person said.

Chris Thompson, CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an event focused on offensive security and AI, said that based on his experience with frontier AI models, guardrails can be inconsistent and work differently every day. This is true even within the broader confines of the Anthropic and OpenAI programs examined.

“I think the practical impact is that you spend a lot of time negotiating with the model rather than working on the core security program,” Thompson said. “Instead of analyzing a vulnerability and reasoning through exploitability, you’re trying to find why you’re getting inconsistent results or why the models are over-sanitizing the result.”

Consequently, researchers rely on or are pushed toward open-source Chinese models like GLM (freely downloadable models that can be run locally without research or usage restrictions), Thompson said.

“There are these responsible researchers who are being moved from US-governed systems to foreign-owned systems,” he said. “I think it’s more harmful than good to have these guardrails in place.”

Instead of tightening restrictions further, Thompson called for frontier AI labs to open up their programs, provide responsible access, and hold accountable those who abuse their tools. Otherwise, he argued, advocates will lose the AI ​​race.

“There is a huge storm coming. There is a huge wave of attacks that will occur at a speed and scale like never before,” Thompson said. “But the same legitimate security consulting firms and researchers that are trying to make a difference are being stifled right now.”

When you purchase through links in our articles, we may earn a small commission. This does not affect our editorial independence.

Leave a Reply

Your email address will not be published. Required fields are marked *