AI guardrails hinder offensive cybersecurity researchers, single source reports

AI guardrails hinder offensive cybersecurity researchers, single source reports

10 reported

According to a TechCrunch report, AI guardrails designed to prevent malicious use are also impeding legitimate offensive cybersecurity researchers. The article states that in June, the U.S. government placed export control restrictions on Anthropic’s Mythos and Fable models, partly due to a report claiming guardrails could be bypassed. Those restrictions have since been lifted, with Fable 5 returning to general access on July 1 and Mythos 5 reintroduced only to vetted U.S. organizations. Researchers quoted in the article describe guardrails as inconsistent and say they sometimes turn to open source models with no restrictions. One researcher said guardrails push responsible researchers away from U.S.-governed systems toward foreign-owned models. The article notes that this is a single-source report and has not been independently verified.

What’s reported

In June, the U.S. government placed export control restrictions on Anthropic’s Mythos and Fable AI models, partly due to a report claiming guardrails could be bypassed.
The export controls on Fable 5 and Mythos 5 have since been lifted; Fable 5 returned to general access on July 1, and Mythos 5 has been reintroduced only to vetted U.S. organizations.
Anthropic and OpenAI offer vetted programs for cybersecurity researchers to access models with fewer restrictions: OpenAI’s Trusted Access for Cyber program and Anthropic’s Cyber Verification Program.
Security researcher Mark Dowd said it is “not really comfortable” that large companies make arbitrary decisions about what is safe in security.
Chris Anley, chief scientist at NCC Group, said guardrails can hurt defenders by refusing to answer prompts that confirm vulnerabilities.
Paolo Stagno, CTO of Crowdfense, said AI companies “essentially treat customers like children who need babysitting.”
Researcher Giuseppe Cali said guardrails are not impeding his work because he does not use AI for offensive work.
One anonymous researcher at a smartphone-component manufacturer said his employer is not part of Anthropic’s CVP program, making the tools barely useful.
Chris Thompson, CEO of RemoteThreat, said guardrails are inconsistent and push researchers toward Chinese open source models like GLM.
Thompson called for AI labs to open up programs and hold abusers accountable, arguing defenders will lose the AI race otherwise.

Key figures

Mark Dowd, security researcher
Chris Anley, chief scientist at NCC Group
Paolo Stagno, chief technology officer at Crowdfense
Giuseppe Cali, security researcher
Chris Thompson, chief executive of RemoteThreat and founder of Offensive AI Con
One anonymous researcher at a smartphone-component manufacturer

Sources: TechCrunch

You may also like...

Leave a Reply

Your email address will not be published. Required fields are marked *