How AI guardrails are impeding the work of offensive cybersecurity researchers
By Jakub Antkiewicz
•2026-07-24T10:16:27Z
The Defender's Dilemma: AI Guardrails Hinder Cybersecurity Research
Safety restrictions and vetted access programs from AI leaders like Anthropic and OpenAI, designed to prevent the malicious use of their models, are now actively impeding the work of legitimate cybersecurity researchers. While intended to stop hackers from building cyberweapons, these guardrails often block the very queries that network defenders use to identify, validate, and patch critical vulnerabilities. This has created a significant operational friction, forcing the security community to question whether the current approach to AI safety is inadvertently disarming the people it's meant to protect.
Vetted Programs and Their Unintended Consequences
AI companies have attempted to address this by creating special access channels, such as OpenAI’s Trusted Access for Cyber program and Anthropic’s Cyber Verification Program. However, researchers report that even within these programs, the guardrails can be inconsistent and overly restrictive. Chris Anley, chief scientist at NCC Group, noted that the same prompt can be both an offensive and a defensive tool, making it difficult to separate legitimate research from malicious intent. This forces security professionals into a frustrating loop of rephrasing queries and 'negotiating' with the model rather than focusing on the core security task. For sensitive work like zero-day vulnerability discovery, some researchers, like Paolo Stagno of Crowdfense, avoid cloud-based frontier models altogether to prevent potential data leaks, opting instead for local open-source alternatives.
- Blocked Queries: Guardrails frequently refuse to process prompts related to exploit generation, which is a key step for researchers to confirm a vulnerability's severity.
- Inconsistent enforcement: Researchers report that safety filters behave differently from day to day, even within vetted programs, making their workflow unreliable.
- Forced Workarounds: To avoid both guardrails and the risk of data leaks, many professionals are turning to open-source models that can be run locally without restrictions.
- Developer Friction: A significant amount of time is wasted trying to bypass or reason with the AI's safety protocols instead of performing security analysis.
A Shift Toward Unregulated Open-Source Models
The practical impact of these strict controls is a migration of security talent towards unregulated and often foreign-developed open-source models. Chris Thompson, founder of the Offensive AI Con, warned that responsible U.S. researchers are being pushed toward systems like China's GLM, which have no such restrictions. This trend risks creating an uneven playing field where malicious actors have unfettered access to powerful tools while defenders are artificially constrained. The consensus among many offensive security experts is that by over-sanitizing access, U.S. AI labs may be stifling the very innovation needed to counter the next wave of AI-powered cyberattacks, potentially losing the AI security race before it has truly begun.
The current AI safety paradigm, which treats offensive security queries as inherently malicious, is creating a critical capability gap. By restricting legitimate researchers, major AI labs are inadvertently pushing the security community toward unregulated foreign models, potentially giving adversaries a permanent advantage in the development of AI-driven exploits.