When AI guardrails inadvertently tie the hands of cybersecurity experts
- Thảo Nguyên

- Jul 24
- 4 min read
In an effort to prevent hackers from leveraging artificial intelligence (AI) to execute cyberattacks, major tech corporations have established highly stringent censorship filters and regulatory guardrails. However, these defensive measures have inadvertently become major barriers for legitimate cybersecurity professionals and offensive security researchers - those who proactively hunt for vulnerabilities to protect systems.
This concern became even more apparent in June when the U.S. government imposed export control restrictions on two of Anthropic's prominent AI models, Mythos and Fable. This decision stemmed from a report showing that users could still bypass the models' protective layers to build and deploy malicious cyberattacks. Nevertheless, this tightening of regulations is leaving significant repercussions for the security community.
The double-edged sword of AI censorship
For professionals working in offensive security, testing the exploitation of a system bug is a crucial step to determine whether it poses a real risk that needs to be remediated. However, when AI models mechanically refuse to respond due to protective guardrails, the work of defensive experts grinds to a halt.

The nature of AI tools in this domain is akin to a "fix this code" command. It serves as both an essential protective solution and, at the same time, a map highlighting critical vulnerabilities within a system. This renders AI an inseparable dual-use tool: it is both a shield and potentially a weapon. It is similar to a hammer - you cannot build a house without one, but no one can deny its destructive power if misused. Faced with these overly rigid barriers, many researchers are forced to abandon large commercial AI models in favor of uncensored, open-source models.
The frustration of "negotiating" with machines
Many industry experts have expressed frustration that AI companies treat professional clients like children in need of a babysitter. Overly sensitive safety systems render AI far less effective. A researcher at a smartphone component manufacturing company shared that because his enterprise was not part of Anthropic's CVP partner program, the company's AI tool was virtually useless for vulnerability hunting. The moment the system detects any security-related elements, it immediately shuts down.
Beyond being overly strict, these guardrails operate inconsistently and change daily, even within more relaxed programs from Anthropic or OpenAI. Consequently, experts waste excessive time "negotiating" and trying to bypass AI filters, rather than focusing on their core professional work: deeply analyzing the architecture of security vulnerabilities.
Data leak risks and the shift to open source
In addition to the annoyance of filters, information security when using cloud-based AI poses a complex challenge. Uploading sensitive vulnerability data to the web can result in leaks or have the information absorbed into future AI training datasets. Consequently, many experts limit their use of advanced AI models to the reverse engineering phase. For vulnerability hunting, they prefer local open-source models running on their own machines to ensure data never leaves the premises.
The excessive tightening of U.S.-regulated AI systems is inadvertently driving responsible researchers toward foreign open-source alternatives, such as China's GLM model, which can be freely downloaded and operated locally without any censorship. Many argue that erecting such strict barriers currently does more harm than good.
A different perspective from experts: Mastering the game
That said, not everyone feels hindered by AI guardrails. Some researchers specializing in hunting zero-day vulnerabilities state they do not rely on AI to discover flaws or build exploit tools. Instead, they view AI merely as an assistant to accelerate initial code analysis and write utility scripts to save time. Discovering and mastering vulnerabilities independently not only safeguards proprietary insights but also preserves the engineer's passion for discovery - something they refuse to surrender to any computer model.
The tightening censorship by AI labs is inadvertently suffocating legitimate cybersecurity professional - those striving to make a difference in protecting the digital world. In the face of an impending wave of cyberattacks of unprecedented scale and speed, instead of tightening restrictions, AI developers need to responsibly expand access and establish clear accountability mechanisms for abusers. Otherwise, cyber defenders may very well end up on the losing side of this grueling technological race.
IPSIP Vietnam - Balancing AI power and a trusted cybersecurity partner
As mentioned, with AI tools creating numerous mechanical barriers for defense, having a flexible, human-led cybersecurity strategy becomes all the more critical. Enterprises cannot leave their entire security at the mercy of rigid algorithms; instead, they need real expert teams to master technology, transcend machine limitations, and proactively hunt for vulnerabilities.
With over 15 years of experience, two 24/7 operations centers, and a team of experts holding top international certifications (Fortinet, SentinelOne, Wallix, AWS...), IPSIP Vietnam is proud to be a strategic cybersecurity partner accompanying businesses.

IPSIP not only intelligently applies the groundbreaking power of AI but also provides flexible, tailor-made services to comprehensively protect enterprise infrastructure against all risks:
AI integration in automated security: Applying advanced technology into systems, such as AI-powered Domain Protection, to accurately and effectively identify and mitigate internet threats.
Penetration Testing (PENTEST) & Vulnerability Scanning: IPSIP's cybersecurity experts proactively act as white-hat hackers to hunt for flaws and test system limits—perfectly compensating for the shortcomings of over-censored AI tools.
Security & Network Operations Centers (24/7 SOC & NOC): Operating as a continuous protective shield, providing proactive monitoring, early incident detection, and real-time 24/7 response to optimize network performance (supporting White Label SOC models).
Advanced security ecosystem: Providing a full suite of leading defensive solutions, including Network Detection and Response (NDR) / Extended Detection and Response (XDR), Privileged Access Management (PAM/BASTION), Next-Gen Firewalls, and Double Data Encryption.
Comprehensive security solutions for SMEs (FlexSecure360): A cost-optimized, all-in-one protection package specifically designed and tailored to fit the budgets of small and medium-sized enterprises.











Comments