Two focused professionals in cyber security team working to prevent security threats, find vulnerability and solve incidents. Woman pointing on a event map.

Can AI guardrails stop cyber threats? Security experts see promise, and limits

For over a year, AI agents seemed to be the hot new talking point when it came to the evolution of AI models. Now, with the release of even more powerful agentic models from companies around the world, and governments attempting to curb their use, “safeguards” is the new buzzword—and it’s one that could leave security experts on the back foot, in an unexpected way.

For the panelists on the latest episode of the Security Intelligence podcast, the real issue is stopping AI agent exploits before they disrupt business operations, and it crops up when not everyone plays by the rules—or safeguards—established.

All eyes on AI safety

Anthropic’s Claude Fable 5 and OpenAI’s GPT-5.6 Sol, the companies’ most capable models to date, have been mired in controversy since the US government imposed, then lifted, its temporary restrictions on access to Anthropic’s new agentic model due to potential security risks highlighted by unspecified national security officials. Both models feature increased reasoning abilities, along with a set of increasingly strict safeguards focused on cybersecurity. But can these new AI models, paired with a new set of governance rules, be enough to keep malicious actors at bay?

For IBM Distinguished Engineer Jeff Crume, the inclusion of safeguards that potentially slow performance is just part of the job. “I think the AI companies are realizing if they don’t do this, there’s going to be pushback,” Crume said. “They’re seeing this as something that’s necessary for them to continue their business. I think we’re fighting a losing battle and always will be on this front. It doesn’t mean that it’s not worth fighting.”

Building a collective defense

Anthropic’s Project Glasswing, where participating companies, including IBM, are developing a shared framework for addressing AI vulnerabilities, is one way Anthropic is combatting the potential for jailbreaking its latest AI models. Diego Matos Martins, IBM X-Force Incident Response Leader for Latin America, said that partnering with other organizations and sharing expertise can speed up the development of security standards and mitigate risk. “Those partnerships—not just with companies, but also the sectors of the government—[are] important,” he said.

While cross-agency collaboration is critical in the effort to safeguard emerging AI models, some experts believe attackers will still find a way around those guardrails. “It’s important for these companies to create these standards, but ... it is a bit of a whack-a-mole,” said Sophie Cunningham, a Cyber Threat Intelligence Analyst at IBM X-Force focused on the dark web. “I don’t think you can create all the standards you’ll possibly need. We can’t protect against everything that any person may ever think to leverage.”

The potential, and practical limits, of AI guardrails

One tool making waves without guardrails? Chinese AI company Z.ai’s GLM-5.2, an open-weight model on par with Anthropic’s Mythos in terms of benchmarks. Not having guardrails doesn’t necessarily make an AI model malicious, but when bad actors have access to powerful AI tools, there is always the risk of vulnerabilities being exploited or threats being executed at a pace defenders can’t keep up with.

“One set of responsible players will put in guardrails, another set won’t,” Crume said.  “We have to be aware that even if we get all of the good guys to put in the guardrails, it only takes one bad guy to [circumvent them].”

For security teams, that doesn’t mean guardrails shouldn’t be prioritized. As Crume sees it, today’s AI models can help give defenders the upper hand when used responsibly and strategically. “The long arc … is that anything we can do to identify vulnerabilities allows us to be more proactive, it allows us to identify the zero days before the bad guys [can] take advantage of it,” he said. “If we find it first, then we can do something to prevent that. That’s the good news story here.”

Security Intelligence | 5 August, episode 45

Your weekly news podcast for cybersecurity pros

Whether you're a builder, defender, business leader or simply want to stay secure in a connected world, you'll find timely updates and timeless principles in a lively, accessible format. New episodes on Wednesdays at 6am EST.

Patrick Lucas Austin

Staff Writer

IBM Think

Related solutions
IBM Guardium

Detect and respond to threats, gain real-time visibility and enforce security and compliance across your data estate.

Explore IBM Guardium®
AI cybersecurity solutions

Improve the speed, accuracy and productivity of security teams with AI-powered solutions.

    Explore AI cybersecurity solutions
    Security services

    Transform your business and manage risk with a global leader in cybersecurity, cloud and managed security services.

    Explore security services
    Take the next step

    Accelerate threat detection and response with AI-powered insights while protecting critical data with real-time visibility, threat detection and automated security controls.

    1. Discover IBM Guardium®
    2. Explore AI cybersecurity solutions