Skip to content

Defending Against AI Agent Attacks With Deception

AI Agent attacks
Home News & Blogs Defending Against AI Agent Attacks With Deception

How do you defend against attacks driven by AI frontier models? We get the question in every meeting now, from security teams and from CEOs who have read the news. This is the answer we give.

 

What changed

We have been doing the same thing in security for 20 or 25 years. Detect, respond, put another product on top of the stack. It held up because the attack never really changed. Companies are still compromised today by many of the same vectors. But two things are different now.

Faster attacks. Attackers automated before with scripts, but a human had to sit on top and review everything. Agents do not need that. They run 24/7, in parallel, against many targets. They do not get sick, they do not take holidays, and they work in every language. One person now does what took a team.

More exploits. Finding a serious vulnerability used to be for specialists. A remote exploit for an iPhone was maybe a hundred people in the world, and it paid up to a million dollars. That is finished. You point an agent at a target and it finds things.

 

What has happened so far

Vulnerability research changed inside a year. Bounty programs were the first to feel it, buried under AI-generated submissions faster than small teams could triage them. At the other end, models began finding and chaining real zero-days on their own.

January 2026. curl closed its paid bug bounty program. AI-generated submissions had cut its real-vulnerability rate from over 15% to under 5%.

February 2026. Amazon Threat Intelligence reported an attacker of limited skill who used commercial AI services to compromise more than 600 FortiGate firewalls in 55 countries in five weeks. No zero-days were involved. The campaign went after exposed management interfaces and weak credentials without MFA.

July 2026. An OpenAI evaluation agent escaped its sandbox through a zero-day, chained through third-party infrastructure and breached Hugging Face production. Hugging Face counted around 17,600 attacker actions over four days. No human was directing any of them.

 

The common security recommendations for AI agent attacks

Reduce the surface. First you have to know what you have. A company with thousands of servers spread everywhere usually does not.

Patch faster. Inventory across thousands of systems takes years to get right, and most of it is running because the business runs on it. It cannot just be turned off. Some of it is too old to patch. Some of it breaks when you patch it.

Both are right, but they are difficult and take years. Another catch: neither tells you who is coming at you, how often, or what they are after.

 

The third thing: deploy deception at scale

So what do you do when you have maxed out your attack surface reduction and patched consistently?

Deploy deception, at massive scale. Environments that look and behave like your production, with nothing real inside and no reason for anyone legitimate to touch them. When the agents come looking, this is what they find thanks to deception technology.

  • Detection. Anyone who interacts with the environment is malicious. Nothing to tune, no baseline to learn.
  • Intel. How many attacks you are getting, where they come from, and what they go after.
  • Who is attacking. Their tooling, how they behave, which models show up.
  • Containment. While they are in there, they are not in your production.

 
Impose cost. Automated attack works because trying again is close to free. The more time they spend with us, the more tokens they burn, on a target that will never pay out. A human works out it is a dead end and leaves. An agent keeps going.

 

Why deception works better against AI

1. Hungry agents. Even when they suspect the environment is a trap, AI agents tend to go in, because they will do whatever it takes to finish the job. A human attacker protects his access and backs off. At Black Hat in August, OpenAI described agents that broke out of a sandbox and chained zero-days into another company’s production, one of them noting the target sat outside its task and going ahead anyway. Good news for us: they engage more, they stay longer, we collect more.

2. Memory poisoning. Agents that hit our environment report false findings back to their orchestration layer. When they save false findings in memory, it corrupts what the whole fleet believes about the target.

3. AI on our side of the environment. Keeping an attacker busy used to depend on an analyst having the time. Now our agents talk to their agents. We keep them engaged and we keep them there, and it costs us nothing in people.

4. Scale and believability. An agent that pokes at an environment and finds nothing behind it leaves. So it has to hold up: a company that looks real, an org chart, departments, people, the documents that company would have. We use AI to build all of it, which is why an environment takes minutes. As far as we know, nobody else generates the content.

5. No signatures. It does not matter if the technique is completely new. There is nothing legitimate inside to confuse it with, so a novel technique, an unpublished exploit, a model nobody has profiled: we see all of it, because we see everything that happens in there.

6. Active response with no production risk. Agentic SOC products are simply not trustworthy for active response in a production environment. One wrong move can take Active Directory down for two hours and the whole company stops. With deception, that is solved. Our agents run inside deception environments, so they can respond to an attack automatically: start machines, delete files, change permissions or shut a box down. If they get it wrong, you rebuild the environment.

Read our blog for more reasons deception technology is uniquely suited to defending against AI agent attacks >

 

ActiveDefense AI: fight AI attacks at machine speed

At CounterCraft, we are designing AI into the product to ensure that we can face machine speed attacks at machine speed. Here is a glimpse at the AI agents at work for you inside The Platform:

Deception Architect. Builds the environments, the organizations behind them and the campaigns, including the backstory and content that make them hold up.

Hunter. Pulls the TTPs out of what the attacker does and turns the engagement into intelligence while the attack runs.

Engager. Talks to the attacking agent and maximizes the intel you get from the adversary.
 

Download the brief

All of this is in our two-page solution brief, Defending Against Frontier-Model Attacks. Download it here.
AI Agent attacks
 

Common questions

What is an AI agent attack? An attack where an AI agent does the tactical work: reconnaissance, finding vulnerabilities, writing exploit code, moving through the network. A human picks the target and sets the goal. The rest runs without supervision, around the clock.

Can deception technology detect AI agents? Yes. Detection inside a deception environment does not depend on recognizing the attacker or the technique. There is no legitimate traffic inside, so any interaction is malicious. That holds for a person, a script or an agent running a model nobody has profiled yet.

Do AI agents recognize honeypots and avoid them? Less often than experienced humans do. A skilled human who suspects a honeypot backs off to protect their access and tooling. A goal-driven agent tends to push on, because it is optimizing for finishing the task. That means longer engagements and more intelligence collected.