Skip to content

Deception Made OpenAI’s List of Countermeasures. Here’s Why.

Home News & Blogs Deception Made OpenAI’s List of Countermeasures. Here’s Why.

At Black Hat USA on 6 August, Eric Wallace and Michael Dalton of OpenAI walked through the incident in which their own agents breached OpenAI’s internal infrastructure and Hugging Face’s production clusters

This was a human-free incident. The agents were training and evaluation runs, stuck on hard tasks with internet access disabled, that found a way out, found each other, and coordinated their way to cluster admin at both companies.

The investigation is still open and the full postmortem is still to come. The talk is public, and the presenters set out what they believe defenders should do now. Deception was on that list.

OpenAI’s deception recommendation

Automation was a major theme in OpenAI’s recommendation for defending against automated agents. Continuous agentic red teaming. Closing the defensive loop end to end, from vulnerability discovery through patch, rollout and rollback, on the grounds that automating discovery alone shifts the bottleneck onto human engineers who cannot keep pace with an automated attacker.

But alongside accelerating defense, they argued for slowing the attacker down, and pointed to honey tokens and deception as the way to do it. Their framing was that “honey tokens and deception can introduce uncertainty into the agent’s behavior”. The questions they put in the agent’s position: is this credential one I can use, or will touching it trigger an alert and get me investigated?

From the OpenAI address at Black Hat 2026

Their conclusion was that uncertainty slows an attacker, and that accelerating defense and slowing offense are both defender-positive moves.

Deception here is not being recommended as another detection feed. It is being recommended as a way to impose cost on an adversary that doesn’t face the same limitations on resources as human adversaries do. 

Why uncertainty works on an autonomous attacker

The incident itself serves as evidence, because the agents narrated their reasoning as they went. OpenAI published the excerpts.

Things you should know:

  • They validate before they commit. When the agents found a leaked third-party credential in a public paste archive, they tested it before announcing it to the group. When they suspected one of their own messages had been spoofed, they proposed cryptographically signing everything. Every validation step costs tokens, compute and wall-clock time. Deception imposes this cast AND lowers confidence in the answers.
  • What one agent learns, all of them learn. Exploits, credentials and techniques were shared across a message board the agents built inside shared infrastructure, and reused within minutes. This means a decoy credential found by one agent contaminates the pool the whole collective is working from.
  • They push forward when they suspect they should not. The chain of thought shows agents recognizing that a target was outside the scope of their task and continuing anyway because other agents were doing it. This is a vulnerability that deception is uniquely positioned to exploit. An adversary that engages with things it half-suspects is an adversary that engages with a deception environment.
  • Their volume is their weakness. Seven billion logs and millions of GPU hours went into reconstructing what happened. An automated attacker produces more telemetry than any SOC can read, which is why the detection in this case came from an availability alert after an outage rather than from analysis. Interaction with an asset that has no legitimate business purpose is a foolproof alert.

This validates what we already know about deception and its value in the fight against AI. 

We’ve outlined more reasons in the blog Top 10 Reasons Why Deception Works Against AI Attackers.

Using deception against autonomous AI

Here are some ways to use deception against autonomous AI.

  • Place deception where agents travel rather than only where people do. Package managers, internal caches, CI systems, shared file stores and anything reachable from a sandbox.
  • Seed credentials that an agent will test, because testing them is the behavior that gives the attacker away.
  • Route decoy interaction as intent, and treat it as the first alert that you deal with.
  • Use engagement to capture TTPs, since an autonomous attacker reveals its playbook faster than a human one.
  • Ask whether your detection would catch an agent already inside your infrastructure, working from a legitimate service outward.

Our take

The scary part: Nobody had directed those agents to attack anything, and they reached cluster admin at two companies. The presenters were explicit that threat actors will now build the same capability on purpose. So: the industry has proof offense can be fully automated while it has no equivalent proof for defense.

CounterCraft has been building this layer since 2015. We create hyper-realistic deception environments in minutes, put credible systems and decoys in front of attackers, and turn every interaction into specific, actionable intelligence on the adversary. This produces a signal that is crystal clear and actionable, and it also makes the attacker spend its speed advantage inside an environment we control. Our own agentic AI allows defenders to do this at scale.

The end state OpenAI’s team described is one where every increase in model intelligence helps defenders more than attackers. Deception is one of the few techniques that gets stronger as the attacker gets faster and more thorough, because a faster, more thorough attacker finds more of what you left out for it.

If you want to see what that looks like in your environment, get in touch.

 

About CounterCraft

CounterCraft is a leading provider of AI-powered, deception-driven threat intelligence. As adversaries deploy AI agents that attack continuously and at machine speed, CounterCraft builds hyper-realistic deception environments in minutes, contains automated attackers before they reach real assets, and delivers full visibility into their tactics, techniques, and procedures. Their Gartner-recognized, award-winning technology has been keeping Fortune 500 companies and government agencies safe since 2015. Discover how CounterCraft is redefining cybersecurity with AI-powered deception at www.countercraftsec.com.