AI Is Entering the Physical World: Cybersecurity Must Change Now

AI Is Entering the Physical World: Cybersecurity Must Change Now
Image generated by ChatGPT, 2026

Part 2 of “AI Is entering the physical world”

Considering everything covered in part 1 of this article, it’s time to explore the importance and relevance of the adversarial mindset.

The Adversarial Mindset Must Extend to Perceived Reality

Security teams cannot limit questions to something like:

Does the AI accurately understand the environment?

Instead, they should ask:

How could I make the AI misunderstand the environment while leaving it confident that its understanding remains accurate?

That shift produces very different security exercises.

For example, the following questions start to gain power:

  • Can conflicting sensor states exist?
  • Can temporal relationships between events be manipulated?
  • Can a legitimate device report technically valid but physically impossible values?
  • Can one AI agent be made to trust information supplied by another compromised agent?
  • Can some AI technology be fooled into selecting a dangerous action that still appears rational?
  • Can several individually low-risk inputs be influenced such that their combined effect changes some system’s interpretation of reality?

This goes far beyond vulnerability scanning.

It requires an understanding of the target environment and the making of educated (adversary informed) attacking assumptions.

Identity Becomes Even More Important in the Physical World

Interestingly, physical AI also amplifies the importance of identity. This is important given that many of the older OT protocols have no notion of a user or identity. Think about that, in many cases older ecosystems would allow network traffic carrying commands that could create physical impact. Yet, no authenticated user was part of the equation. We must do better now.

Every participant in a physical environment needs an identity or an attributable source of authority. That includes people, sensors, robots, cameras, controllers, applications, models, workloads, autonomous agents, and external systems.

The AI must know more than what information arrived.

It needs context about where that information came from, which identity produced it, and whether the organization should trust that source.

Likewise, when some AI technology decides to act, the receiving system needs to understand the authority behind that action.

Organizations should preserve a hierarchical chain such as:

  • Human owner
  • AI system
  • Model decision
  • Agent identity
  • Delegated authority
  • Physical command
  • Machine action

When anything within that chain breaks, accountability breaks with it.

More importantly, security loses the ability to determine whether something with legitimate authority produced some physical action.

In IT, a compromised identity can expose information or disrupt systems.

In physical AI, a compromised identity may eventually move something.

Cybersecurity Must Now Protect State, Not Just Systems

As physical AI develops, cybersecurity architecture will need to focus increasingly on state. Some of the types of questions that need answers as they relate to state are:

  • What is true right now?
  • Which entities exist at the moment?
  • What are these entities doing?
  • Which relationships connect them?
  • What actions have already occurred?
  • Which future states remain plausible?
  • What authority exists to change the current state?

And critically:

How confident are we that the data describing current state is trustworthy?

Traditional alerts often examine individual events.

Physical AI security will need to understand sequences, relationships, causality, and physical context.

Some examples are:

  • A temperature reading of 190 degrees may be safe in one operating state and extremely dangerous in another.
  • A valve opening may be normal after one event and malicious after another.
  • A robot entering an area may present little risk until a person enters the same physical space.

In the physical domain, context determines risk.

Therefore, security platforms will need stronger temporal models, dynamic graphs, event streams, behavioral baselines, identity relationships, process awareness, and state prediction.

The goal cannot remain simply detecting what has already happened.

We need to understand what is happening, why it is happening, and what is likely to happen next.

This represents another important shift for cybersecurity.

Historically, security operations have been overwhelmingly reactive. An event occurs, a signal appears, an alert fires, and analysts investigate. The entire incident response industry exists because of this reactive model.

Physical AI will demand more predictive security.

If some AI technology controlling or influencing an environment can reason about what happens next, defenders must develop comparable capabilities to identify dangerous future states before systems can reach them.

The objective becomes more than detecting malicious activity.

It becomes preventing the environment from reaching an unsafe state.

Six Security Principles for Physical AI

This is not television, and it will be a bit before humanoid robots begin to appear throughout the enterprise. Cybersecurity leaders do not need to wait for that day before they start preparing. That preparation can begin now and here are six relevant suggestions:

1. Protect the Data That Defines Reality

Identify the data that physical AI systems use to understand their environment.

Establish provenance, integrity controls, behavioral baselines, cross-source validation, and clear ownership. This data needs to be protected as it will be the basis of important truths.

Furthermore, treat manipulation of physical-state data as a high-consequence security event.

We have spent decades protecting sensitive data from exposure. Physical AI requires equal attention to protecting data from malicious alteration.

2. Understand Semantics, Not Just Traffic

Do not stop at network visibility as that is simply not enough.

Understand what commands and values actually mean to consuming physical processes.

This was central to the data-centric approach I advocated in OT security years ago, and it becomes even more important when AI consumes that information to understand some environment.

Allowed communication and safe action are not synonymous.

3. Bind Identity to Physical Authority

Every human and non-human actor capable of influencing the physical environment needs an attributable identity, constrained authority, and accountable owner. We have to do better than the OT protocols of the past where no identity was bound to commands and changes flowing via network communications.

Organizations must know who or what caused every consequential action.

They also need to continuously evaluate whether that identity remains trustworthy.

4. Model the Blast Radius Before Granting Autonomy

Before giving an AI system authority to act, determine what happens if it makes the wrong decision. This requires proper testing, consideration of edge cases, and careful attention to the design and enforcement of boundaries.

Ask how far one incorrect action can propagate through interconnected machines, systems, and physical processes.

Then constrain autonomy accordingly.

The greater the physical consequence, the smaller the acceptable gap between authority and accountability.

5. Use Simulation as a Security Tool

Digital twins and simulated environments should do more than optimize operations or train models.

Security teams can use them to test adversarial scenarios, evaluate “what-if” conditions, attempt to predict attack paths, and observe potential physical consequences without endangering production environments.

However, teams must also secure the simulation itself.

If the digital twin becomes an input into training, planning, or decision-making, poisoned simulation data can eventually influence downstream real-world behavior.

6. Design for Safe Failure

Every physical AI system needs an independently enforceable path to a safe state.

Security teams should be able to dynamically revoke authority, isolate compromised components, reject untrusted data, switch to manual control, and stop physical action.

Most importantly, do not assume that the AI responsible for normal operation should also control its own emergency containment.

Leadership needs to Understand the Physical AI Transition

Boards and executive teams do not need to become experts in AI technologies or the designing of world-class architectures.

However, they do need to understand what happens when AI crosses the boundary between recommendation and action as that can have a direct impact on business operations.

Here are a few questions leadership should start asking:

  • Where can AI already influence physical processes in our organization?
  • Which AI systems can issue commands that trigger physical action?
  • Which data sources clearly define their understanding of physical state?
  • Can we establish the integrity and provenance of that data?
  • Can data manipulation create unsafe environments?
  • Which human and machine identities possess authority over systems that have physical capabilities?
  • Have we tested how the systems respond to intentionally manipulated environments?
  • What physical consequences could follow an incorrect decision?
  • Can we quickly force the environment into a safe state when trust disappears?

These are not robotics questions.

They are enterprise-risk and governance questions.

We Have Seen Part of This Future Before

World models and physical AI introduce powerful new technology.

Yet one of their central security problems brings me directly back to the OT environments we worked to protect at Bayshore Networks.

In 2020, I argued that protecting industrial environments required us to understand more than who communicated with whom.

We had to understand the actual values moving through industrial protocols.

Then we had to understand what would happen in the physical world when a PLC, controller, drive, or other system acted upon those values.

That progression was:

Data → Context → Command → Physical Consequence

World models extend that to:

Data → Perceived Reality → Predicted Future → Decision → Physical Consequence

That additional intelligence does not eliminate the old security problem.

It magnifies it.

World models will increasingly use huge volumes of data to construct representations of reality, predict future states, and select actions.

Therefore, cybersecurity must protect much more than the model.

We must protect the integrity of the world the model believes it inhabits.

That requires trusted data, attributable identities, semantic understanding, adversarial testing, state awareness, predictive security, and tight control over physical authority.

The GenAI era taught organizations that machines can create.

The agentic AI era is teaching us that machines can act.

World models and physical AI will force us to confront the next question:

What happens when machines can understand enough of the physical world to predict it, and possess enough authority to change it?

Cybersecurity leaders should start considering that question now.

Because AI is entering the physical world.

And once cyber risk becomes physical risk, we no longer get to treat a corrupted view of reality as merely a bad AI output.

AI Is Entering the Physical World: Cybersecurity Must Change Now

AI Is Entering the Physical World: Cybersecurity Must Change Now
Image generated by ChatGPT, 2026

Part 1 of “AI Is entering the physical world”

For the last several years, most organizations have experienced Artificial Intelligence through a screen. Come to think of it, so have many of the recently self-appointed AI experts. I consider most of these people users, not experts. Things are changing on levels these folks are not prepared for. AI Is entering the physical world. Why cybersecurity must change now.

Typical “AI” usage at the moment equates to Generative AI (GenAI). This means someone types a prompt and an engine generates content. The engine can write code, analyze a document, create an image, summarize data, or recommend an action.

That model of AI is already changing.

The next major evolution will push AI beyond understanding language and digital information. AI systems will increasingly model environments, predict how those environments may change, reason about physical objects, and take actions in the real world.

World models, embodied AI, robotics, autonomous systems, digital twins, and increasingly capable agentic ecosystems are moving us in that direction.

Consequently, cybersecurity leaders need to understand that this transition changes the security problem dramatically.

Potential Physical Impact

When AI exists primarily inside a digital environment, a bad decision may generate incorrect information, expose data, execute malicious code, or compromise a business process.

When AI can perceive and act upon the physical world, a bad decision can move a machine.

It can alter a manufacturing process or the behavior of a robot. It can influence an autonomous vehicle or manipulate an industrial control process.

Ultimately, it can create physical consequences.

That is why cybersecurity must change. Now.

Traditional cybersecurity primarily protects systems, identities, networks, applications, and information. Physical AI adds something fundamentally different. In most cases, foreign. Cybersecurity must now protect an AI system’s perception of physical reality, the data used to construct that reality, the authority to act upon it, and the resulting physical state.

An attacker may no longer need to compromise the AI model itself.

Manipulating the world the model sees may be enough.

That shifts cybersecurity beyond protecting systems and information. We must increasingly protect state, perception, prediction, authority, and physical consequence.

Physical AI changes cybersecurity because we must protect not only the AI, but the integrity of the world the AI believes it inhabits.

World Models Change What AI Understands

Large Language Models (LLMs) became powerful by learning relationships across enormous amounts of data.

World models pursue a different capability.

At a high level, a world model develops a representation of an environment and uses that representation to reason about how the environment may change over time.

Instead of merely asking, “What should come next in this sequence?” as LLMs do, world model based systems begins answering questions such as:

  • What exists in this environment?
  • How are these objects related?
  • What state are they currently in?
  • What happens if something moves or changes?
  • What will the environment probably look like next?
  • How will my actions affect this environment?
  • What action is necessary to reach a desired state?

This capability matters enormously for robotics and autonomous systems.

For example, a robot operating in a warehouse cannot simply identify a forklift. It needs to understand where the forklift is, is it currently being operated, how quickly it is moving, where it will probably go next, and what obstacles surround it.

Likewise, an industrial AI system cannot simply recognize that a valve exists. It may need to understand the valve’s current state, its relationship to pressure elsewhere in the process, what normally happens after the valve changes state, and which physical consequences could follow.

In other words, the AI must build and continuously update a representation of reality.

That representation becomes extraordinarily valuable.

It also becomes an extraordinarily attractive target.

Why Cybersecurity Must Change When AI Becomes Physical

Cybersecurity traditionally focuses on protecting identities, systems, applications, networks, APIs, and data.

Physical AI forces us to extend that thinking.

We now have to protect the system’s understanding of reality.

If an attacker manipulates the information an AI system uses to construct that reality, the attacker may never need to compromise the model itself. That is a dynamic the security industry has yet to contend with.

Consider an autonomous system that continuously processes sensor readings, environmental conditions, machine states, visual information, historical behavior, operator commands, and other telemetry.

The AI uses those inputs to determine what exists, what is happening, what will probably happen next, and what action it should take.

Now change one of those inputs.

Then change several.

Make the changes subtle enough that no individual result looks catastrophic.

An attacker can gradually create a false version of reality inside that target system. If the approach is slow and low the end result can be rather complex.

Along that journey, AI systems could make completely rational decisions based on completely corrupted context.

The model did not necessarily fail.

Its understanding of the world failed.

That distinction will become one of the defining problems in physical AI security.

I Wrote About This Problem Before World Models Entered the Conversation

This problem feels new to many because technology has changed and those people have likely not dealt with these types of environments.

But, to some of us the underlying security principle is not new at all.

In January 2020, while I was one of the original members and CTO at Bayshore Networks, I published an article in Network Security titled “Operational Technology Security – A Data Perspective.” (https://www.sciencedirect.com/science/article/abs/pii/S1353485820300088)

The central argument was straightforward: OT cybersecurity was concentrating too heavily on network-level visibility while overlooking something far more consequential – the actual values inside the data.

Knowing the following mattered:

  • That one IP address communicated with another.
  • Which network protocol was used.
  • That a particular workstation communicated with a Programmable Logic Controller (PLC).

However, none of those facts necessarily told us what happened to the physical process.

For that, we had to understand the data itself. We needed to understand the command, the register, the setpoint value.

Most importantly, we needed to understand what changing certain values would mean in the physical domain.

That was the data-centric security problem in OT. To an extent that is still a problem today.

An attacker did not necessarily need to break the network connection. The connection could remain completely legitimate.

An authenticated engineering workstation could communicate with an approved controller over an expected industrial protocol.

Yet if the attacker changed the right value inside that legitimate communication, the physical result could become dangerous.

In OT, the packet can be legitimate while the value inside it is hostile.

That concept drove much of the thinking behind the technology we built at Bayshore Networks.

We pushed inspection beyond basic network metadata and deeper into industrial protocols, transactions, commands, and values. We wanted security controls to understand what the industrial communication meant, not simply observe that the communication occurred.

Why?

Because data was not simply information.

Data could become physical action.

World Models Extend the Data-Centric OT Problem

This is where my earlier OT work and today’s world-model discussion converge.

The problem I described in 2020 focused on protecting data values because industrial systems could act upon those values with potential physical impact.

World models take that concept significantly further.

A physical AI system does not simply receive a single value and execute a command. Increasingly, it will consume enormous amounts of data to construct an internal representation of its environment.

It will correlate inputs, infer relationships, estimate current state, and predict future state.

Then it may select an action based on that representation.

Therefore, take the old OT question: what does this data value mean to the physical process? This now becomes an even more consequential AI security question: what reality is this data causing the AI to believe?

That is the intellectual bridge between data-centric OT security and physical AI security.

In the OT environments we protected years ago, manipulating a register or setpoint could change a physical process.

In a world-model-driven environment, manipulating enough trusted data could change the AI’s model of the entire process.

At that point AI itself may determine which action should follow.

This gives the adversary an entirely new level of leverage.

Data Becomes Part of the Physical Control Surface

Security leaders need to stop thinking about data solely as something they must prevent attackers from stealing.

In physical AI environments, adversaries may gain more value by changing data than by stealing it.

Consider the information that could feed a physical AI system:

  • Camera and video streams.
  • Temperature and pressure measurements.
  • Location and proximity data.
  • Motor position and velocity.
  • Machine state.
  • Human presence and movement.
  • Production telemetry.
  • Historical operating patterns.
  • Digital-twin state.
  • Maintenance information.
  • Identity and authorization context.
  • Commands from machines or other AI agents.

If some AI technology uses those inputs to determine current state, predict future states, and choose actions, the integrity of those data points becomes part of the physical control surface.

As a result, security teams must ask something much more sophisticated than:

Can this system access the sensor?

They must ask:

Should the system trust what that sensor is telling it right now?

That requires context. As such, questions such as these become very relevant:

  • Is this the expected device?
  • Has the configuration been modified?
  • Does value X make sense within the current operating state?
  • Did an authorized identity make some change?
  • Does the sequence of events match expected process behavior?
  • Could the data be syntactically valid but operationally impossible?

This is precisely where the lessons from my OT past become invaluable.

Seeing the communication is not enough.

Understanding the data is not enough.

Security must understand the data in the context of the physical process, and now in the context of some AI’s evolving model of that process.

A Hostile World Does Not Require a Compromised Model

Much of the current AI security conversation concentrates on attacking models.

We discuss prompt injection, jailbreaks, model theft, adversarial inputs, training-data poisoning, and manipulated outputs.

Obviously, those threats remain important.

However, physical AI creates another powerful adversarial strategy:

Do not attack the intelligence. Attack the world that the intelligence sees.

An adversary could target:

  • Perception – change what sensors, cameras, or other inputs report.
  • State – alter the data describing the current condition of a machine or environment.
  • History – corrupt the historical context the system uses to recognize normal behavior.
  • Identity – impersonate a trusted operator, sensor, machine, or workload.
  • Relationships – manipulate the system’s understanding of how physical entities depend upon one another.
  • Prediction – distort enough contextual information to influence the system’s expected future state.
  • Action – abuse the mechanism that translates AI decisions into physical commands.

This attack model should concern cybersecurity leaders because the attacker can work around the intelligence rather than directly against it.

Imagine an AI system correctly concluding:

Given everything I currently know about this environment, action X represents the safest response.

Now imagine that an adversary manipulated what the system knows.

The reasoning may remain sound, but the action can still become dangerous.

See part 2 of this write-up here.

AI Is Undeniably Weaponized Now. The Human Is the Adversary.

AI Is Undeniably Weaponized Now. The Human Is the Adversary.

Artificial Intelligence (AI) is undeniably weaponized now. But the human is still the adversary. AI changes the speed, scale, sophistication, and autonomy of cyberattacks, while in most AI-enabled attacks a human still defines the objective, determines the desired outcome, directs or delegates activity to the technology, and benefits from success.

AI has changed cybersecurity at extraordinary speed. Attackers now use AI as both a force multiplier and a capability multiplier. They can accelerate reconnaissance, generate and refine malware, build highly targeted phishing campaigns, impersonate executives, analyze enormous volumes of stolen data, discover relationships between data points, identify exploitable weaknesses, and increasingly execute sequences of actions through autonomous agents.

Yet those capabilities do not eliminate the human element. AI may execute the action. An agent may navigate the application. A model may create the campaign material. But behind most malicious AI activity, a human still defines the objective, decides what outcome matters, and benefits when the operation succeeds. Last I checked there wasn’t some AI technology cashing out some Bitcoin from a ransom and partying on a yacht.

Consequently, understanding The Adversarial Mindset matters more today than in the past.

Does AI Eliminate Human Intent From Cyberattacks?

No, AI does not eliminate human intent from cyberattacks. It can dramatically change how an attack is executed while a human adversary still defines the objective the technology is pursuing.

Too often, it feels like we talk about AI-powered attacks as though AI itself has suddenly become the adversary.

That framing can be misleading.

Consider the difference between traditional Generative AI (GenAI) and Agentic AI.

With traditional GenAI, the relationship remains relatively obvious. A human asks a model to do things such as identifying vulnerabilities, improving code, analyzing data, translating messages, performing research, or solving some other element of an operation.

The system provides the power. The human provides the objective.

Agentic AI creates more distance between those two elements.

Instead of asking AI to perform one task, a human can increasingly define an objective and allow an agent to determine how to accomplish it. The agent can browse websites, invoke tools, query data, make decisions, evaluate responses, select subsequent actions, and continue working toward a defined goal.

In other words, the human moves farther away from each individual action.

However, distance from execution does not automatically remove intent.

That distinction matters enormously for cybersecurity.

An attacker does not need to personally enumerate every endpoint, craft every request, write every line of malicious code, or send every social-engineering message to remain the adversary behind an operation.

AI gives that nefarious actor both abstraction and leverage.

Agentic AI gives that same human a certain level of delegation.

Neither automatically removes the human from the equation.

Who Is Acting When an AI Agent Accesses a Computer?

When an AI agent accesses a computer on a user’s behalf, the human user can remain the party performing the access. In the Ninth Circuit’s August 2026 Perplexity decision, the court treated the AI assistant as a tool and the human user as the party accessing Amazon’s systems for purposes of the federal Computer Fraud and Abuse Act (CFAA).

The dispute involved Perplexity’s Comet browser and its AI Assistant. Users could direct the Assistant to perform tasks on Amazon.com. Amazon argued that Perplexity’s technology accessed Amazon’s systems without authorization and sought relief under the CFAA, and its California counterpart (the Comprehensive Computer Data Access and Fraud Act – CDAFA).

The Ninth Circuit rejected Amazon’s theory at the preliminary-injunction stage.

More importantly, the court focused on a remarkably significant question:

Who actually accesses the computer?

On the record before it, the court concluded that the AI Assistant functioned as a tool. The court described the Assistant as a “tool, not a person for statutory purposes.” It then concluded that the user accessed Amazon’s computers while using the Assistant to carry out specific actions.

The decision marks the first federal appellate ruling addressing whether AI agents acting on behalf of users can legally access online platforms.

That distinction carries enormous significance beyond this particular dispute.

The court did not treat the AI agent as an independent legal actor simply because it could perform actions on behalf of a user. Instead, it looked through the technology to determine who actually performed the access for purposes of the statute.

At the same time, something important surfaced by way of a limitation.

The Ninth Circuit DID NOT create a sweeping legal doctrine that makes humans universally responsible for everything an AI system does. In fact, the opinion expressly states that it does not establish a new legal regime for agentic AI. The court limited its holding to the CFAA and CDAFA “access” issue, the technology at issue, and the factual record before it. Different facts, different levels of control, different laws, or different AI architectures could produce different outcomes.

Nevertheless, from a cybersecurity perspective, a much broader lesson remains powerful: technology can sit between a human and some action without rendering the human irrelevant (or innocent by default).

Should Security Programs Defend Against AI or the Adversary?

Security programs should defend against the adversary, not AI in isolation. AI mechanisms such as prompt injection, model poisoning, tool abuse, and MCP attacks matter, but they do not explain who wants to attack you, why they are targeting you, or how they will adapt.

The industry has become obsessed with AI security. Both RSAC and BlackHat this year showcased that obsession with great fanfare.

To answer the questions of who, why, and how, you need to understand the adversary, not just the technology at hand.

For example, imagine two attackers with access to exactly the same AI model and exactly the same agentic capabilities.

One is a teenager experimenting, testing boundaries.

The other operates inside an organized cybercriminal enterprise with millions of stolen identities, infostealer logs, credential collections, years of operational experience, and a clear understanding of how to monetize access.

The AI may be identical.

The threat is not.

The adversary behind the technology creates that difference.

Should Analysts and Frameworks Define a Security Program?

No, analysts and frameworks should not define a security program or its security strategy. They can inform both, but market intelligence about technologies, vendors, categories, and industry trends is not the same as understanding the adversary targeting your organization.

An industry analyst publishes some analysis. Vendors push categories. A maturity model emerges. Boards ask where the company sits relative to peers. CISOs then purchase technologies to fill perceived gaps. Eventually, the organization builds an architecture that looks remarkably similar to the architectures of dozens of other companies that consumed the same analyst research. And along the way end up with tons of tools whose true capabilities are not fully utilized.

To be clear, industry analysts provide value.

They can deliver market intelligence, technology comparisons, vendor analysis, spending benchmarks, maturity models, and useful observations about where the industry is heading.

However, organizations make a serious mistake when they use analyst research as the foundation of a security program.

Market intelligence is not adversary intelligence.

An analyst may understand the cybersecurity industry exceptionally well while possessing little firsthand understanding of the people trying to defeat your security program.

They may understand industry sectors, products, categories, vendors, differentiators and even what other CISOs are spending on.

Yet none of those things necessarily means they understand how a real adversary thinks.

More importantly, an adversary does not care whether your program aligns with an analyst’s reference architecture.

The adversary cares whether your defenses prevent the desired outcome.

Therefore, security leaders should never stop at this question:

What does the industry say a modern security program should contain?

They must also ask:

If I were a competent, cunning, determined attacker targeting this organization, how would I defeat what we have built?

What Blind Spot Do Many CISOs Have?

The blind spot many CISOs have is a limited understanding of the real adversaries their security programs are supposed to defeat. Managing risk, compliance, architecture, technology, and incident response is not the same as understanding how a determined adversary thinks, adapts, combines weaknesses, and pursues an objective.

I am not referring to understanding ethical hackers, penetration testers, or red-teamers. These professionals absolutely add value, but they ultimately operate within constraints established by rules of engagement.

I am talking about a real adversary with ill intent, whose motivations may be financial, ideological, geopolitical, personal, or simply opportunistic; and who feels no obligation to respect rules, scope, policy, business hours, budgets, architecture diagrams, or organizational boundaries.

That distinction matters.

A penetration tester typically asks whether something can be compromised within an agreed scope.

An adversary asks a very different question:

How do I achieve my objective despite everything this organization has done to stop me?

That question requires a fundamentally different way of thinking.

Why Should Defenders Start With the Human Behind the Machine?

Defenders should start with the human behind the machine because AI amplifies adversarial capability without automatically replacing adversarial intent. Less sophisticated attackers can now access capabilities that once required specialists, while sophisticated adversaries can operate faster, analyze more data, uncover hidden relationships, and adapt more efficiently.

As an example, consider that AI technologies create conditions in which enormous quantities of stolen identity data can be ingested and analyzed, revealing relationships humans would otherwise miss.

It can perform actions such as:

  • transforming OSINT into targeted and strategic intelligence
  • generating individualized social-engineering content based on attackable profiles across thousands of targets
  • creating strategic campaigns rapidly
  • refining malicious code
  • analyzing defensive responses and adaptively creating alternatives

But it does so at the request of some human element. Consequently, we should stop thinking only in terms of “AI attacks.” What we increasingly face are human adversaries with machine-scale leverage.

That represents a much more consequential problem.

How Does The Adversarial Mindset Change Security Strategy?

The Adversarial Mindset changes security strategy by making the adversary, not the framework, product, analyst, or compliance requirement, the starting point. Security leaders first ask what an adversary wants, what that adversary already knows, which assumptions and relationships can be exploited, and how the attacker will adapt when defenses interfere.

Ask questions like:

  • Who would want what we possess?
  • What exactly would they want?
  • What information about our people, systems, suppliers, executives, and customers do they already possibly have?
  • Which assumptions are we making that they would immediately challenge?
  • Where do identities, relationships, privileges, and trust create nefarious opportunities?
  • How could they combine several individually minor weaknesses into one viable attack path?
  • How would they adapt after encountering resistance to their techniques?
  • How could AI make each of those steps of adaptability cheaper, faster, or more precise?

At that point, you begin designing security from the adversary backward.

That is the essence of The Adversarial Mindset.

Moreover, this approach does not require organizations to abandon frameworks, compliance obligations, analyst research, or established security architectures. Those tools still serve important purposes.

However, they should support your security strategy rather than define it.

The adversary should help define it.

Why Does AI Make The Adversarial Mindset More Important?

AI makes The Adversarial Mindset more important because it gives human adversaries greater speed, scale, precision, leverage, and increasingly autonomous execution. Security teams therefore need to understand not only what AI can do, but what a motivated adversary can now accomplish because those capabilities exist.

The cybersecurity industry will inevitably spend enormous amounts of time debating how autonomous AI will become.

That discussion absolutely matters.

Eventually, increasingly autonomous systems may force us to confront genuinely difficult questions about intent, accountability, responsibility, control, and attribution.

However, that conversation leaves gaps. Organizations cannot afford to wait for those philosophical and legal questions to reach resolution. This is especially so for larger organizations that are not exactly agile.

Today, humans are discovering what AI can do for them.

Some of those humans are defenders, others are researchers, and still others are innovators.

Realistically, some are adversaries.

The last group does not care whether your AI strategy appears in an analyst report. They do not care which security technologies occupy a leader quadrant, or how mature your program looks against some industry benchmark.

They care whether they can accomplish their objective.

The Ninth Circuit’s Perplexity decision gives us an important legal manifestation of a broader technological reality: an intelligent tool can become increasingly capable while still operating in service of human direction.

Therefore, defenders should resist the temptation to focus exclusively on the technology.

That distinction changes the questions security leaders should be asking.

“What can AI do?”

This has to start migrating towards something like:

“What can an adversary now do because AI exists?”

Those are very different questions.

Ultimately, the second question is the one our security programs need to answer. AI is undeniably weaponized now. The human is the adversary.


Note: This article discusses the cybersecurity implications of Amazon.com Services, LLC v. Perplexity AI, Inc. and does not provide legal advice. The Ninth Circuit’s August 4, 2026 decision concerned a preliminary injunction and a specific interpretation of “access” under the CFAA and CDAFA based on the record before the court.

Reasons AI Governance Fails Without The Adversarial Mindset

AI Governance Fails Without The Adversarial Mindset
Image generated by Jetpack AI, 2026, via WordPress

I genuinely believe that most AI governance programs begin with good intentions. But that doesn’t instantly equate to a successful program. Often, AI governance fails without the adversarial mindset.

These types of governance programs typically define acceptable use, establish review committees, classify risk, document models, assign owners, and publish principles around fairness, privacy, transparency, and human oversight.

Undoubtedly, those activities matter.

Those programs also tend to assume that people, systems, data, and models will operate within the boundaries the organization designed.

An adversary makes no such assumption. Moreover, the boundaries organizations have designed mean nothing to an adversary.

Attackers will generally search for paths of least resistance that lead to success. This includes ways to manipulate inputs, compromise identities, poison data, exploit integrations, misuse legitimate capabilities, and confuse and manipulate humans.

Employees, contractors, customers, partners, activists, fraudsters, competitors, and nation-state actors may all test the distance between what an AI system was intended to do and what it can be manipulated to do.

Governance that considers only intended behavior is policy.

Governance that anticipates intentional manipulation becomes resilience.

AI governance without an adversarial mindset documents how a system should behave. It does not prepare the organization for how the system can be made to behave.

The Adversary Has a Vote

Executives often discuss AI risk as though the organization controls all relevant variables.

Leaders choose the model. Engineers establish the architecture. Data teams manage information. Security implements controls. Legal writes policy. Users receive training.

Then the system enters the real world.

Customers provide unexpected inputs. Employees find shortcuts. Vendors go out of business. Models drift. Credentials become exposed. Attackers study architectures and systems. Data sources become contaminated.

Ultimately, the adversary gets a vote in how the system operates.

This principle has shaped cybersecurity for decades. A secure architecture cannot assume that users will follow instructions, data will remain trustworthy, or that controls will continuously operate exactly as designed.

AI governance must adopt a similar reality.

MITRE ATLAS (https://atlas.mitre.org/) documents tactics and techniques used against predictive, generative, and agentic AI systems. NIST has developed a taxonomy for adversarial machine learning (https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-2e2025.pdf). The NCSC’s secure AI guidance (https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines/secure-design) emphasizes threat modeling across design, development, deployment, and operation.

Those movements create a message that is consistent: AI risk cannot be governed solely through compliance reviews and intended-use documentation.

AI Governance Must Cover More Than Model Failure

Executives commonly focus on whether an AI system will produce an incorrect answer.

Yet that is only one failure mode.

Interestingly, an AI system may produce an accurate answer for the wrong person. It may also follow a valid command issued through a compromised identity or expose sensitive information while correctly completing a task. Sound recommendations may be generated based on poisoned data.

The model may operate exactly as designed while the larger system fails.

This distinction matters because AI is not merely a model.

It is an architecture of identities, data, software, infrastructure, integrations, humans, workflows, and delegated authority.

An adversary does not need to attack the most sophisticated component. The adversary will attack the component that produces the greatest advantage or outcome for the least effort.

That may very well be the model itself.

But, it may also be an exposed API key, overprivileged service account, manipulated document, compromised developer, insecure plugin, careless employee, or trusted third-party data source.

An Adversarial Mindset Is Not Merely Pessimism

Some leaders resist adversarial thinking because it can sound negative or obstructive. Others claim that doing so empowers and/or validates the adversary.

Those notions misunderstand the purpose of the adversarial mindset.

An adversarial mindset does not assume that every person is malicious or that every AI initiative will fail. It assumes that valuable systems will attract manipulation and that unintended behavior becomes more likely as complexity, scope, and authority increase.

It asks the organization to examine its assumptions before an attacker does.

This mindset challenges statements and/or beliefs such as:

  • Only employees can access the system.
  • The model does not have access to sensitive data.
  • A human reviews every important decision.
  • The agent can only use approved tools.
  • The training data comes from trusted sources.
  • We have a kill switch, and it works.

Each statement may be technically accurate while concealing dangerous assumptions. As such, further probing may look like:

  • Which employees?
  • Through which identities?
  • Does the human reviewer understand enough to challenge the model or system?
  • Can approved tools be combined to create an unapproved outcome?
  • Who determines whether a source remains trustworthy?
  • How long does shutdown actually take?

The adversarial mindset turns reassuring claims into testable questions.

Seven Questions Leaders Should Ask

1. How Could Someone Intentionally Misuse This System?

Governance reviews often begin with the approved business use case.

Adversarial governance begins with the abuse case.

Leaders should scrutinize the angles. Ask how an employee, customer, contractor, criminal, or competitor could use some legitimate capability for a different objective.

A customer-service assistant might help employees retrieve account information. Could it also help a malicious insider assemble customer profiles?

A fraud model might identify suspicious payments. Could someone probe its thresholds and learn how to avoid detection?

A security agent might isolate compromised systems. Could an attacker manipulate it into disrupting legitimate operations?

Security teams should create abuse stories alongside legitimate user stories. Every material use case should identify who may benefit from subverting it and which capabilities they would seek out in pursuit of that benefit.

2. What Happens When the Data Becomes Hostile?

Organizations tend to view data as an input.

Adversaries may use it as code, instructions, a weapon, or a persistence mechanism.

Manipulated training data can alter future model behavior. Poisoned reference material can corrupt retrieval-augmented systems. Malicious instructions embedded in emails, websites, documents, or images can influence agents that consume external content.

The source may appear trusted while the content is not.

Leaders should ask:

  • Which data sources can influence the system?
  • Who can change those sources?
  • How does the organization establish provenance?
  • Can the system separate data from instructions?
  • What happens when sources conflict?
  • How quickly can poisoned information be detected and removed?

Data governance must assess not only quality and privacy but also hostility.

3. Which Identity Creates the Greatest Blast Radius?

Eventually, many attacks turn out to have some intersection with identity.

A compromised developer can change code. A stolen service credential can invoke functionality. An overprivileged agent can call dangerous tools. An administrator can modify thresholds or do things like suppress logs.

Leaders should identify the human and non-human identities capable of influencing high-consequence AI systems.

They should understand:

  • Who can alter data, especially data used for training models.
  • Who can change policies or guardrails.
  • Who can deploy or replace models.
  • Which agents can invoke production tools.
  • Who can disable monitoring.
  • Which identities can approve their own changes.

The most dangerous identity may not possess the most obvious administrative title. It may be an automation account that quietly connects elements such as models, data, and production systems.

4. Can the Human in the Loop Be Manipulated?

Organizations frequently rely on “human in the loop” as their final safeguard.

Adversarial thinking asks whether the human loop actually works.

People may over-trust AI recommendations, especially when a system appears confident or technically sophisticated. Reviewers may approve outputs automatically because of enormous volumes of data to review, time pressure, weak interfaces, inadequate context, or fear of challenging a system the organization has heavily promoted.

An attacker may also target the reviewer, a human, directly. After all, whether purposely or not, humans are the source of many unfortunate cyber events.

Manipulated systems could present targeted evidence, conceal uncertainty, overwhelm the reviewer with volume, or frame the decision in a way that encourages the desired behavior (of the nefarious actor).

Meaningful human oversight requires authority, context, time, training, and psychological permission to disagree.

A person who can technically click “reject” but is organizationally discouraged from doing so is not an effective control.

5. What Happens When Trusted Components Become Untrustworthy?

AI systems depend on external data, models, open-source libraries, cloud platforms, plugins, APIs, and vendors.

Each dependency extends the trust boundary.

A provider may change a model. An integration may gain new capabilities, or lose existing ones. A library may become compromised. A data source may introduce manipulated content. A vendor may retain more information than expected.

Third-party risk assessments often occur before deployment and then fade into GRC oblivion.

Adversarial governance treats trust as temporary.

Leaders should know which external components can influence decisions or actions, what changes providers are making, how the organization detects those changes, and whether it can continue operating safely when the dependency becomes unavailable or untrustworthy. These are elements that make up a resilient ecosystem.

6. How Would We Detect a Quiet Failure?

Not every AI incident will produce blinking lights, loud alarms, an obvious outage or some catastrophic result.

Some of the most damaging failures will appear gradually.

A model may become slightly less accurate for a specific population. An agent may begin retrieving more data than it needs. A recommendation system may slowly favor manipulated content. A fraud model may become less sensitive to a criminal technique.

The system continues to operate, and traditional availability metrics remain healthy.

Leaders need indicators that reveal changes in context, behavior, authority, data access, confidence, exception rates, and human impact.

They should also monitor near misses. An action stopped by a human or compensating control still reveals a weakness in the system.

Simply put, governance cannot measure only uptime and adoption. It must measure whether the system remains within its intended behavioral boundaries.

7. Can We Contain the System Before We Understand the Incident?

Executives often assume that teams can shut down an AI system when something goes wrong.

That assumption should be tested.

A system may be embedded in business workflows. Multiple applications may depend on this integration. Agents may retain active sessions or credentials. Third-party components may continue processing information. Teams may hesitate because shutting some system down could create operational consequences.

Incident response requires the ability to reduce authority quickly, even before investigators understand the complete failure.

Organizations should be able to:

  • Revoke agent and service-account access.
  • Disable specific tools or integrations.
  • Quarantine suspicious data sources.
  • Roll back models and configurations.
  • Move automated decisions into manual review.
  • Preserve evidence for investigation.
  • Continue critical operations in a degraded mode.

The ability to stop an AI system safely needs to be a design requirement, not an emergency improvisation.

Turn Governance Into an Adversarial Operating Model

An adversarial mindset must produce more than provocative questions. When pursued properly this mindset should guide and mold entire security programs.

Organizations should embed this mindset into their operating model.

That includes:

  • Threat modeling – examine the complete AI architecture, not only the model.
  • Abuse-case development – document how legitimate capabilities could support illegitimate objectives.
  • Red teaming – without boundaries (attackers have none) test models, identities, integrations, users, and workflows.
  • Control validation – under realistic conditions, prove that guardrails work rather than accepting that they exist.
  • Behavioral monitoring – actively detect changes in data access, authority, tool use, and outcomes.
  • Incident exercises – regularly rehearse containment, rollback, investigation, communication, and recovery. This can follow the standard tabletop model even though many of those exercises introduce boundaries and constraints that take them out of the realm of realistic conditions.
  • Continuous reassessment – review risk whenever the system’s data, model, tools, users, or authority change.

Governance should operate as a feedback loop:

Assume → Challenge → Test → Observe → Adapt

Typically, a policy changes only when someone updates the document. An adversarial governance program needs to change when evidence reveals that an assumption no longer holds.

Policy Describes the Organization You Hope Exists

AI governance policies describe approved behavior. They define responsibilities, expectations, controls, and boundaries.

These definitions are necessary.

Sadly, they are not the same as actual readiness.

The real organization includes shortcuts, legacy access, human bias, compromised credentials, conflicting incentives, third-party dependencies, weak integrations, and determined adversaries.

An adversarial mindset closes the distance between the organization described by policy and the one that actually operates under pressure.

Leaders must ask more than whether AI is accurate, compliant, or useful.

They must ask how someone could manipulate it, misuse it, impersonate a trusted identity, corrupt its data, exploit its authority, or influence the humans responsible for oversight.

AI governance without an adversarial mindset is just policy.

Policy defines the rules.

The adversarial mindset determines whether those rules survive contact with reality.

Why AI Governance Is Now a Critical Leadership Responsibility

Why AI Governance Is Now a Critical Leadership Responsibility.
Image generated by Jetpack AI, 2026, via WordPress

Unfortunately, many organizations are on the path to repeat one of the most consequential mistakes that we have made in this industry. AI Governance is now a critical leadership responsibility.

For years, executive leaders treated cybersecurity as a technical issue. It was convenient to tuck it away under Information Technology (IT) and it became someone else’s problem. They delegated it to specialists, discussed it only when budgets or incidents demanded attention, and assumed that technical teams could contain the risk.

Then the breaches became business disruptions. Regulatory consequences reached the CFO as well as the boardroom. Trust degraded, especially from customers. Operations stopped. Executives discovered that although they could delegate security work, they could not delegate accountability for the outcome. Tucking it conveniently inside of IT was no longer an option.

Artificial Intelligence (AI) is now following an eerily similar path, only much faster.

Many organizations still treat AI governance as a collection of technical controls, acceptable-use policies, legal reviews, and model assessments. They assign it to IT, data science, security, privacy, or compliance and assume those functions can govern the technology on behalf of the enterprise.

Simply put, they cannot.

Those teams can implement controls, evaluate models, monitor systems, and advise the business. They cannot independently decide which risks the organization should accept, which decisions should be influenced by AI or automation, where humans must retain authority, or who remains accountable when an AI-enabled processes have a negative impact.

Those are leadership decisions.

AI governance is not a technical specialization that executives can delegate. It is a leadership capability that executives must develop.

Leadership Cannot Outsource Accountability

I have spent much of my career moving between deeply technical responsibilities and executive leadership. I have worked in federal law enforcement technology, application architecture, offensive security, cybersecurity, the CISO function, the CTO function, and the CEO role.

Those experiences repeatedly reinforced the same lesson: technology may create the mechanism, but leadership creates the consequence.

For me, it took a while but that had to sink in as I lived my professional journey.

An algorithm does not determine whether an organization should use AI to evaluate employees, prioritize customers, detect fraud, approve transactions, recommend medical actions, or automate security responses. Leaders make those decisions.

The system may generate a recommendation, classification, or action. It does not absorb responsibility for the result. It simply generates an output.

A model cannot accept enterprise risk.

A chatbot is likely to not be able to explain a decision to an auditor.

An autonomous agent cannot appear before the board and defend the authority it was granted.

The human signature may become less visible as the footprints of AI and automation increase, but it does not disappear. It moves upward through the organization until it reaches the leaders who authorized these systems, established their boundaries (hopefully), funded their deployments, and accepted the risks at hand (again, hopefully).

AI Means More Than Generative AI

One reason organizations misunderstand AI governance is that most current conversations concentrate so heavily on Generative AI (GenAI). And the notion of “AI” in those conversations is incorrectly used to mean “GenAI”.

Large Language Models (LLMs), copilots, image generators, and conversational interfaces have made a subset of AI (GenAI) visible to almost everyone. They have also narrowed the discussion.

AI as a field extends far beyond generated text and images. Some organizations already use AI to:

  • Detect financial fraud and account takeover.
  • Score credit and insurance risk.
  • Identify cyber threats and automate containment.
  • Rank candidates and evaluate employee performance.
  • Recognize faces, objects, behaviors, and anomalies.
  • Predict equipment failures and optimize industrial processes.
  • Recommend products, services, prices, and content.
  • Route vehicles, shipments, and supply-chain resources.
  • Support medical diagnosis and clinical decisions.
  • Operate robots, sensors, and autonomous systems.

These systems may never generate a paragraph, but they can still shape someone’s employment, financial access, safety, privacy, or treatment.

Leaders who define AI governance as a policy for using ChatGPT will govern only the most visible layer of a much larger technology landscape.

Every system that predicts, accepts, rejects, classifies, recommends, prioritizes, optimizes, or acts should fall within the governance conversation. Yet, the limited understanding of where AI actually exists within organizations does not make that proper conversation possible.

AI Governance Begins With Ownership

Every material AI system needs an accountable owner. It doesn’t need a committee or some vague reference to “the business.”

A named leader must own the business purpose, risk, performance, and consequences of the system.

Technical ownership also matters, but it is not the same as business accountability. A data science team may build a model. A cloud team may host it. Security may monitor it. Legal may review it. None of those activities answers this central question:

Who has the authority to decide that this system should operate?

Ownership must extend across the AI lifecycle:

  • Who approved the use case?
  • Who authorized the data?
  • Who selected or developed the model?
  • Who defined acceptable performance?
  • Who approved production deployment?
  • Who monitors changes in behavior?
  • Who can suspend the system?
  • Who is accountable for the outcome?

When organizations cannot answer those questions, they do not have governance. They have the illusion of governance via distributed activity and no focused accountability.

Leaders Must Establish AI Risk Appetite

Many organizations speak about AI principles. Fewer define their AI risk appetite.

Principles describe what an organization values. Risk appetite determines what it will permit.

Effective leadership demands decisions around where AI may operate autonomously, where it may only recommend, and where it should not participate at all.

That requires decisions about:

  • Which data AI systems may access.
  • Which decisions may be automated.
  • Which decisions require human approval.
  • How much uncertainty the organization will tolerate.
  • What level of explainability a use case requires.
  • How much authority and/or autonomy an AI agent may receive.
  • Which failures require immediate shutdown.
  • When efficiency cannot outweigh safety, fairness, privacy, or trust.

For example, a fraud-detection model and an autonomous industrial controller should not operate under identical tolerance levels. Neither should a marketing assistant and a system that affects employment or access to employee resources.

Optimally, governance reflects potential consequence.

That judgment cannot come exclusively from a technical scoring system. It requires leaders who understand the organization’s culture, strategy, customers, obligations, operations, and values.

Human Oversight Must Be Real

“Human in the loop” has become a very overused phrase in AI governance.

Organizations often point to human review as evidence that a system remains under control. But placing a person near an automated decision does not guarantee meaningful oversight. Nor does it even reflect reality in some cases. The sheer volume of what AI powered systems can generate make human intervention questionable.

The human may lack sufficient time and/or information to challenge the system. The interface may encourage automatic approval. Time pressure may make careful review impossible. Employees may assume that the model is more accurate than they are. Responsibility may become so distributed that nobody feels empowered to intervene.

Realistically, human oversight requires more than a final approval button.

The reviewer must have:

  • Enough context to understand the decision.
  • Enough authority to reject or override it.
  • Enough time to exercise independent judgment.
  • Enough technical literacy to recognize uncertainty.
  • Enough organizational protection to challenge the system.

Effectively, leaders must also consider automation bias. This is the natural tendency for people to trust the output of a system that appears objective, complex, or authoritative.

Ultimately, the human factor does not disappear when AI enters a workflow. It becomes more complicated.

Identity and Authority Form the AI Control Plane

Oddly, many AI governance discussions often focus on models and data while overlooking identity.

That is a serious mistake.

People build AI systems. Service accounts train them. Pipelines deploy them. Applications invoke them. Administrators change them. Agents increasingly act through them.

Every step involves an identity exercising authority.

An organization must know:

  • Who or what is acting.
  • Which identity the actor represents.
  • What authority that identity possesses.
  • What constraints exist on that authority.
  • Who granted that authority.
  • Whether the authority remains appropriate.
  • Whether the identity remains trustworthy.

This becomes especially important with autonomous agents. An agent may retrieve information, call APIs, create accounts, modify configurations, communicate with customers, or initiate actions.

An agent should not receive unrestricted access simply because an authenticated employee launched it.

It needs its own identity, constrained privileges, defined purpose, limited duration, attributable owner, and immediate revocation path.

The organization should preserve the full chain of authority:

  1. Human initiator
  2. Agent identity
  3. Delegated permission
  4. Tool invocation
  5. Impacted resource

Without that chain, the organization cannot distinguish legitimate automation from compromised autonomy.

The Adversary Gets a Vote

AI governance cannot operate only under the assumption that people and systems will behave as intended.

Adversaries couldn’t care less about the rules. They will manipulate models, compromise identities, poison data, steal credentials, exploit integrations, and misuse legitimate functionality.

They will search for the gap between what leaders think the system does and how it actually behaves under pressure.

This is where an adversarial mindset becomes essential.

Leaders should not ask only, “Does the system work?” In a headspace where there are no limits, they should also ask:

  • How could someone intentionally misuse it?
  • What happens if its data becomes untrustworthy?
  • Could a compromised identity change its behavior?
  • Can an attacker manipulate the human reviewer?
  • What authority could the system silently accumulate over time?
  • How would we detect subtle rather than catastrophic failure?
  • Can we stop it before we fully understand the incident?

Governance that assumes normal behavior is policy. And look at how effective policies are at stopping nefarious actors.

Governance that anticipates manipulation is a healthy step towards resilience.

AI Governance Must Become an Operating Rhythm

Organizations will not govern AI effectively through a policy, its annual review, or a one-time model assessment.

AI systems change. Their data changes. As do their users. Their integrations expand while authority grows. Their behavior may also shift as the environment around them changes. All of this is also happening at a rate of speed many organizations are not prepared for.

Governance must therefore become part of the organization’s operating rhythm.

Executive teams should receive recurring visibility into:

  • The inventory of blindly discovered (approved and unapproved) AI systems.
  • High-consequence use cases.
  • Detected changes in model behavior or authority.
  • Exceptions to established guardrails.
  • Third-party and supply-chain dependencies.
  • Identity exposure affecting AI environments.
  • Evidence that human oversight remains effective.

The objective is not to force leaders to review algorithms. One is to ensure that leadership understands where the organization has transferred decision-making power to machines and what could happen if that transfer fails. Another is to build a cadence of readiness preparation so that negative surprises are minimized.

Five Questions Executive Leaders Should Ask Now

Every executive team should be able to answer five questions:

  1. Where is AI already influencing decisions or actions across the organization?
  2. Who owns each material AI system and remains accountable for its outcomes?
  3. Which decisions may AI make autonomously, recommend to a human, or never influence?
  4. Can we trace every important AI action to a human or non-human identity and its delegated authority?
  5. Can we suspend the system quickly when its behavior, data, identity, or operating environment becomes untrustworthy?

If leadership cannot answer those questions, the organization is not ready to deploy and/or scale AI responsibly.

Leadership Is the Ultimate AI Control

Technical teams will remain essential to AI governance. Organizations need skilled architects, data scientists, security professionals, privacy experts, engineers, and legal counsel.

But expertise does not replace executive accountability.

AI will compress the distance between a leadership decision and its technological consequences. A policy choice can become an automated workflow. A risk tolerance can become a model threshold. A poorly governed identity can become an autonomous actor.

The organizations that succeed will not necessarily be those that adopt AI fastest.

They will be the organizations whose leaders understand where AI should have authority, establish clear boundaries around that authority, demand attributable ownership, anticipate adversarial behavior, and retain the ability to intervene.

We eventually learned that cybersecurity was not merely a technology problem.

We should not need another decade of incidents to learn the same lesson about AI.

AI Governance is now a critical leadership responsibility, it is also a leadership test.

The outcomes will reveal who studied and prepared for that test.

“The Artificial Adversary” – a New Operating Model for Cybercrime

The Artificial Adversary - a New Operating Model for Cybercrime

“Artificial adversaries don’t have egos, suffer burnout, or deal with corporate drama. Your defenses do.” – Andres Andreu

In the spring of 2026, a handful of engineers with little security background ran an experiment. They pointed an Artificial Intelligence (AI) model at thousands of software codebases and asked it to identify issues. Over the course of one night it did more than find decades-old flaws hiding in plain sight. It created working exploits for them. The model was Claude Mythos Preview. In fact, its creator judged it so capable at weaponizing vulnerabilities that it chose not to release the model at that time.

For most of our field’s history, the adversary was human. Clever and motivated but bounded by sleep, attention, money, and skill. Now, however, that adversary is being augmented, and sometimes replaced. The replacement does not tire or hesitate. Moreover, it ignores the operational rhythms our defenses quietly assume. I call it “The Artificial Adversary.” Essentially, it takes one of two forms:

  • A human operator empowered by an AI stack.
  • An autonomous AI system acting toward malicious ends.

At this stage these have stopped being thought experiments and are now turning up in incident reports.

An Inflection Point, Not a Trend Line

Three things are happening at once. Together, they mark an inflection point rather than an incremental shift:

  • AI has lowered the barrier for entry to sophisticated crime.
  • Synthetic media is collapsing our ability to trust digital signals. A familiar face or a known voice, after all, no longer proves what it once did.
  • The volume and speed of AI-enabled activity now outpaces the manual, static defenses built for a slower era.

The numbers are no longer speculative

SoSafe’s 2025 research found that roughly 87% of organizations worldwide faced an AI-powered cyberattack in the prior year. Direct attacks aside, model evaluations are just as concerning. For instance, the UK’s AI Security Institute (AISI) tested Claude Mythos Preview. It solved expert-level CTF challenges about 73% of the time. Notably, no model could complete those challenges at all before April 2025. Mythos went further still. In fact, it became the first model to solve the AISI’s 32-step simulated network takeover, from reconnaissance to full compromise. Anthropic’s red team reported even broader findings. Working alongside the AISI, it watched the model surface thousands of zero-day flaws. These included a dormant 27-year-old vulnerability in OpenBSD and a 16-year-old bug in FFmpeg. In Firefox alone, Mythos found 271 vulnerabilities and wrote exploits for 181 of them.

A signal, not the threat itself

Anthropic withheld Mythos from public release. Instead, it granted limited access to a small set of organizations that build and maintain critical software and infrastructure. The program is called Project Glasswing. Launch partners reportedly include Amazon Web Services, Apple, Cisco, CrowdStrike, Google, JPMorgan Chase, the Linux Foundation, Microsoft, NVIDIA, and Palo Alto Networks. Officially, the intent was to give defenders a head start. Yet Mythos isn’t the only game in town. For example, things such as OpenAI’s GPT-5.4-Cyber, OWASP CVE Lite CLI, and Google’s Big Sleep already show great promise and in some cases comparable capability. When competition rises the cost of entry keeps falling. Regulators noticed quickly. Within weeks, the Bank of England intensified its AI risk testing, and German banks consulted regulators and cyber experts. The lesson, therefore, is the one Bain and others drew immediately. In short, assume your adversaries are building equivalent capabilities, nation-states, criminal enterprises, and rogue actors alike. Mythos is a signal, not the threat itself.

Defining the Artificial Adversary

It helps to name the archetype precisely, because precision changes how we defend. So picture an AI-enhanced human actor. Here, the human sets the strategic objectives. The machine, in turn, executes the great majority of the tactical workload. The consequence is direct. As a result, offensive cycles compress, and defenders can no longer assume a human-speed response on the other side of the keyboard.

Human adversaries operate within cognitive, temporal, and logistical limits. An autonomous AI-based adversary does not. Needing no sleep, it carries no emotional baggage and runs continuously across global digital environments. Moreover, it can analyze vast data stores and reason probabilistically in real time. Such a system can also coordinate through decentralized, agentic architectures that resist any single point of shutdown. Its capacity for deception, mimicry, and adaptation, therefore, creates a new category of risk. Consequently, detection, attribution, and deterrence all become far harder. The asymmetry, however, is not only technological. It is also cognitive. In the end, defenders must prepare for opponents that do not tire, hesitate, or follow any rules.

The Artificial Adversary Taxonomy

A practical taxonomy has five levels.

  • AI-assisted human operator – a human attacker uses AI for discrete tasks such as phishing, translation, research, script generation, or stolen-data summarization.
  • AI-augmented threat crew – a criminal or nation-state team embeds AI into reconnaissance, exploit research, identity profiling, malware development, infrastructure staging, data exfiltration, and victim communications.
  • AI-orchestrated campaign – agentic systems coordinate personas, assign tasks, monitor responses, tune timing, and manage parallel workflows while humans supervise outcomes.
  • Semi-autonomous adversarial agent – the system conducts meaningful parts of the intrusion chain itself, including asset discovery, service testing, response analysis, and attack path modification.
  • Autonomous malicious AI system – an AI system pursues malicious objectives with limited or delayed human direction, raising harder questions around attribution, containment, predictability, and control.

This taxonomy matters because an AI-assisted phishing actor requires different defenses than an autonomous agent probing applications, manipulating identities, and adapting to telemetry in real time.

Facilitation – Lowering the Barrier

The first way AI empowers adversaries is the least glamorous and the most pervasive. Simply put, it removes friction. For a few years now, the underground has marketed “Dark LLMs.” The roster includes WormGPT, FraudGPT, KawaiiGPT, and imitators such as MalwareGPT, SpamGPT, and Xanthorox. Each promises jailbreaks, malware help, and ready-made scam playbooks. Some are functional. Many, however, are simply scams that prey on aspiring criminals. Either way, the real significance is not any single tool. Rather, it is the normalization of the idea. A capable, on-demand junior developer is now available to anyone with a few GPUs, a wallet of API keys, and some patience.

Malware that writes itself

Proof-of-concept work made the threat concrete. Researchers, for instance, demonstrated BlackMamba, a keylogger that built its malicious code at runtime by calling a Large Language Model (LLM). That approach neatly sidesteps the static signatures defenders rely on. By late 2025, the threat had moved from the lab to the wild. Google’s threat intelligence team documented two malware families: PROMPTFLUX and PROMPTSTEAL. Both query LLMs during execution. One rewrites itself, while the other generates fresh commands mid-attack. This is “Just-In-Time” (JIT) malicious code. In other words, the software does not carry its full payload. Instead, it assembles the payload on demand, from a model that does not know it is being conscripted.

When the face on the call is fake

Facilitation also reaches the human layer through synthetic media. Convincing face and voice clones, for example, can now be mass-produced. So can cross-lingual conversion and studio-quality content. Better yet for the attacker, agent teams run these operations around the clock, iterating on failures without fatigue. As a result, the multi-party deepfake video call is no longer hypothetical. Picture a finance employee walked through an “urgent” wire transfer by a “CFO” and “general counsel” who are both synthetic. Clearly, the attack surface is no longer just endpoints and identities. It now also includes the emotional tone around those identities. And does so across collaboration tools, social media, and internal communications.

Vibe Hacking – Psychological Warfare at Machine Speed

This last point deserves its own name. After all, it is where AI-enabled social engineering becomes something new. Vibe hacking is social engineering supercharged with a full AI stack. Here, the adversary does not send a single phishing email or place one deepfake call. Instead, models shape the emotional context around a target over time. The goal, therefore, is not to trick a victim once. Rather, it is to tune the “vibe” of their human state along with their digital environment, so that risky actions feel natural, familiar, and self-initiated.

Sensing, profiling, persistence

A campaign begins with sensing and profiling. To start, adversaries point AI at everything they can scrape. These sources include OSINT, LinkedIn activity, public Slack and Discord communities, conference talks, support tickets, and marketing emails. Sentiment analysis is important here and models infer mood, personality, stress levels, decision style, and trust anchors. That attackable profile, in turn, feeds a working model of the target’s context. Things like a looming quarter, a key project, the likely sources of anxiety or excitement all become real and exploitable. Generative models subsequently produce content tuned to the target’s state. The real weaponization, however, comes from scale and persistence. One artificial adversary can run dozens of long conversations at once. Each hides behind a distinct persona, the sympathetic colleague, the urgent executive, the overworked vendor. Meanwhile, it A/B tests tone, timing, and channel to learn what lowers resistance and/or skepticism. By the time the critical ask arrives, therefore, the victim feels they are accommodating a relationship, not responding to an attack.

This is the reframing that matters:

Vibe hacking isn’t better phishing. It’s your own people, profiled and played at machine scale – we hardened the edges and left the nervous system exposed.

Andres Andreu

From theory to a real victim list

None of this is a forecast. In August 2025, in fact, Anthropic’s Threat Intelligence team disclosed a case it tracked as GTG-2002. A single actor used an agentic coding tool to run a data-extortion operation. In total, the targets numbered at least 17 organizations, spanning healthcare, emergency services, government, and religious institutions. A defense contractor was among the victims, too. Remarkably, the whole campaign ran in roughly a month. To pull it off, the attacker embedded an operational playbook in a configuration file, so the AI could make tactical decisions during live intrusions. From there, the model automated reconnaissance and credential harvesting. It even generated ransom notes tailored to each victim, with demands reported between roughly $75,000 and more than $500,000. Ultimately, one person, with an AI operator alongside, did the work of a coordinated crew.

Scale – From Assistant to Operator

Facilitation lowers the barrier to entry; scale changes the magnitude. For example, the same agentic models that help an enterprise automate work can be organized into adversarial swarms. A planner agent sets the goals. Meanwhile, sub-agents run in parallel performing actions such as OSINT scraping, phishing and deepfake generation, code generation, and dropper construction. Because they share memory and data from feedback loops, the whole system improves with each iteration.

The criminal supply chain, in turn, has matured around this model. Telegram, for instance, serves as a resilient “dark social layer”, encrypted, anti-censorship, easy to churn and burn, and slow to take down. There, automated bots stream stolen credit card data and run validation checks at a pace no human team could sustain. Increasingly, the same architecture is aimed at availability, too. Agentic orchestrators break a Layer-7 denial-of-service goal into reconnaissance, traffic generation, and adaptive evasion, while coordinated worker nodes handle individual parts of the overall campaign.

The first autonomous espionage campaign

A defining incident arrived in November 2025. Anthropic reported disrupting a campaign it attributed, with high confidence, to a Chinese state-sponsored group tracked as GTG-1002. Notably, it was the first publicly documented, largely autonomous AI-orchestrated cyber-espionage campaign. It was detected in mid-September. In all, the operation targeted roughly thirty high-value organizations across technology, finance, chemical manufacturing, and government.

To pursue their objectives, the attackers manipulated an agentic coding tool into acting as a fleet of autonomous penetration-testing orchestrators and agents. First, they jailbroke its safeguards by role-playing a defensive security firm. Then they broke malicious objectives into benign-looking subtasks. From that point, the AI handled reconnaissance, vulnerability discovery, exploitation, credential harvesting, lateral movement, and exfiltration. In total, that came to an estimated 80 to 90% of tactical operations, issued at thousands of requests per second. Human operators, by contrast, stepped in only at a few strategic chokepoints. This wasn’t as clean as a Hollywood movie scene as the model’s hallucinations sometimes invented credentials or overstated findings. Those errors were among the few things keeping the operation from full autonomy.

A Real Incident, End to End – The NPD Sextortion Wave

To see these capabilities combine into one industrialized pipeline, consider the extortion spam that followed the National Public Data (NPD) breach. The underlying breach was staggering. Systems were first compromised in December 2023. By April 2024, the data had surfaced on the dark web. The company, however, acknowledged the incident only in August 2024. All told, it affected up to 170 million people and exposed as many as three billion records. The follow-on campaign was instructive less for its novelty than for its assembly. Specifically, attackers used GPT-based code generation to operationalize the stolen data end to end. The result was personalized extortion content. Each message addressed the victim by name, referenced a real home address, and embedded street-view imagery of the respective house. Then it demanded payment in Bitcoin, usually between $1,900 and $2,000, for the sake of tranquility or peace of mind.

None of the individual techniques were sophisticated. The sophistication, instead, lay in the orchestration. Consider the parts, a breach corpus, a code-generating model, a templating layer that fused public records with mapping imagery, and a delivery pipeline. Stitched together, these produced a campaign with a scale and personalization no manual operation could match. That, in essence, is the pattern security leaders should internalize. The artificial adversary rarely wins with one brilliant exploit. Instead, it wins by removing friction from every step, and running the whole chain faster than defenders can detect and respond.

Turning the Tables – Disrupting Malicious Automation

The very properties that make AI dangerous on offense also make it invaluable on defense. Better still, they open a counter-strategy that purely human teams never had. If attackers automate, then defenders can engineer the environment to exploit that automation. In practice, deception engineering and adversarial intelligence combine well.

The single goal is to convert the attacker’s automation into your early-warning system. Synthetic credentials, decoy services, and AI-generated traffic, for instance, all look irresistible to an autonomous agent. As such, they become tripwires. Because the agent probes tirelessly and indiscriminately, it hits the decoys long before a careful human would. Consequently, it can surface a campaign while it is still in an early stage.

Red teaming with autonomous agents

AI-augmented red teaming has a strong place here. In a 2024 experiment reported by WIRED, for example, a journalist let autonomous AI agents from the startup RunSybil attack a custom web app. The agents collaborated in real time. Specifically, they used SQL injection, brute-force authentication, form-field manipulation, and path traversal. Most importantly, they iterated on their failures. Without human direction, they re-planned and adjusted strategies, surfacing logic flaws that traditional scanners had missed. The agents were not malicious; their behavior, however, was. It was adversarial, coordinated, and effective. The takeaway, then, is fairly straightforward. First, adopt autonomous red-teaming agents to pressure-test your defenses against continuous, iterative, logic-driven attacks. Then pair them with high-fidelity telemetry and behavioral anomaly detection. Together, they can flag AI-like probing even when individual requests looks benign.

Governing the Machine and the People Around It

Speed without governance introduces its own risk. As defenders deploy autonomous and semi-autonomous capabilities, they take on an obligation. Those capabilities must be fast where they must be, careful where they should be, and always controllable by competent humans. Fortunately, a workable program can borrow from frameworks now maturing across the industry. For a foundation, anchor on NIST’s AI Risk Management Framework or ISO/IEC 42001. To turn principles into adversarial test cases, layer in MITRE ATLAS and the OWASP Top 10 for LLM applications. To harden the model lifecycle, draw on ISO/IEC 23894 and Google’s Secure AI Framework. Finally, add a staged maturity model to move from reactive to adaptive.

High-impact automated actions, meanwhile, need extra care. By default, mass credential revocation, large-scale connection throttling or tarpitting, and account lockouts should sit behind human-in-the-loop gates. In addition, back them with immutable audit logs, explainability proportional to impact, and fast paths to appeal and rollback.

Two cautions

Two cautions deserve emphasis.

First, treat AI models and their supply chains as critical software assets. In practice, that means validating provenance, verifying integrity, and monitoring runtime behavior. After all, data and model poisoning are now first-class threat vectors.

Second, resist the urge to fight fire with fire across legal lines. Attacker AIs, remember, routinely route through innocent third parties. As a result, heavy-handed countermeasures invite escalation and cross-border legal exposure, among them hack-back, automated counter-intrusion, and poisoning someone else’s ecosystem. Privacy by design, data minimization, auditability, and human oversight should not be compliance theater. On the contrary, they should be focused on what keeps a fast defense lawful and trusted.

What Security Leaders Must Do Now

The artificial adversary does not need to be sentient to change the game. Instead, it only needs to make capable attackers faster, more iterative, and less dependent on rare human skill. Accordingly, defenders should architect for that reality:

  • Treat AI as both adversary and ally – regularly run hybrid threat scenarios, machine-augmented attackers against machine-augmented defenders, so that you find your blind spots first.
  • Shift from signatures to behavior – static, content-based controls cannot anticipate self-modifying code or agentic chaining. Instead, invest in behavioral analytics, high-fidelity logging, and context-aware security that reads relationships, not keywords.
  • Stand up real AI governance – name a single accountable owner and convene a cross-functional oversight board. Then keep a model and agent registry, and define rules of engagement and rollback paths before you enable automation.
  • Secure the model supply chain – audit data lineage and model integrity, and assume third-party datasets, weights, and components can be poisoned upstream.
  • Deploy deception as early warning – use AI honeypots and synthetic assets to turn the adversary’s tireless automation into your early detection advantage.
  • Compress your defensive cycle – above all, adopt AI-augmented red teaming and threat hunting so that you out-learn the adversary. Then measure what matters – detection accuracy, false-positive and false-negative rates, model drift, autonomy and override rates, and time to contain.

The Pivotal Question

The pivotal question about any adversary has changed. No longer is it simply who they are or what they want. Instead, it is “what can they assemble and operationalize with AI faster than we can detect and respond?” Once, the human attacker was the central concern. Now, by contrast, security leaders face intelligent, scalable opponents that run as close to machine speed as the hardware allows. Confronting them takes more than static controls and periodic red teaming. Rather, it takes continuous learning, dynamic simulation, and AI-augmented defense. Above all, it takes one hard admission, the next major breach may not be human at all.

Awareness is the beginning; action defines resilience. The Artificial Adversary is here. The only question is whether we will be ready when it decides to strike.

Adversarial Intelligence: How AI Powers the Next Wave of Cybercrime

Adversarial Intelligence: How AI Powers the Next Wave of Cybercrime

AI Summit New York City – December 11, 2025

On December 11, 2025, I spoke at the AI Summit in New York City on a topic that is becoming unavoidable for every security leader: AI is not just improving cyber attacks, it is transforming cybercrime into an intelligence discipline. Adversarial Intelligence: How AI Powers the Next Wave of Cybercrime.

The premise of the talk was simple: adversaries are no longer running isolated campaigns with a clear beginning and end. They are building living, learning models of target organizations (e.g., your people, workflows, identity fabric, operational rhythms) and then using generative-class models and autonomous agents to probe, personalize, adapt, and persist.

The core shift: AI gives attackers decision advantage

In an AI-accelerated threat environment, the attacker’s edge often comes down to decision advantage. They see you earlier, target you more precisely, and adapt in real time when controls block them. In a pre-AI world, that level of precision required time and rare talent. Now it is becoming repeatable, automated, scalable, and accessible to people with no real skill.

Where AI shows up in the modern attack lifecycle

When people think about “AI in cybercrime”, they often jump straight to malware generation. That is not wrong, but it is incomplete. In practice, AI technologies are being applied across the attack lifecycle.

Reconnaissance becomes continuous

Autonomous agents can enumerate exposed assets, map third-party relationships, and monitor public signals that reveal how teams operate. Recon becomes less like a phase and more like a background process, always learning and always refreshing the target model.

Social engineering becomes high-context

Generative models do not just write better phishing emails. They enable sentiment analysis, tone and context matching, multi-step pretexting, and persuasion that mirrors internal language and business cadence. The outcome is fewer “obvious” lures and more synthetic conversations that simply feel real.

Identity attacks scale faster than traditional controls

Identity is the front door to modern enterprises (e.g., SaaS, SSO, MFA workflows, help desk interactions, API keys). AI-powered adversaries can probe identity systems at scale, adapt-ably test variants, and blend into normal traffic patterns, especially when enforcement is inconsistent.

“Proof” gets cheaper: impersonation goes operational

Deepfakes and impersonation have moved from novelty to operational enablement. They can be used for vibe hacking (e.g., pressure targets, accelerate trust, push high-risk decisions), especially in finance, vendor-payment, and administrative workflows.

The defensive answer is not “more AI“. It is better strategy.

A common trap is thinking, “attackers are using AI, so we need AI too”. Yes some AI is necessary, but alone it is not enough. Winning here requires adversary-informed security: security designed to shape attacker behavior, increase attacker cost, and force outcomes.

Three tactics that disrupt malicious automation

Deception Engineering: make the attacker waste time … on purpose

Deception is no longer just honeypots and honeytokens. Done well, it is environment design: believable paths that look like privilege or data access, instrumented to capture telemetry and shaped to slow, misdirect, and segment adversary activity. The goal is not only detection. It is decision disruption, raising uncertainty and forcing changes within the adversary’s ecosystem.

Adversarial Counterintelligence: treat your enterprise as contested information space

Assume adversaries are collecting, correlating, and modeling your ecosystem, then design against that reality. Practical counterintelligence includes reducing open-source signal leakage, hardening executive and finance workflows against impersonation, and introducing verification into high-risk decisions without paralyzing the business.

AI honeypots and canary systems: fight automation with instrumented ambiguity

AI-enabled adversaries love clean feedback loops. So do not give them any. Modern deception systems can present plausible but fake assets (APIs, credentials, source code repositories, data stores), generate dynamic content, and create unique fingerprints per interaction so automation becomes a liability.

What this means for CISOs: measure money, not security activity

If you are briefing a board, do not frame this as anything like “AI is scary”. Frame it as: AI changes loss-event frequency, loss magnitude, and time-to-detection/time-to-containment. These can directly impact revenue, downtime, regulatory exposure, and brand trust. If attackers can industrialize reconnaissance and/or persuasion, then defenders must industrialize identity visibility, verification controls, detection-to-decision workflows, and deception at scale.

Key takeaways

  • Assume continuous and automated recon.
  • Harden verification workflows against synthetic content; train executive and administrative teams regularly.
  • Deploy deception at scale; raise attacker cost to reduce downtime.
  • Operationalize counterintelligence; aim to avoid blind spots to reduce exposure.
  • Quantify decision advantage to accelerate funding decisions and defend revenue/margins.

Closing thought

AI is accelerating the adversary, no question. It has also lowered the entry barrier to cybercrime. But it is also giving defenders a chance to re-architect advantage: to move from passive defense to active disruption, from generic controls to adversary-shaped environments, and from security activity to measurable business outcomes.

The real message behind adversarial intelligence is this: the winners will not be the organizations that merely “adopt AI”. They will be the organizations that use it to deny attackers decision advantage, and can in turn prove it with metrics the business understands and values.