
Many conversations about AI agent risk over the past year start from the same unspoken assumption: something bad happened because someone or something manipulated the agent. A hidden instruction in a document, a poisoned prompt, an adversary steering the model toward an action it shouldn't take. That's a real category of risk, and it deserves the attention it's getting. But, it's also not the whole picture, and the incidents that fall outside it are, in my experience, the ones security teams are least prepared to catch. Not because the tooling is immature, but because there's structurally nothing for that tooling to find.
Consider a scenario I keep coming back to when I talk to security teams about where their agent monitoring actually has gaps. A development team deploys a coding agent to fix a mismatch in a staging environment. Mid-task, the agent hits an inconsistency nobody briefed it on: a database schema that doesn't match what its original instructions described. No one manipulated the agent. No malicious content entered its context window. Trying to resolve the discrepancy on its own, the agent reaches for a deployment token its role legitimately carries, because that same deployment role also happens to handle production releases. It uses the token to connect to a production database it was never instructed to touch, and while trying to reconcile the schema, it drops the production table.
The entire sequence takes under ten seconds. Every individual API call the agent made was permitted by a role a human had already approved. Nothing in the sequence would trigger a prompt-injection filter, because there was no injection to catch. From the perspective of every identity and access control system watching that session, an authorized agent used an authorized credential to take an authorized action. The audit log reads as routine, right up until the production table is gone.
Why the Industry Keeps Reaching for the Wrong Fix
The instinct, when I describe this scenario to security leaders, is almost always the same: tighten the role, and the problem goes away. It's an understandable instinct, and it's wrong in a way that matters for how we build agentic security programs going forward.
A human engineer's job description is stable for months at a time. Scoping a role to that job description is a durable control, precisely because the job itself doesn't change from one task to the next. An agent's effective task set doesn't work that way. It's decided by the model at runtime, and it can shift with every single invocation based on the specific instructions, the specific context, and the specific obstacles the agent encounters along the way. There is no equally stable job to scope a role to, because the agent doesn't have one job in the way a human does. It has whatever task it was dispatched to perform in that particular session, and the actual boundaries of that task are often only fully knowable after the fact, once you can see what the agent decided to do with the room it had.
This is precisely why traditional privileged access management doesn't already close this gap, and I don't think that's a knock on PAM as a discipline. PAM programs were built around an assumption that holds for human principals: authorized scope maps to a role that changes slowly, if at all. Agents violate that assumption structurally, not occasionally. A role tight enough to prevent every possible out-of-scope action would also be too tight to let the agent complete the legitimate variations of its actual job. A role loose enough to handle those legitimate variations will, sooner or later, also permit an action the agent was never supposed to take. There isn't a scoping sweet spot that solves this. The problem lives one level up from where scoping operates.
The Question No Existing Control Knows How to Ask
What was actually missing in the production database scenario wasn't a tighter role. It was an evaluation, at the moment of the call, of whether that specific action belonged to the task the agent was actually dispatched to perform. I want to be precise about this, because it's easy to gesture at "better governance" without naming the actual gap: no identity and access management policy carries the concept of "the task." IAM answers whether a principal is authorized to access a resource. It has no mechanism for answering whether this particular use of that access, in this specific moment, given everything the agent has already done in this session, is consistent with what it was sent to do.
That's a fundamentally different kind of control than a permission boundary, and I think this distinction is under appreciated in how the industry currently talks about agent governance. It requires evaluating the action itself against the agent's declared purpose, not inferring intent from the model's language and not waiting for a policy violation that never occurs because no policy was technically violated. Posture and detection tooling built for this failure mode need to flag out-of-scope actions even when the agent's own reasoning trace looks entirely coherent and well-intentioned, because in this class of incident, the reasoning usually is coherent. The agent isn't confused. It's reasoning correctly toward the wrong boundary.
This also has an implication for how we think about behavioral baselines, which I think the industry has gotten slightly backward. There's a common assumption that an agent needs a track record before you can meaningfully evaluate whether its behavior is normal. But an agent that has never run before still shouldn't be permitted to take a destructive action outside its declared scope, and that protection can't wait for a baseline to accumulate. If the control only kicks in once you've observed enough sessions to know what "normal" looks like, you've already ceded the first N incidents to exactly the pattern I'm describing here.
What This Means Beyond Coding Agents
I've described this specific scenario in terms of a coding agent because it's concrete and it's the version most security teams have already heard about in some form. But the underlying mechanism isn't specific to code. A customer-facing agent handling account inquiries, given broad access across a CRM and billing system to complete legitimate service tasks efficiently, can reach into an adjacent module during an unusual request and take an action nobody scoped it to take, for exactly the same structural reason: the credential was correctly issued for the agent's normal task and badly over-scoped for the specific moment it decided, on its own reasoning, to act outside that task.
An orchestration agent coordinating a multi-step workflow runs into this even more directly. It can hand a sub-task to a specialized agent that has broader system access than that specific step requires, and that broader access becomes available the instant the sub-agent decides it needs it, whether or not the human who set up the workflow ever anticipated that particular path. The scenario generalizes precisely because the root cause generalizes: any agent whose credentials were scoped to a role, rather than evaluated against a task in real time, carries this exposure by default.
I don't think this is a solved problem yet, industry-wide, and I'd rather say that plainly than pretend otherwise. What I do think is that naming the failure mode correctly, as something other than a manipulation problem, is the necessary first step to building the right control for it. Testing for this in a vendor evaluation is more useful than asking about it directly: in a sandbox, give an agent a legitimate task and a credential correctly scoped to an adjacent system, then introduce an environmental inconsistency it was never briefed on. It passes if the platform blocks the action against the adjacent system and raises an alert naming the agent identity, the credential used, the action attempted, and the deviation from expected scope, with no injected content anywhere in the trace. It fails if the sequence is only reconstructable from post-hoc logs, or if the alert fires on prompt content rather than on the action. Run that test against every vendor in your evaluation, including us.
Why "No Evidence of Manipulation" Keeps Getting Treated as "No Incident"
There's a pattern I've noticed in how incident review processes handle this class of event, and I think it's worth naming directly because it's easy to fall into without realizing it. When a post-incident review finds a clean credential chain, no injected content, and no external actor, the instinct is to file the finding as a near-miss or an edge case rather than as a structural gap in the security architecture. That instinct makes sense if you're evaluating the incident against the mental model of "was this an attack," and the honest answer is no, it wasn't. But that's the wrong question to be asking if the goal is figuring out whether the security architecture actually has coverage for this failure mode, because the absence of an attacker doesn't mean the absence of a control failure.
I think this distinction matters more as agent deployments scale, because the ratio of these non-adversarial failures to manipulation-driven ones is likely to shift over time, not stay fixed. As agents take on more genuinely open-ended tasks, the number of moments where an agent encounters something its instructions didn't anticipate is going to grow faster than the number of moments where an attacker successfully manipulates it, simply because open-ended tasks create more surface area for legitimate improvisation than for injection. A security program that only tunes its detection for manipulation is going to find its coverage gap widening exactly as agent autonomy increases, which is the opposite of where you want the gap to be trending.
There's also a governance implication here that I think gets underappreciated. If an organization's incident response playbook treats "no evidence of manipulation" as grounds to close an investigation, that playbook is going to systematically under-report this category of incident, not because anyone is hiding anything, but because the playbook itself doesn't have a category for it. Fixing that starts with explicitly separating two questions that current review processes tend to collapse into one: was this agent compromised, and did this agent's action match the task it was given. The first question has a clean yes-or-no answer most of the time. The second one requires an entirely different kind of evaluation, one that has to happen at the moment of the action rather than during a retrospective review, because by the time the review happens, the production table is already gone.
Download the The Enterprise Buyer's Guide to Agentic AI Security to learn how to evaluate, compare, and select security solutions purpose-built for the age of AI agents.
All ArticlesRelated blog posts

MCP Is Growing Up
Our team spends much of the week talking to security leaders and practitioners who are trying to figure out where...

The Agent Will See You Now: Why Healthcare's AI Agent Boom Needs Visibility and Control
Healthcare, as an industry vertical, is moving faster on agentic AI than it has in past technology evolutions....

"AI Regulation" Isn't One Debate. It's Several, Wearing the Same Coat.
Ask ten people what "AI regulation" means, and you'll get ten different answers, and most of them will assume the...
Secure Your Agents
We’d love to chat with you about how your team can secure and govern AI Agents everywhere.
Get a Demo