Testing AI Agents for Excessive Agency: Where Can AI Act Without Approval?

Contributors

Shantanoo Govilkar
Shantanoo Govilkar
SVP Strategic Solutions Risk & Cybersecurity Solutions

AI security testing in 2026 has moved beyond asking whether a model can generate an unsafe response. For agentic AI, the more important question is what the system can do after receiving an instruction.

An AI agent may have access to CRM records, email, databases, cloud services, ticketing platforms, code repositories, or financial workflows. If that agent can invoke tools autonomously, a manipulated instruction can become an actual security event.

OWASP's 2026 guidance highlights risks associated with excessive permissions, functionality, and autonomy in agentic systems.

A penetration test should therefore establish the agent's actual authority rather than simply reviewing its configuration.

Testers can map the tools available to the agent and identify the credentials or service identities those tools use. They can then test whether the agent can perform actions outside its intended business function.

For example, an agent designed to retrieve customer information may legitimately require CRM read access. That does not mean it should be able to modify customer records, create users, change permissions, or export an entire database.

Testing should also introduce adversarial inputs designed to cross these boundaries. These can include direct prompt injection, malicious instructions embedded in retrieved content, poisoned documents, or attempts to manipulate the agent into chaining legitimate tools in an unauthorized sequence.

The important finding is not simply that prompt injection worked. It is what the successful attack allowed the attacker to accomplish.

An AI system that generates an inappropriate answer has a model-safety problem. An AI agent that can be manipulated into accessing sensitive systems has an exploitable security boundary.

As organizations give AI systems greater operational authority, penetration testing needs to validate both sides of that boundary: whether the agent can be manipulated and whether that manipulation can be converted into unauthorized action.

Back
to Top