AI agent security
Agentic AI security review and adversarial testing. The failure is rarely in the code; it is in what the agent is permitted to do and what it is willing to believe.
-
Agent red teaming
Adversarial testing of agent behaviour across multi-turn attack chains.
-
Prompt injection research
Direct and indirect injection reaching privileged tools and data.
-
Tool and permission abuse
Privilege escalation through the capabilities an agent already holds.
-
Model supply chain review
Model, MCP server and dependency provenance and trust.
What is tested
Prompt injection reaching a privileged action, direct and indirect. Jailbreak paths that survive multiple turns. Tool manipulation and agent privilege escalation, where an agent is persuaded to reach for a capability it holds but should not use here. Context poisoning through retrieved documents, and RAG exfiltration where the retrieval layer becomes the leak.
Then the plumbing underneath: MCP server trust and poisoning, credential and tool scope, unsafe autonomy in build and deploy pipelines, and the model supply chain the whole system inherits. Coverage maps to the OWASP Top 10 for LLM Applications, extended for agentic behaviour that the list predates.
Why adversarial testing is continuous
A point-in-time assessment describes a system that no longer exists. Agent behaviour changes when the prompt changes, when a tool is added, when the model version moves, and when the data it retrieves changes underneath it — none of which are code deployments, and none of which trigger a review.
Testing runs continuously against a defined scope, so a finding arrives when the behaviour changes rather than at the next audit window.
Tested by a system that runs agents
This discipline is performed by agents that are themselves operating in production across other professions. The failure modes are not read from a paper; they are the ones encountered while running a large number of agents against each other continuously.
Output is adjudicated by a separate agent before it leaves the system, and every finding carries the path that produced it.
Engagement
- Pre-deployment review of an agent system
- Continuous adversarial testing against a defined scope
- Standing agent security capability
Questions
- What is agentic AI security?
- Agentic AI security is the review and adversarial testing of systems where a model takes actions through tools rather than only producing text. The risk moves from what the model says to what it is permitted to do: which tools it can reach, which credentials those tools carry, and whether untrusted input can steer either.
- How is AI red teaming different from a penetration test?
- A penetration test targets deterministic systems where the same input produces the same result. Agent behaviour is probabilistic and changes with prompt, tool set, model version and retrieved context. Red teaming an agent means attacking behaviour across multi-turn sessions rather than enumerating endpoints, and it has to run continuously because the target changes without a deployment.
- What is prompt injection testing?
- Prompt injection testing determines whether text an agent reads — from a user, a document, a web page or a tool response — can override its instructions and cause a privileged action. Indirect injection, where the payload arrives through retrieved content rather than the user, is the case most systems fail.
- Does this cover MCP servers and RAG pipelines?
- Yes. MCP server trust and poisoning, and exfiltration through the retrieval layer, are both in scope. They are among the most common paths by which an otherwise sound agent leaks data or acts outside its remit.
- When should an agent system be reviewed?
- Before it reaches production, and continuously afterwards. The pre-deployment review catches permission and trust design faults while they are cheap to change; continuous testing catches the behaviour drift that follows every prompt, tool and model change after launch.
- Security Continuous offensive security and defensive operations run by autonomous agents: vulnerability research, exploit development, adversary emulation and detection engineering.
- Compliance Compliance automation by autonomous agents: continuous control testing, audit evidence collection and third-party risk review for SOC 2, ISO 27001 and DORA.
- Revenue operations Agents performing clinical coding, denial appeals and prior authorisation, at a volume and consistency staffing cannot reach.
- Bid and proposal Agents performing requirement extraction, compliance matrices and response drafting for government and enterprise tenders.
Bring a market. We run the discipline.
For operators and organisations deploying agents at scale. Not a hire. Not an enquiry.
partner@boutrig.com