Capability grants
Your agent can be fooled. Its grants cannot
Detection is a probability. A grant is a fact about the call.
The problem
Every text-based guard has the same ceiling: it is trying to decide whether a string is an attack, and attackers get a fresh attempt every request. A capability check asks a different question — was this agent ever allowed to do this? — and the answer does not change because the prose was persuasive.
Two ladders, and a call has a rung on each
What the tool can do, and how much the data reaching it can be trusted. A refusal is usually the pair, not either one alone.
Tool impact
readReturns something. Changes nothing.writeChanges state that can be changed back.high_impactWide blast radius, or spawns something with its own tools.irreversibleMoney moved, data deleted, mail sent. No undo.
Argument taint
noneConstant, from your own code.userThe operator typed it.retrievedIt came out of a document.tool_resultA tool returned it. A third party wrote it.subagentAnother agent asserted it.memoryIt was written to memory earlier, by something.
- 01
Declare what each tool actually does
A tool is
read,write,high_impactorirreversible. That is the floor for what a call can be reasoned about as — and for a shell tool it really is only the floor, becauselsandrm -rfare the same tool.agentfox tools declare - 02
Grant the capability, with its limits
A grant carries constraints — a value ceiling, an environment, a maximum taint for the data that may reach it. Anything not granted is refused; that is the default, not a rule somebody has to remember to write.
- 03
Track where each argument came from
Arguments are tainted by origin: something the operator typed, something a document said, something a tool returned. A transfer whose amount came out of a retrieved page is a different call from one the user asked for, and the record says which it was.
- 04
Untrusted content may fill a value, never choose an action
This is the line the whole model rests on. A retrieved document can supply an account number that then gets checked; it cannot decide that a transfer is the next step. Control flow belongs to the operator’s intent, and data belongs to the data.
- 05
Read the chain, not only the step
Two harmless calls can compose into a privilege escalation, and a destructive action is often reached rather than requested. Cascades, blast radius and loops are evaluated across the run, so an outcome nobody asked for in one step is still caught.
What grants do not do
A grant is only as good as the declaration behind it, and the impact we infer for an undeclared tool is a guess we label as one. We also cannot bound what a tool does on the other side of its own API: if a tool you declared as a read deletes something, containment believed you.
Try to break it before you trust it
No account, no install, and the same enforcement code as the product.
pip install agentfox agentfox init && agentfox demo
Offline: no API key, no downloaded weights, no network egress.