Start
Concepts
The vocabulary every other page uses, defined precisely enough to predict what AgentFox will decide.
When to use this
Read it once after the Quickstart, and come back when a verdict surprises you. Each term says where it is used, so you can jump to the guide or command that works with it.
How a tool call is decided
Most of the concepts below meet in one place: the check that runs before a tool call executes. In order:
- Which agent is calling, and is it stopped (quarantine or kill switch)?
- Does it hold a grant for this tool, and do the argument values fit the grant's limits? No grant: refused (default deny).
- What is the tool's impact, and what is the provenance of the arguments? Untrusted values in an irreversible call: escalated to a person.
- Was a value produced by one tool passed into a higher-impact one (composition)? Was a task intent declared? Is the run looping or over budget?
- What did the detectors find in the arguments? Detection rules apply on top, in whatever mode their policy is in.
The strongest effect wins, it becomes the verdict, and the decision is written to the audit chain. Steps 2 and 3 read no text at all, which is why they hold when a detector misses.
Agent
A program that calls a model and can take actions. AgentFox identifies each one by a slug such as support-triage or payments-ops. Every grant, policy binding, finding and decision is attached to an agent.
| Registered | Created by agentfox.auto("support-triage"), by the hooks installer, or in the web app, with an owner and a risk tier. |
|---|---|
| Shadow | Registered by its own traffic: the first call through the gateway from an unknown agent creates it, unowned. It shows up in agentfox agents list and as a shadow_agent finding. |
| Stopped | agentfox agents quarantine or agentfox agents kill; every call is refused until agentfox agents resume. |
Used in: Agents in the web app, Approvals and the kill switch.
Tool
Anything an agent can call that is not the model: a Python function, an HTTP API, an MCP server method, a shell. A tool has a key (tickets.close, crm_lookup, billing.export), an impact, an output trust, and optionally the downstream effects it triggers. Tool declarations are organisation-wide.
Used in: agentfox declare tool, agentfox declare list tools, Contain tool calls.
Impact
What a tool can do to the world. It is the axis every containment rule reasons over. Four tiers:
| Impact | Meaning | With untrusted arguments |
|---|---|---|
read | Reads data, changes nothing. | Allowed, if granted. |
write | Changes data, and the change can be undone. | Escalated when an argument came from a tool result or worse. |
high_impact | Reversible, but costly if wrong. | Escalated when anything is untrusted. |
irreversible | Cannot be taken back: sending, paying, deleting. | Escalated when anything is untrusted; also escalated when no task intent was declared. |
Inferred or declared. A tool agentfox.auto() sees for the first time is registered with an impact guessed from its name, and marked inferred until you confirm it:
agentfox declare list tools
agentfox declare tool email_send --impact irreversible tool impact output triggers
crm_lookup read (inferred — confirm with `agentfox declare tool`) untrusted —
email_send irreversible (inferred — confirm with `agentfox untrusted —
declare tool`)
…
email_send declared — impact irreversible, output untrusted
arguments carrying untrusted provenance now require approval or are refused, whether or not a
detector firesCapability grant
Permission for one agent to call one tool (or a glob such as tickets.*), with conditions. Default is deny: once an agent is under containment, a call no grant covers is refused with capability.denied, whether or not anything looked suspicious.
| Part | Flag | What it does |
|---|---|---|
| Argument limits | --limit rows:lt=1000, --limit currency:in=USD,EUR | A call outside a limit is refused with capability.constraint_violated. |
| Provenance ceiling | --max-taint user (default) | The worst provenance an argument may carry and still go through without an approval. |
| Approval | --requires-approval | Every matching call goes to a person first. |
| Expiry | --expires-in-days 30 | The grant withdraws itself. |
| Accountability | --granted-by | Recorded in the audit chain with the grant. |
The same billing.export call, against a grant of --limit rows:lt=1000 --max-taint user, asked three ways through POST /v1/guard/tool_call:
| Arguments | Provenance | Verdict | Rules fired |
|---|---|---|---|
rows: 250 | all user | allow | none |
rows: 250 | to from tool_result | escalate, with an approval_id | taint.irreversible_tool, capability.approval_required |
rows: 5000 | all user | block | capability.constraint_violated |
The middle row is the point: same tool, same values, same agent; only where one value came from differs. Grants are made by hand with agentfox permit grant (it asks before it writes; --yes in scripts) or drafted from traffic as proposals.
Used in: Contain tool calls, the gateway.
Provenance (taint)
Where an argument's value came from. AgentFox tracks it per value and compares it against the grant's ceiling and the tool's impact. Ordered from most to least trusted:
| Source | Means |
|---|---|
none | No external origin. |
user | A person typed it: the user's message. |
retrieved | It came out of a retrieved document. |
tool_result | It came out of another tool's output, such as a fetched web page. |
subagent | Another agent produced it. |
memory | It was read back from the agent's long-term memory. |
Everything after user is untrusted. With agentfox.auto(), provenance is read from the conversation: a value copied out of a role="tool" message counts as tool output. Over HTTP, the caller says it in the request's provenance map.
Session or argument scope
One setting, taint_scope, decides what a call's provenance is:
taint_scope | A call's provenance is | Trade-off |
|---|---|---|
session (default) | The worst untrusted content anywhere in the run so far, or in the call's own arguments. Once the agent has read a web page, every later irreversible call carries it. | Contains attacks whose payload never lands in an argument. Escalates much legitimate work. |
argument | Only what this call's own arguments were copied from. | Lets more legitimate work through. Misses an attack whose values are not found in the arguments: identifiers shorter than six characters are never matched, nor is attacker text inside a longer argument. |
Measured on AgentDojo (97 user tasks, 949 attack pairs, provenance inferred from the real tool outputs, every detector off), from benchmarks/agentdojo/README.md:
| Provenance | Benign tasks run without escalation | Attack pairs contained |
|---|---|---|
| Session-level (shipped default) | 24/97 (24.7% [17.2, 34.2]) | 588/588 |
| Argument-level | 37/97 (38.1% [29.1, 48.1]) | 527/588 |
| Session-level, read-only tools exempt | 43/97 | 588/588 |
| Argument-level, read-only tools exempt | 62/97 (63.9% [54.0, 72.8]) | 527/588 |
The Quickstart shows the same trade on four tickets. Under session, two legitimate replies sent after the agent read a status page were escalated, and the injected email was refused. With AGENTFOX_TAINT_SCOPE=argument, the same run gave:
Ticket T-311 from customer c -> T-311 handled.
Ticket T-312 from customer c -> T-312 handled.
Ticket T-313 from customer c -> T-313 handled.
Ticket T-314 from customer c -> refused: agentfox: tool call email_send was refused by composition.escalation: argument 'to' carries a value produced by tool 'web_fetch' (read), now passed into 'email_send' (irreversible) — …Set it in agentfox.toml or the environment (Configuration). A value other than session or argument is an error at startup, not a silent fallback.
Output trust
Whether values copied out of a tool's output count as untrusted. Every tool is untrusted unless declared otherwise. Declare trusted only for a system of record you control: a CRM read, your own mail service's confirmation. Then an email address copied out of the CRM into email_send is not treated like one scraped from a web page, and the tool no longer raises the run's provenance.
agentfox declare tool crm_lookup --impact read --output-trust trustedcrm_lookup declared — impact read, output trusted
values an agent copies out of this tool's output no longer count as untrusted input, and no longer
raise the run's provenanceThis is also the fix for composition.escalation, which refuses a value produced by one tool and passed into a higher-impact one, unless the producing tool is trusted.
Policy, pack, rule and mode
| Rule | A condition and an effect: taint.irreversible_tool says an irreversible tool with arguments above user escalates. Each rule maps to controls in the catalog. |
|---|---|
| Policy | A named, versioned set of rules, bound to agents, teams or environments. |
| Pack | A policy shipped with the product. baseline (detector rules), eu-ai-act-high-risk, tool-containment (grants, provenance, composition, intent, loops, budgets), and coding-agent, which is only bound to agents that coding-agent hooks govern. |
| Mode | observe records what the policy would have done and changes nothing; enforce does it. The mode belongs to the policy, not to the integration. |
agentfox policy listpolicy version mode rules
baseline v1 observe 13
eu-ai-act-high-risk v1 observe 7
tool-containment v1 enforce 25tool-containment enforces from agentfox init because its rules are structural facts (no grant, an untrusted value in an irreversible call), not classifier scores. The detector packs start in observe because a false block is how guardrails get switched off. agentfox policy enforce baseline is the one step that starts blocking model traffic, and agentfox policy observe baseline reverses it.
Intent is the agent's task in a sentence, declared with auto(intent=…) or the X-AgentFox-Intent header. An irreversible call with no declared intent is escalated by intent.undeclared_irreversible.
The integration has its own mode too, for agentfox.auto(): "policy" (default) raises exactly when an enforcing policy refuses; "observe" never raises; "enforce" raises whenever the effective verdict blocks, even for a policy in observe, for tests and CI.
Used in: Policy language, Policies in the web app, Tune detectors.
Verdict
What a check returns. Every rule that fires has an effect; the strongest one wins. From weakest to strongest:
| Effect | What happens |
|---|---|
allow | The call or text goes through. |
tokenize, mask, redact | It goes through with the matched values replaced. |
abstain | The answer is withheld and a declared refusal is said instead. |
escalate | Held for a person; an approval is created. |
block | Refused. |
Over HTTP the response is a 200 carrying the verdict, so your code branches on verdict rather than on an error status. In Python, an enforced block or escalate on a tool call raises agentfox.Blocked. In observe mode two verdicts are recorded: the one applied, and the one the policy would have applied (x-agentfox-verdict and x-agentfox-effective-verdict on a proxied call). The gap between them is what you watch before enforcing.
Decision and trace
A decision is one evaluation: the surface checked, the verdict, every rule that fired with its reason, the detector results, and the provenance it was judged on. It has an id (dec_…) and lands in the audit chain.
A trace (trc_…) is one request end to end: the input guard, the model call, the output guard, each tool call, with timings. The decisions hang off it. agentfox.auto() creates one per model call; the gateway returns its id in x-agentfox-trace.
Checks run on nine surfaces: input, output, tool_args, tool_result, retrieved, memory_write, agent_message, completion (the agent claiming it is finished) and reasoning.
Used in: Traces in the web app, Traces and integrations.
Finding
Something a person should look at. Findings are deduplicated by fingerprint, so a problem that recurs is one finding with a count (4x), not four.
| Type | Raised when |
|---|---|
containment | A permission, provenance, composition or blast-radius rule stopped (or would have stopped) a tool call. Titled by what stopped it, for example "tried to pass the output of web_fetch into email_send, a higher-impact action". |
guardrail_detection | A detector matched (injection, personal data, a secret) and a rule acted on it. |
shadow_agent, unowned or stale agents | Inventory problems from traffic and the registry. |
Contained or would have been. A finding ending (contained) or (held for approval) was stopped by an enforcing rule. One ending (would have been contained) was decided by a rule in observe mode and let through. agentfox report keeps the two in separate sections.
Used in: agentfox findings, Findings in the web app, Detectors and findings.
Approval
The record an escalate verdict creates (apr_…): the call, its arguments and provenance, and why it needs a person. Someone approves or denies it in the web app or through POST /api/approvals/{id}/approve and /deny. An approved call counts as clean evidence the next time permissions are drafted from traffic.
Used in: Approvals and the kill switch, the approval queue.
Proposal (learned permissions)
A change to governance configuration that the product drafts and a person decides. The kind you meet first is a grant or tool declaration drafted from traffic by agentfox policy proposals from-traffic. Limits and the provenance ceiling are read only from clean calls: ones refused only because nothing was configured, or approved by a person. Calls held for provenance shape neither, so an attacker's recipient never becomes a limit.
| Status | Means | Moves to |
|---|---|---|
proposed | Filed. | proven, rejected, superseded |
proven | Replayed against the recorded calls it was drawn from. | approved, rejected, superseded |
approved | A person said yes (proposals approve). | canary, applied, rejected, superseded |
canary | Live for a cohort only. | applied, rolled_back |
applied | In force (proposals apply). | verified, rolled_back |
verified, rejected, rolled_back, superseded | Final. | — |
A proposal that loosens a control is never applied by automation. Grants are scoped to one agent and need one approver; tool declarations apply to the whole organisation and need two different people.
Used in: Quickstart, Contain tool calls, Tune detectors.
Control points
The same policies and the same decision record, bound in six places. Each sees a different surface, so connecting one does not cover the others.
| Control point | Connect with | What it sees | What it does not |
|---|---|---|---|
| SDK (Python) | agentfox.auto() | Every OpenAI, Anthropic, LiteLLM and LangChain model call in the process (sync, async, streamed): messages, response text, and each tool call in the response before your code runs it. | The OpenAI Responses API; tools your code calls without the model asking. |
| Gateway | agentfox serve | Model calls proxied through /v1, and whatever you ask about on /v1/guard/*: input, output, tool calls, memory writes, agent messages. | Anything your code does not send to it. |
| LangGraph | AgentFoxGuard node wrappers | Retrieval, model and tool nodes; state survives checkpoints. | Nodes you did not wrap. |
| MCP | agentfox scan mcp and the governor | Server configs before anything starts; at call time, each tool's digest against the one reviewed, and results as untrusted input. | Servers no config declares. |
| Hook (coding agents) | agentfox admin hooks install | Claude Code's UserPromptSubmit, PreToolUse (can stop the call) and PostToolUse (cannot withdraw it; marks the result untrusted). | Anything off this machine, including a session running in the vendor's cloud. |
| CLI | agentfox scan, agentfox test, agentfox policy simulate | Source code, configs, and recorded traffic replayed against a candidate. | Live calls; it decides nothing at runtime. |
Guides: Python, gateway, LangGraph, MCP, coding agents, CLI and CI.
Audit chain and evidence
Every decision, grant, approval and operator action is appended to a hash chain: each entry carries the digest of the one before, and signed checkpoints are written over the head. Editing a row breaks every digest after it.
agentfox report verifychain: 215 entries, head seq 215, 2 checkpoints
CHAIN INTACT — 215 entries verified (seq 1..215)An evidence package (agentfox report evidence) is a zip of the agents, traces, decisions, approvals, findings and chain rows for a period, with the one-page summary on top and a standalone verify_chain.py that re-derives every digest using only the Python standard library. Framework mappings inside it are labelled DRAFT — UNVERIFIED / NOT LEGAL ADVICE until a qualified reviewer signs them off.
Used in: Prove it to an auditor, Compliance in the web app. The signing key is AGENTFOX_AUDIT_SIGNING_KEY; change it before any real deployment (Self-hosting).