Start
AgentFox documentation
AgentFox decides what an AI agent may do from what you declared, what you granted, and where each argument came from, and records every decision so you can prove it later.
Most agent-security tools are detectors: they read the text going into and out of a model and try to recognise an attack. A detector that misses lets the action through. AgentFox starts from the other end. Each tool carries a declared impact (read, write, high impact, irreversible), each agent holds explicit grants with argument limits, and every argument carries the provenance of its value. An email whose recipient was copied out of a web page is refused because of where the address came from, whether or not anything recognised the page as hostile. Default deny means a tool nobody granted is refused on the first call.
Detectors are still there, as a second layer: injection, personal data, secrets and unsafe content on every surface a request touches. They ship in observe mode, they raise the cost of an attack, and the product is built so that their failing is survivable. Containment was measured with every detector switched off; the numbers, and what they cost in escalated legitimate work, are on Benchmarks. What it does not do yet is on Limits.
The path: see, watch, contain, prove
Four steps, in this order. Each is one command or one line, and each has a guide. The Quickstart walks all four against a sample agent in about ten minutes.
| Step | What you get | Guide |
|---|---|---|
Seeagentfox scan | A static read of the repository: every model call site and whether it is governed, every tool and MCP server, and each place where one agent can read private data, read untrusted content and send data out (the lethal trifecta). Nothing leaves the machine. | Audit a repository |
Watchagentfox.auto(mode="observe", intent="…") | One line at your entry point. Every model call and every tool call the model asks for is traced, checked and recorded. In observe mode nothing is refused; you get the list of what would have been. | One line in Python |
Containagentfox policy proposals from-traffic | Tool declarations and grants drafted from the calls the agent actually made, with limits read only from clean calls. A person approves each one; nothing widens on its own. Or write them by hand with agentfox permit grant and agentfox declare tool. | Contain tool calls |
Proveagentfox report | One page for whoever signs off: what ran, what was contained and why, what observe mode would have stopped. The same page opens every evidence package, alongside a hash chain an auditor re-derives with a standalone script. | Prove it to an auditor |
No agent of your own yet? The offline demo needs nothing but the install:
pip install agentfox
agentfox init && agentfox demoinit creates a local database and loads the controls and policy packs. demo runs a thirteen-step walkthrough against three seeded agents in about five seconds. Use a scratch state directory for it; see Quickstart.
What you can do
Five capability areas. Containment is the one that holds when the others miss, so it is the one to configure first.
See
What agents, tools and servers do you already have?
| You want to | Run |
|---|---|
| Inventory a repository: model calls, tools, MCP servers, the lethal trifecta | agentfox scan |
| First look at this machine, including local AI-tool sessions | agentfox scan --sessions |
| MCP servers: reach, version pinning, auth, tool descriptions | agentfox scan mcp |
| Agent skills: planted instructions and declared danger | agentfox scan skills |
| Shadow and unowned agents seen in traffic | agentfox agents list |
Contain
What may each agent do, with which values, from which sources?
| You want to | Run |
|---|---|
| Tool permissions: declare impact, grant with argument limits | agentfox permit grant |
| Learned permissions: grants drafted from recorded calls, approved by a person | agentfox policy proposals from-traffic |
| Provenance: refuse untrusted values in irreversible calls | agentfox declare tool |
| Blast radius: what an agent reaches, and what a SQL or shell call would do | agentfox agents lineage |
| Approvals for escalated calls, and the kill switch | agentfox agents quarantine |
| Loop and budget limits on a run | agentfox policy effective |
Screen
Does the text itself look like an attack, a secret or personal data?
| You want to | Run |
|---|---|
| Detectors on nine surfaces: injection, PII, secrets, safety, schema | agentfox doctor |
| Tune them: feedback, suppressions, simulate, canary | agentfox policy simulate |
Ground
Is the answer inside what the agent knows, for someone allowed to see it?
| You want to | Run |
|---|---|
| Answerability: a knowledge boundary and forced abstention | agentfox declare boundary |
| Entitlement: what each end user may retrieve | agentfox permit user |
| Sources: authority tiers and freshness | agentfox declare source |
Prove
Can you show someone else what happened, and that the controls hold?
| You want to | Run |
|---|---|
| One-page summary for whoever signs off | agentfox report |
| Tamper-evident audit chain | agentfox report verify |
| Evidence package an auditor verifies without you | agentfox report evidence |
| Compliance posture (draft mappings) | agentfox report status |
| Red team the deployed configuration | agentfox test redteam |
| Eval suites and a CI regression gate | agentfox test gate |
Pick your integration
The same policy and the same decision record, bound wherever your agent already runs. Each place sees a different surface, so connecting one does not cover the others. Control points lists what each one sees.
| Your agent is | Use | What it governs |
|---|---|---|
| Python, any framework that calls OpenAI, Anthropic, LiteLLM or LangChain | agentfox.auto() | Every model call in the process, and every tool call in a response before your code runs it. Not the OpenAI Responses API, and not tools your code calls without the model asking. |
| A LangGraph graph | AgentFoxGuard node wrappers | Retrieval, model and tool nodes. An escalation becomes LangGraph's own interrupt(). |
| An MCP client or server | agentfox scan mcp and the MCP governor | Server reach and pinning before anything runs; at call time, tool digests (a server that changed after review) and results checked as untrusted input. |
| A coding agent (Claude Code) | agentfox admin hooks install | The prompt, each tool call before it runs, and each result. On this machine only. |
| Any language | agentfox serve, the gateway | A drop-in OpenAI or Anthropic endpoint, and POST /v1/guard/tool_call to ask about one call before you run it. |
Every page
- Overview: What AgentFox does, and the path through it.
- Quickstart: Scan, watch, contain and report in ten minutes.
- Concepts: Agents, tools, grants, provenance, policies and findings.
- Install and configure: Extras, where state lives, agentfox.toml.
- Audit a repository: Inventory, the lethal trifecta, and a CI gate.
- Monitor connected sources: Rescan repos, APIs and MCP servers on a schedule and on push.
- One line in Python: agentfox.auto(): observe, then enforce.
- Contain tool calls: Declarations, grants, provenance, learned permissions.
- LangGraph: Guard the retrieval, model and tool nodes.
- MCP servers: Scan configs, pin tools, govern calls.
- Coding agents: Claude Code hooks and the coding-agent pack.
- Any language: the gateway: The proxy and the guard API over HTTP.
- Retrieval and answers: Who may see what, and when to say I don't know.
- Approvals and the kill switch: Escalations, hand-offs, quarantine.
- Business rules: Threshold ladders, and policy compiled from prose.
- Red team and evals in CI: Probe the deployment; fail the build on regression.
- Probe deployed agents: Scheduled, opt-in attacks on a running agent.
- Tune detectors: Feedback, suppressions, simulate, canary.
- Prove it to an auditor: The report, evidence packages, compliance.
- Traces and integrations: OpenTelemetry, Langfuse, LangSmith, SIEM, webhooks.
- Tour of the web app: Sign in, workspaces, and where everything is.
- Start here and connect: The checklist, GitHub repos, API tokens.
- Agents: Registry, risk, knowledge boundary, kill switch.
- Findings: What needs a person, and why.
- Traces: One request, every check, and why it was blocked.
- Policies and tuning: Rules, modes, simulation, canary, detector tuning.
- Approvals and escalation: The queue a person works from.
- Access control and sources: End-user entitlement and verified sources.
- Evaluation: Suites, runs, SLOs and red-team campaigns.
- Compliance: Controls, frameworks, evidence, board snapshot.
- Playground: Attack a live agent with no account.
- CLI: Every command and option, generated from the CLI.
- HTTP API: Every route, generated from the gateway.
- Python SDK: auto(), AgentFox, sessions, integrations.
- Policy language: Rule schema, packs, modes, hierarchy.
- Detectors and findings: Surfaces, detectors, finding types.
- Configuration: Environment variables and agentfox.toml.
- Self-hosting: Docker Compose, Render, or a Python app.
- Claude Code plugin: Skills, slash commands and a read-only MCP server.
- Benchmarks: What was measured, and where each result stops.
- Limits: What is only as good as your declarations.
- Support: What to run before you file an issue.