Operate
Limits
What AgentFox does not do, what is only partly built, and what depends entirely on what you declare. Written down so you do not discover it later.
When to use this
Before you rely on a control, and before you tell someone else what AgentFox covers. The per-capability status below comes from docs/status.md in the repository, which is generated by probing the code rather than written by hand. Detection's measured misses are on Benchmarks.
It is only as good as your declarations
- Impact is a floor. A destructive tool declared
readis treated as read by everything downstream. A shell is one tool for bothlsandrm -rf. Tools thatagentfox.auto()registers get an impact guessed from their name, markedinferreduntil you confirm it withagentfox declare tool. - Output trust is a claim you make. Declaring a tool's output
trustedstops values copied out of it counting as untrusted. Declare it for a system of record you control, never for anything a stranger can write. - The scan reads names.
agentfox scanclassifies tools from names and descriptions. It can flag a harmless tool and miss one nameddo_thing; an unrecognised MCP server is reported as unknown, never as safe. - You declare the estate yourself. Principals, grants, source tiers and row-scoped tables live in AgentFox. The seams for Okta, DataHub and OpenFGA exist; the integrations do not.
agentfox doctor grades the declarations you have, and agentfox declare list tools shows which impacts are still guesses:
agentfox declare list tools tool impact output triggers
crm_lookup read (inferred — confirm with `agentfox declare tool`) untrusted —
…Containment costs legitimate work
Under the default session-level provenance, once an agent has read untrusted content, every later irreversible call needs a person. On AgentDojo only 24 of 97 benign tasks ran without an escalation. Per-argument provenance lets more through and misses attacks whose values do not appear in an argument. Neither is free; see the utility table.
What each integration does not see
agentfox.auto(): not the OpenAI Responses API, and not a tool your code calls without the model asking for it.- The gateway: only what your code sends to it.
- Coding-agent hooks: only the agent on the machine they are installed on. A session running in the vendor's cloud is invisible to them, and
PostToolUsecannot withdraw a call that already ran. - Detectors: text only. No images, audio or files.
Compliance mappings are drafts
The 43 controls are mapped to the EU AI Act, ISO 42001, NIST AI RMF, SOC 2, OWASP LLM, OWASP Agentic and MITRE ATLAS by engineers, from the framework texts. Counsel has not reviewed them. Evidence packages include them labelled DRAFT — UNVERIFIED / NOT LEGAL ADVICE. Control status is computed from what agents did; it is not a certification. See Prove it to an auditor.
Version 0.3
- No SSO, OIDC or SCIM. The first operator comes from GitHub sign-in or
agentfox admin seed. - One organisation per deployment, enforced at the database session.
- Text only.
Partly built
Each of these works, with a specific piece missing. From the generated status table (25 of 41 tracked capabilities built, 16 partial, none absent):
| Area | What is missing |
|---|---|
| Registry and discovery | No connector-based estate discovery. Agents are found by scanning code, by their traffic, and by registration. |
| Identity | No live identity provider. No Entra or Okta integration; principals and grants live in AgentFox. |
| Authentication | Operator tokens and the development-mode gate ship. OIDC and SCIM do not. |
| Action assurance | Cascade analysis is only as good as the trigger declarations it is given: an undeclared webhook stays invisible. Shell analysis is a deny-list, not a parser. |
| Entitlement | The OpenFGA adapter is a declared seam, not an implementation. |
| Source provenance | No catalog ingestion (DataHub, OpenMetadata, Unity). You declare sources yourself. |
| Context integrity | Ingestion and retrieval gates report problems. Semantic chunk repair and automatic re-extraction of a corrupt document are not built; dropping the document is the operator's decision. |
| Evaluation | No annotation queue for human review of borderline eval results. |
| Failure attribution | Attributes a failure only against constraints that were written down. An expectation nobody typed is invisible. |
| Compliance | No dynamic risk scoring and no workflow engine. |
| Loop governance | Catches identical repeats, alternating cycles and steps with no new observation. An agent that is wrong but varied still looks like an agent working. |
| Async work | Jobs run in process with retries and a dead letter. No Redis or SQS backend. |
| Persistence at scale | Pooling ships, and SQLite is refused for multi-worker servers. Scale-out under load is untested. |
| Fail-open budgets | Budgets and rate limits are per process, so N workers get N times the declared budget. |
| Tool contracts | Data-access scoping is only as good as the table declarations; an undeclared table is reported, never assumed safe. Register checks are lexical. |
| Business rules | Policy compiled from prose is deterministic; prose with no parseable structure is reported as inexpressible rather than guessed. No model-assisted path. |
Detection
Detectors raise an attacker's cost; they are not a defence on their own. An adaptive attacker that reads the verdict and retries gets 73% of the attacks the default detectors catch through within 50 attempts. The model-backed detectors are opt-in, need weights, and can time out on long inputs; with fail_mode left at open, a timeout lets the request through and records the gap. agentfox doctor says so on every run.