Guide
Prove it to an auditor
Turn what AgentFox recorded into something another person can check: a one-page summary, an evidence package with a standard-library verifier, and control status computed from what your agents actually did.
When to use this
- Someone asks what your agents did last week and what was stopped. Send them
agentfox report. - An auditor, a customer's security team or your own risk function wants records they can verify without trusting you. Build an evidence package.
- You need to know where you stand against the EU AI Act, ISO/IEC 42001, NIST AI RMF, SOC 2, OWASP or MITRE ATLAS, and what a reviewer has to sign off before you can make that claim.
Everything here reads what was recorded while your agents ran. If nothing was recorded, there is nothing to prove: start with One line in Python or the gateway, and come back once traffic has flowed.
| You want to | Run |
|---|---|
| One-page summary of what happened | agentfox report |
| Evidence package for an auditor | agentfox report evidence --agent payments-ops --since-days 30 |
| Check the audit chain | agentfox report verify |
| Anchor the chain with a signed checkpoint | agentfox admin checkpoint |
| Control posture for one framework | agentfox report status --framework eu-ai-act |
| Hand a reviewer what they must sign | agentfox report review-packet --framework soc2 --out soc2-review.md |
The one-page summary
agentfox report with no subcommand prints report summary: what is running, what was contained, what observe-mode rules would have stopped, and coverage against the OWASP agentic threat list. It is written for a person who has never opened AgentFox. Narrow it to one agent and a shorter window:
agentfox report summary --agent payments-ops --since 24h# AgentFox summary
*payments-ops, 2026-10-04 to 2026-10-05. Generated 2026-10-05T15:01:09+00:00.*
**In short:** 1 agent(s) under management; 2 risky action(s) stopped or held for approval; 3 more that observe-mode rules recorded but did not stop; 1 agent(s) with a risky combination.
## What is running
- **Agents:** 1
- payments-ops — production, high risk, marcus@example.com, demo data
…
## What was contained
2 action(s) stopped or held for approval. By cause:
- Needs a person's sign-off: 1
- Data from an untrusted source: 1
- Outside the limits of its permission: 1
Examples:
- payments-ops tried to payments.transfer with data that came from another tool's output (held for approval)
- payments-ops tried to payments.transfer outside the limits of its permission (contained)
- payments-ops tried to payments.transfer, which needs a person's sign-off first (held for approval)
## Still in observe mode
These policies record what they would do without acting on it: Baseline runtime guardrails, EU AI Act — high-risk system controls.
3 action(s) would have been stopped or held. By cause:
- Needs a person's sign-off: 3
…
## Coverage against OWASP Agentic AI — Threats and Mitigations (T1–T15) — DRAFT — UNVERIFIED
> The mapping from these threats to our controls is an engineering draft. It has not been reviewed by compliance counsel or an auditor and is not a certification.
| Threat | Status |
| --- | --- |
| Memory Poisoning | Watching only — nothing blocked |
| Tool Misuse | Blocking |
…
Controls with working evidence: 35 of 43; 2 have no evidence yet.One action can have more than one cause, so the cause counts can add up to more than the headline number.
The options:
--agent/-a: only this agent. Repeat it for several. Default: all agents.--since: how far back, as24h,7d,2wor30d. Default7d.--format/-f:md(default) orhtml. Anything else exits 2 withunknown format 'pdf' — use md or html.--out/-o: write to a file instead of printing.
agentfox report summary --since 30d --format html --out summary.htmlwrote summary.htmlThe HTML file is self-contained (inline styles, light and dark), so you can attach it to an email or a ticket.
Build an evidence package
Build it
bash agentfox report evidence --agent payments-ops --since-days 30 --requested-by auditor@example.comOutput evidence package …/var/evidence/evd_01m469avnz63xaz363.zip period 2026-09-05 → 2026-10-05 agents 1 traces 0 decisions 3 audit entries 27 audit payloads withheld 22 checkpoints 1 approvals 2 eval runs 0 findings 4 control statuses 43 reviewed mappings 0 draft mappings included 317 chain verification valid Verify independently: unzip, then `python3 verify_chain.py`The zip lands in
var/evidence/under your state directory (AGENTFOX_STATE_DIR; see Install and configure). Building a package writes anevidence.exportedentry to the audit chain, with--requested-byas the actor: who pulled the evidence is itself evidence.Look inside
bash unzip -l evd_01m469avnz63xaz363.zipOutput Archive: evd_01m469avnz63xaz363.zip Length Date Time Name --------- ---------- ----- ---- 2939 10-05-2026 20:31 SUMMARY.md 4651 10-05-2026 20:31 SUMMARY.html 34265 10-05-2026 20:31 audit_entries.json 268 10-05-2026 20:31 audit_checkpoints.json 2 10-05-2026 20:31 traces.json 19176 10-05-2026 20:31 decisions.json 7130 10-05-2026 20:31 policy_versions.json 372 10-05-2026 20:31 agents.json 1301 10-05-2026 20:31 approvals.json 1956 10-05-2026 20:31 eval_runs.json 8989 10-05-2026 20:31 findings.json 29862 10-05-2026 20:31 control_status.json 69441 10-05-2026 20:31 framework_mappings.json 234 10-05-2026 20:31 risk_assessments.json 176 10-05-2026 20:31 chain_verification.json 3545 10-05-2026 20:31 verify_chain.py 3565 10-05-2026 20:31 README.txt 3225 10-05-2026 20:31 manifest.json --------- ------- 191097 18 files
| File | What it holds |
|---|---|
README.txt | What each file is, how to verify, and the package's limitations. Start here. |
SUMMARY.md, SUMMARY.html | The same one-page summary as agentfox report, for the package's scope and period. |
audit_entries.json, audit_checkpoints.json | The hash-chained audit log for the period, and the signed checkpoints over it. |
decisions.json, policy_versions.json | Every policy decision with the rules that fired, and the exact policy version in force. Policy versions are immutable. |
traces.json | Full execution paths for the agents in scope. Content is redacted at capture. |
approvals.json, findings.json, eval_runs.json | Who approved what, what the platform raised, and evaluation results. |
control_status.json, framework_mappings.json, risk_assessments.json | Control status computed from telemetry, every control-to-clause mapping (each with a chip of REVIEWED or DRAFT — UNVERIFIED / NOT LEGAL ADVICE), and per-agent risk. |
chain_verification.json | AgentFox's own verification result. The auditor should not rely on it; that is what the next file is for. |
verify_chain.py | A standard-library verifier that re-derives every digest from audit_entries.json. |
manifest.json | Scope, counts, code and catalog versions, and the SHA-256 of every other file. |
A draft mapping row in framework_mappings.json looks like this:
{
"control_key": "…",
"framework": "eu-ai-act",
"reference": "Art. 11 — technical documentation",
"review_status": "draft",
"reviewed_by": null,
"chip": "DRAFT — UNVERIFIED / NOT LEGAL ADVICE"
}How an auditor verifies it without AgentFox
The recipient needs the zip and any Python 3. They do not install AgentFox, and nothing calls your API. This was run with the system Python 3.9, in isolated mode and with an empty environment, where import agentfox fails:
unzip evd_01m469avnz63xaz363.zip -d evidence && cd evidence
env -i /usr/bin/python3 -I -c "import agentfox"
env -i /usr/bin/python3 -I verify_chain.py…
ModuleNotFoundError: No module named 'agentfox'
note: AGENTFOX_AUDIT_KEY not set - checkpoint signatures not verified
entries checked: 26 (seq 1..26)
CHAIN INTACTThe script exits 0 when the chain is intact and 1 otherwise. It recomputes each entry's payload digest and entry digest, checks every entry links to the one before it, and reports gaps in the sequence. Here is what an edit looks like. This changes a recorded block into an allow:
import json
rows = json.load(open("audit_entries.json"))
row = next(r for r in rows if r["seq"] == 8)
row["payload"]["verdict"] = "allow" # pretend the block never happened
json.dump(rows, open("audit_entries.json", "w"), indent=2)python3 tamper.py && python3 -I verify_chain.py; echo "exit $?"note: AGENTFOX_AUDIT_KEY not set - checkpoint signatures not verified
entries checked: 26 (seq 1..26)
CHAIN INVALID - 1 break(s):
seq 8 payload_mismatch: payload does not match its digest
exit 1Check the manifest too
verify_chain.py covers the audit log. To check that no other file in the package changed after it was built, compare each file against the SHA-256 in manifest.json:
import hashlib, json, sys
manifest = json.load(open("manifest.json"))
bad = [
name for name, meta in manifest["files"].items()
if hashlib.sha256(open(name, "rb").read()).hexdigest() != meta["sha256"]
]
print("manifest:", "all files match" if not bad else "changed: " + ", ".join(bad))
sys.exit(1 if bad else 0)manifest: all files matchAfter the edit above, the same script prints manifest: changed: audit_entries.json and exits 1.
Checkpoint signatures: the part that needs a key
Someone who can rewrite the whole package can also recompute every digest and the manifest. What they cannot forge without the signing key is a checkpoint: an HMAC over the chain head at a point in time. Give the auditor the key through a separate channel and they verify the checkpoints as well:
AGENTFOX_AUDIT_KEY="$KEY_FROM_YOUR_OPERATOR" python3 -I verify_chain.pyWith the right key the output is the same as above, without the note. With the wrong key:
entries checked: 26 (seq 1..26)
CHAIN INVALID - 1 break(s):
seq 25 checkpoint_signature: checkpoint signature mismatchCheck the chain in place
On the machine that runs AgentFox, agentfox report verify re-derives the live chain from the database and checks every checkpoint with the configured key. Use --start and --end for a range of sequence numbers.
agentfox report verify
agentfox report verify --start 5 --end 12chain: 28 entries, head seq 28, 1 checkpoints
CHAIN INTACT — 28 entries verified (seq 1..28)
chain: 28 entries, head seq 28, 1 checkpoints
CHAIN INTACT — 8 entries verified (seq 5..12)After someone edits a row directly in the database (here, the same change from block to allow on sequence 8), it exits 1:
chain: 28 entries, head seq 28, 2 checkpoints
CHAIN TAMPERED — 1 break(s)
seq 8 payload_mismatch: payload does not match its recorded digestIf the signing key changed since a checkpoint was written, it reports checkpoint seq 25 signature mismatch — checkpoint forged or key changed and also exits 1. Run it in CI or on a schedule and alert on a non-zero exit.
Write a checkpoint
AgentFox writes a checkpoint automatically every AGENTFOX_AUDIT_CHECKPOINT_INTERVAL entries (100 by default). agentfox admin checkpoint writes one now, over the current head. Do it right before you build a package, so the package ends on a signed anchor:
agentfox admin checkpointcheckpoint seq 28 digest a69d2a0f38be35e3…Control posture and frameworks
AgentFox ships a catalog of 43 controls and maps them to seven frameworks. A control's status comes from recorded activity, not from a form: whether the evidence it needs exists, whether its rule passed, and whether the chain is intact. A control whose evidence source produced nothing reads not_implemented.
Recompute status from telemetry
bash agentfox admin catalog computeOutput computed 43 controls over 30 days 35 effective · 3 degraded · 2 failing · 2 not implemented--window-dayschanges the window (default 30). The remaining control in this example isnot_applicable: its capability was never used.Read the posture
bash agentfox report status agentfox report status --framework eu-ai-actOutput all frameworks — 43 controls 35 effective · 3 degraded · 2 failing · 2 not implemented effectiveness 88%--verboseadds one row per control with the rationale behind its status, for exampleowned_agents / total_agents = 2/4 = 50.0% (effective at 100%), and, with--framework, the gaps the product declares it does not cover (for the EU AI Act: conformity assessment, registration in the EU database, and the fundamental rights impact assessment filing).Framework keys:
eu-ai-act,iso-42001,nist-ai-rmf,soc2,owasp-llm,owasp-agentic,mitre-atlas. A misspelled key is not an error: it prints0 controlsand exits 0.See coverage and review status per framework
bash agentfox report frameworksOutput framework controls mapped mappings reviewed status EU AI Act 43/43 70 0 draft NIST AI RMF 43/43 74 0 draft ISO/IEC 42001 43/43 51 0 draft SOC 2 43/43 64 0 draft OWASP LLM Top 10 22/43 26 0 draft OWASP Agentic Threats 19/43 22 0 draft MITRE ATLAS 9/43 10 0 draft All mappings are engineering drafts. They are not legal advice and ship inside evidence packages tagged DRAFT — UNVERIFIED / NOT LEGAL ADVICE until reviewed.
Risk, obligations and the board view
agentfox report riskagent risk tier EU class residual assessed review
support-triage limited — — no —
payments-ops high high medium yes 2027-09-30
hr-screening limited — — no —
marketing-copy-bot limited — — no —agentfox report obligationsdate framework obligation status agents build by
2025-02-02 eu-ai-act Prohibited AI practices live 4 —
2026-08-02 eu-ai-act Transparency obligations live 3 —
2026-12-02 eu-ai-act General-purpose AI model obligations upcoming 4 2025-12-02
…
2027-12-02 eu-ai-act Standalone high-risk AI systems upcoming 1 2026-12-02
2028-08-02 eu-ai-act High-risk AI embedded in regulated products upcoming 1 2027-08-03The obligation calendar is part of the same draft catalog. A row is a date to plan for, not a record that the duty was met, and its dates are not legal advice: check them against the regulation's current text.
agentfox report board╭─────────────────╮
│ AI risk posture │
╰─────────────────╯
agents under management 4 (1 shadow, 2 unowned)
high-risk agents 1 ['payments-ops']
unassessed agents 3 ['support-triage', 'hr-screening', 'marketing-copy-bot']
open findings 26 {'high': 7, 'critical': 9, 'medium': 10}
controls with evidence 35 of 43 effective (40 assessed, 2 with no evidence yet)
live obligations 2
upcoming (24mo) 5
Control statuses are computed from telemetry over the stated window. Framework mappings are DRAFT and have not
been reviewed by compliance counsel — see the coverage and gap declarations per framework.Sign off a mapping
A draft mapping becomes reviewed when the person accountable for the claim signs it off. Give them the review packet first: one section per control, with what it is meant to achieve, what evidence it produces, and every clause it is mapped to.
agentfox report review-packet --framework soc2 --out soc2-review.mdwrote soc2-review.md 64 mapping(s), 64 draft# Compliance mapping review packet — soc2
64 mapping(s), 64 awaiting review.
For each row: does this control, as implemented, support the clause claimed? Approve with
`agentfox report signoff <control> --framework soc2 --reviewer "<your name>"`, optionally `--reference` for a single clause.
…The reviewer records their decision with agentfox report signoff, using the control key from the packet's section heading. --reference signs off one clause only; without it, every clause for that control in that framework is marked reviewed.
agentfox report signoff <control> --framework soc2 --reviewer "Dana Reyes, external auditor" --reference CC7.2It prints how many mappings it marked reviewed. Afterwards agentfox report frameworks shows SOC 2 with 1 under reviewed, the packet says 63 awaiting review, and the next evidence package carries that row with the REVIEWED chip and the reviewer's name.
In the web app Compliance → Frameworks → Review mappings
Retention and legal hold
What exists today is a record, not an enforcement mechanism. Be precise about this with an auditor.
- Retention policies are rows (data class, days to keep, fields to redact) that you read with
GET /api/retentionor on the Retention tab. A seeded environment has four. There is no command, route or screen to add or change one, and nothing in AgentFox deletes or redacts data when a period runs out. - Legal holds are placed with
POST /api/legal-holds(rolecompliance) or from the Retention tab. Placing one writes alegal_hold.placedentry to the audit chain. There is no route to release a hold, and because nothing purges data on a schedule, a hold does not change what is kept. - Redaction at capture is on by default (
AGENTFOX_REDACT_AT_CAPTURE). In the words of every package's README: content in traces is redacted at capture, and detector findings record entity type, location and a masked sample rather than the sensitive value.
Placing a hold over the API, with a token issued by agentfox admin auth issue for a user with the compliance role:
curl -s -X POST http://127.0.0.1:8080/api/legal-holds \
-H "Authorization: Bearer $AGENTFOX_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"scope":{"agents":["payments-ops"],"from":"2026-09-01"},"reason":"Preservation request, matter 2026-114"}'{"id":"hld_01m469qszk6fk4h23b","placed_at":"2026-10-05T15:08:26.483424+00:00"}curl -s http://127.0.0.1:8080/api/retention -H "Authorization: Bearer $AGENTFOX_TOKEN" | python3 -m json.tool{
"policies": [
{
"data_class": "prompt_content",
"retain_days": 90,
"redact_fields": [
"content",
"messages"
]
},
…
{
"data_class": "audit",
"retain_days": 2555,
"redact_fields": []
},
…
],
"legal_holds": [
{
"id": "hld_01m469qszk6fk4h23b",
"scope": {
"agents": [
"payments-ops"
],
"from": "2026-09-01"
},
"reason": "Preservation request, matter 2026-114",
"placed_by": "dana@example.com",
"placed_at": "2026-10-05T15:08:26.483424+00:00",
"released_at": null
}
]
}In the web app Compliance → Retention & legal hold
What can go wrong
- The package is nearly empty. Nothing was recorded in the period. Check
--since-days, or--fromand--to(YYYY-MM-DDor ISO-8601; a bare--todate includes the whole day), and that your agents run through AgentFox. audit payloads withheldis not zero. Expected with--agent. The audit chain is exported for the whole period so it still verifies, but entries about other agents carry"payload_withheld": trueand no payload or subject. Decisions, findings, approvals and eval runs are limited to the agents named.verify_chain.pysayscheckpoint_signature. The key you passed is not the one the checkpoints were signed with, or the key was changed after they were written.report status --frameworkshows0 controls. The framework key is misspelled. Use one of the keys listed above.- A control reads
not_implemented. Its capability was never exercised in the window. Runagentfox report status --verboseto see which records it looked for.
Limits
- A package shows what was recorded. It cannot show what never passed through AgentFox, and it does not establish that the controls in scope were the right ones.
effectivemeans a control's evidence was present and its rule passed over this scope and window, not that the control is effective in general.- Framework mappings, the threat coverage table and the obligation calendar are drafts until signed off. They are a starting point for a reviewer, not a legal conclusion.
- Without the signing key, an auditor can detect accidental damage and partial edits, but not a complete, consistent rewrite of the package.
More on what AgentFox cannot see: Limits.