Web app
Policies and tuning
Policies has three tabs: Rules (what each policy checks and whether it observes or enforces), Guardrail tuning (whether the detectors behind those rules are right, fast and not silenced), and Judgment posture (whether any text may leave the deployment to be judged).
In the web app Policies
When to use this
- To promote a policy from observe to enforce, after checking what it would block.
- To add or change a rule, for everyone or for one agent.
- To act on a false positive: suppress it for a while, or propose a threshold change.
Rules tab
/app/policies (filter with ?agent=). From top to bottom:
- Pending review: policies a scan proposed. They are already bound in observe, so they block nothing. Approve acknowledges one; Reject discards it.
- All policies: name and key, description, latest version, mode (
observe,enforce, orunbound), rule count, and Review & edit →. One policy usually covers many agents by name pattern, so filtering by agent shows every policy whose scope matches it. FromGET /api/policies. - Detectors: turned on, available but off, and not installed. Under it, Every check available (the detector catalogue) lists each check, the surfaces it reads, its state, its average time on your traffic, and what turns it on:
add <key> to enabled_detectorsfor an installed one, or thepip installline for one that is not. Checks wrapped from Guardrails AI Hub carry their own licences and none ships enabled. - Judgment tiers: a read-only summary of the judgment posture (see the third tab).
- Attack simulations: the built-in probe library (22 probes in this build) with category, surface, severity, OWASP and ATLAS ids, and a button to run them on Evaluation.
- Change proposals: there is no screen for proposals. This section is a table of the commands and HTTP routes (below).
policy version mode rules
baseline v1 observe 13
eu-ai-act-high-risk v1 observe 7
tool-containment v1 enforce 25A policy: rules, editor, versions, canary
In the web app Policies → a policy → Review & edit
/app/policies/<key> loads GET /api/policies/{key}. It shows the current mode and rule count, then Rules in this policy: each rule's description and id, when it fires in words ("on input · INJECTION detected, score ≥ 0.85"), effect and severity. Rules are checked in order; the first whose conditions match decides.
Edit these rules — builder or YAML opens the editor. A policy that has no rules yet (a new one, or a scan proposal) opens with one real example rule and a "Nothing is saved yet" banner.
- Rule builder: When this happens (injection, personal data, secrets, unsafe content, wrong output format), how confident (0.85, 0.7 or 0.5), where to check (input, output, retrieved, tool arguments, tool results), then do this (block, redact, mask, ask a human, decline to answer, just log), severity, and a note. Add this rule writes YAML into the editor below; nothing is saved.
- The YAML: every builder choice compiles to it, and you can edit it directly. The schema is in Policy language.
- Where does this apply? Level (
org,team,agent,user), a scope id or glob such aspayments-*(not used at org level), and compose:extendadds to broader rules,restrictmay only tighten,overridemay loosen rules the broader level marked overridable. The narrowest level wins on ties. - Validate runs the same check the engine runs at enforcement time (
POST /api/policies/validate) and reports the rule count and controls. It also runs the lint, so a rule that can never fire, or an unknown value such assurface: [toolargs], is reported as invalid. - Save new version stores an immutable new version (
POST /api/policies) and reloads. It never changes what is in force: the live version and its mode stay as they are, whatevermode:the YAML says, and the header shows the new one as saved, not live. A policy saved for the first time goes live in observe, which records and blocks nothing. - Simulate against recent traffic replays recorded decisions against the YAML in the editor (
POST /api/policies/simulate) and says how many would be newly blocked, newly escalated, newly allowed, or unchanged. - Promote to enforce makes the newest saved version live in enforce; Make vN live in observe (shown when a saved version is not live) makes it live in observe; Set observe demotes the live version (
POST /api/policies/{key}/mode, withversionwhen a saved version is being made live).
Worked example: an agent-level redaction rule
Write or build the rule
support-triage-pii.yaml key: support-triage-pii name: Support triage PII description: Redact personal data in what support-triage sends back. version: 1 mode: observe default_effect: allow fail_mode: open scope: agents: ["support-triage"] rules: - id: support.pii.output description: "The content contains personal information (names, emails, SSNs, phone numbers…)" when: surface: [output] detection: {entity_prefix: PII, min_score: 0.7} effect: redact severity: high reason: "Personal data in a support reply."Validate, then simulate
In the editor, or with the same checks from the CLI:
bash agentfox policy validate support-triage-pii.yaml agentfox policy simulate -f support-triage-pii.yaml --agent support-triageOutput valid — support-triage-pii v1, 1 rules, mode=observe controls: [] compiles to 33 lines of Rego support-triage-pii simulated against 7 decisions unchanged 7 newly blocked 0 newly escalated 0 newly allowed 0 No production traffic would newly block.The editor shows the same result as "Safe to promote — replayed 7 recent decision(s): 0 newly blocked, 0 newly escalated, 0 newly allowed, 7 unchanged." Each recorded decision is replayed with the candidate in place of its own pack (whichever version of it was in force then); rules that fired from other packs still count, so a change in the counts is a change in the whole outcome. Newly blocked, escalated and allowed decisions are each listed.
policy simulateexits non-zero when anything would be newly blocked, so it can gate a pull request.Save at agent level, then promote
Set Where does this apply? to
agent/support-triage/extend, press Save new version, then Simulate and Promote to enforce. From the CLI, promoting the live version is:bash agentfox policy enforce support-triage-pii
Version history
Every saved version, newest first: author, notes ("edited from the dashboard"), rule count, when.
Canary rollout
Shown once a policy has two or more versions. Pick a version and Start canary rollout (POST /api/policies/{key}/canary/start). The candidate takes a share of traffic in steps of 10%, 25%, 50% and 100%, against the stable version. The panel shows decisions and block rate for each side.
- Check health & advance (
/canary/advance) moves to the next step only when both sides have at least 20 decisions and at least an hour has passed at this step; if the candidate's block rate exceeds stable's by more than 15 points, it rolls back automatically. Until then it reports what it is waiting for. - Roll back now (
/canary/rollback) ends it, recording "manual rollback by <you>".
{"id":"cny_01m469s5s681tc1gyb","status":"rolling","percent":10,"step_index":0,"steps":[10,25,50,100],
"stable_version":1,"candidate_version":2,"max_block_rate_delta":0.15,…,"min_dwell_seconds":3600,
…,"min_sample":20,"started_by":"admin@example.com",…}There is no CLI command for canaries; use the page or the routes above.
Guardrail tuning tab
In the web app Policies → Guardrail tuning
Are the checks catching real problems, slowing agents down, or silenced? Filter by agent. Warning tiles appear only when something needs a person: checks that ran late or were skipped above 1%, and suppressions expiring this week. Then a strip: checks on, checks run in the last 7 days, the late-or-skipped rate, active suppressions.
- Suppressions: detector, scope (an agent or all agents, and an entity type), reason, hits (with a
never usedtag at zero), expiry countdown, and revoke (DELETE /api/guardrails/suppressions/{id}). Every suppression expires; a permanent exception looks exactly like a detector that stopped working. - Feedback log: the last 20 labels filed from traces: detector, entity, label, note, who, status, a link to the trace. An open false positive has suppress 30d, which creates a suppression for that detector, agent and entity for 30 days (
POST /api/guardrails/suppressions) and marks the feedback applied. Creating and revoking suppressions needs owner, admin or security. - How often each check is actually right (collapsed): labelled count (
too fewbelow five), false positives, precision, and a recommendation such asraise thresholdorinsufficient data, with its rationale. - How much each check slows things down (collapsed): per detector, version, runs, p50, p95 and worst case, with availability, against the budget (300 ms per call and 40 ms per detector by default).
detector scope reason hits expires
injection.heuristic support-triage INJECTION.INSTRUCTION_OVERRIDE quoted text in a ticket, not an instruction 0 never used 2026-11-04 revokeJudgment posture tab
In the web app Policies → Judgment posture
Whether optional evaluators may judge what code cannot, and whether a customer's text may leave the deployment for that. Loads and saves GET/PUT /api/judgment/posture.
- A banner says either "Nothing leaves this deployment" or "This tenant sends payloads to a third party". With egress off at deployment level (
allow_egress = false, the default), tiers that send data cannot be selected here. - Tiers:
deterministic,local_model,local_llm,jev,llm. Tiers that send data off the machine are marked; a disabled tier says why. - Personal data on the way out:
block,redactorallow. Options looser than the deployment's floor are disabled. - Backend and Fail closed (a tier's outage denies the request instead of passing it unjudged).
- Why is required and goes to the audit chain with the before and after.
- I am turning on a tier that sends data off this machine… must be ticked for any change in that direction. Tightening needs no confirmation.
- What each decision kind would use: for each kind of decision, which tiers decide it and which are refused, with the reason.
Only owner, admin and security can save; other roles see the form read-only. An unticked Fail closed saves as fail open.
Change proposals
AgentFox proposes changes to its own configuration rather than making them: threshold changes from your labels, and grants and tool declarations learned from traffic. A loosening is never applied automatically. There is no screen for proposals; the Rules tab lists the commands, which are:
| You want to | Run |
|---|---|
| Draft grants and declarations from recorded tool calls | agentfox policy proposals from-traffic --since 7d |
| Draft threshold changes from false-positive labels | agentfox policy proposals from-labels --days 30 |
| See what is waiting | agentfox policy proposals list --status proposed |
| Read one in full | agentfox policy proposals show PROPOSAL_ID |
| Approve one | agentfox policy proposals approve PROPOSAL_ID --actor you@example.com --note "…" |
| Put it into effect | agentfox policy proposals apply PROPOSAL_ID --actor you@example.com |
| Undo it | agentfox policy proposals rollback PROPOSAL_ID --actor you@example.com --reason "…" |
| Record whether it worked | agentfox policy proposals verify PROPOSAL_ID --actor you@example.com --note "…" |
Apply, rollback and verify need the policy_production permission (owner, admin, security). The CLI asks for --actor on every decision; over HTTP the actor is whoever the token belongs to. The in-app table shows the commands under their older proposals name; the current names are above. The full loop is in Contain tool calls.
Common tasks
| You want to | Run |
|---|---|
| List policies and modes | agentfox policy list |
| Validate a policy file | agentfox policy validate support-triage-pii.yaml |
| Replay traffic against a candidate | agentfox policy simulate -f support-triage-pii.yaml |
| Promote / demote | agentfox policy enforce baseline |
| Demote | agentfox policy observe baseline |
| What applies to one agent | agentfox policy effective --agent payments-ops |
What can go wrong
- "Run Simulate against the currently saved rules first" on Promote: the editor text changed since the last simulation. Simulate again.
- A policy went back to observe after you saved. The YAML said
mode: observe. Promote it again, or save with the mode you want. - The mode column says
unboundafter a canary. In this build the list reads the latest version's binding, which a canary start or rollback replaces.agentfox policy effective --agent …shows what is actually in force. - "Change these on Settings → Judgment posture" under Judgment tiers leads to a missing page; the setting is the Judgment posture tab.
- Tool containment blocks with no detector involved. It reads no text; it checks grants, provenance and blast radius. Tune it with grants, not suppressions.
Limits
- No screen for proposals, grants or tool declarations.
- Simulation replays recorded decisions; it cannot predict traffic you have not seen.
- Suppression length from the web app is fixed at 30 days.