Runtime guardrails
Detectors read the text. A policy decides
Nine places content can enter or leave an agent, and a rule set you can read.
The problem
A scanner returns a score. A score is not a decision, and the gap between them is where most guardrail products leave the hard part to you: which surface was this, how much do we trust its source, what is this agent allowed to do, and is the answer different because a check was down.
Nine places content enters or leaves
The same sentence means different things depending on where it turned up, so the surface is part of the decision rather than metadata attached to it.
inputWhat the operator typed or pasted
outputWhat the model is about to say
tool_argsThe arguments of a call about to run
tool_resultWhat a tool sent back
retrievedA document pulled into context
memory_writeSomething being written to long-term memory
agent_messageA claim from another agent
reasoningThe model's own thinking, before it acts
RarecompletionThe claim that it finished
Rare- 01
Nine surfaces, not one
Input, output, tool arguments, tool results, retrieved documents, memory writes and messages from other agents — plus two most products do not have: the model’s own
reasoning, and itscompletionclaim that it finished. The same text means different things on different surfaces. - 02
The same detection, weighted by where it landed
Injection-shaped text in a retrieved document is an attempt. The same text in the model’s own reasoning is a compromise in progress, so it is caught at a lower confidence. That asymmetry is the argument for having surfaces at all.
- 03
Rules you can read, in packs you can choose
50 rules across four packs — a detector baseline, tool containment, an EU AI Act pack, and one tuned for coding agents. YAML, in the repository, with a lint that catches a rule shadowed by another and a rule whose conditions can never all hold.
agentfox policy lint - 04
Nothing blocks until you say so
Every pack ships in observe. It records the verdict it would have returned against real traffic, and you promote it when the counterfactual looks right rather than when the documentation says to.
agentfox policy observe baseline - 05
A check that cannot run says so
350ms for the whole request, 300ms for the pipeline, 40ms for any one detector. Over budget, a detector is marked degraded on that decision rather than quietly skipped — and four controls, including tenant isolation and the audit chain, may not be configured to fail open at all.
What the detectors do not do
They are pattern and classifier based, and a sufficiently novel phrasing gets through — which is the entire reason grants exist underneath them and why the benchmark is run with every detector switched off. On some individual detection tasks a specialised scanner is more precise than ours, and /compare names which.
Try to break it before you trust it
No account, no install, and the same enforcement code as the product.
pip install agentfox agentfox init && agentfox demo
Offline: no API key, no downloaded weights, no network egress.