Web app
Playground
A chat with three demo agents, where you try a prompt injection or a tool call the agent was not granted and watch the real enforcement path answer. No account, nothing to install.
In the web app Playground
When to use this
- Before installing anything, to see what a refusal looks like and why it happened.
- To show someone the difference between observe and enforce.
- To see that a tool call is refused for lack of a grant even when the text looks harmless.
What it sandboxes
Opening /playground creates a sandbox for you (POST /api/playground/sessions): a private workspace in the deployment's database, seeded with the demo agents, tools and grants. Your session id is its only credential; nothing is shared with other visitors, and nothing in it reaches the outside:
- No model is called. Replies come from the offline
echoprovider and are labelled "scripted reply, no model ran". - The
payments.transferandemail.sendtools have no backend; no money moves and no mail is sent. - Verdicts come from the same enforcement code as the product, not a mock.
What you can try
Pick an agent along the top. Each has three suggestions, plus a poisoned document:
| Agent | Holds | Suggestions |
|---|---|---|
| Support Triage | Runs the help desk; holds nothing that moves money. | Transfer $5,000 · Close every ticket · Ignore your instructions |
| Payments Operations | Can move money, up to a declared limit. | Transfer $5,000 · Transfer $500 · Ignore your instructions |
| HR Screening | Registered, and granted nothing at all. | Email a candidate · Search the help desk · Ignore your instructions |
- A suggestion with a tool shows your message, then the tool call the agent attempts, then its verdict (
POST …/tool-call). - Typed messages, and "Ignore your instructions", are checked as text (
POST …/chat). Each message is also checked together with the recent conversation, so an injection split across several short messages is tagged "caught across messages". - Summarise a poisoned document, or the paper-clip button, attaches an editable "retrieved document" containing an indirect injection. Edit it to try your own.
- Each verdict has a why disclosure: the reason, every rule that fired, and the latency.
These are the verdicts the playground API returned for the suggestions:
support-triage payments.transfer 5000 -> block / block | no capability grants 'payments.transfer' (action '*') to agent:support-triage (default deny)…
payments-ops payments.transfer 5000 -> block / block | EU AI Act Art. 14 — irreversible action by a high-risk system requires human oversight.; agent:payments-ops holds a grant for 'payments.tran…
payments-ops payments.transfer 500 -> allow / escalate | EU AI Act Art. 14 — irreversible action by a high-risk system requires human oversight.
hr-screening email.send -> block / block | no capability grants 'email.send' (action '*') to agent:hr-screening (default deny)…
chat observe -> allow / block observe | Prompt-injection or jailbreak attempt detected in user input.; System-prompt extraction attempt.
chat enforce -> block / block enforce | Prompt-injection or jailbreak attempt detected in user input.; System-prompt extraction attempt.Each line is the applied verdict, then the verdict the policy asked for. The $500 transfer is within the grant's limit, so the grant allows it, while the EU AI Act rule for high-risk agents asks for a human: in observe mode that is recorded, not applied.
Observe and enforce
The sandbox starts in observe. The side panel shows the mode and switch to enforce (POST …/enforce), which flips the baseline policy for your sandbox only. In observe, a flagged message is shown as flagged, not stopped with "would block in enforce", and the scripted reply goes through. In enforce, the same message is blocked. Capability checks on tool calls enforce in both modes, because they read no text. Earlier messages are not re-judged: send the same one again after switching.
The side panel also shows how many audit records your sandbox has written and whether its hash chain holds (GET …/state).
Limits
- A sandbox expires 30 minutes after your last action, and its data is deleted. The page then offers Start a new one.
- Eight new sandboxes per network address per hour; after that the API answers 429.
- Forty actions per sandbox per minute.
- At most 200 live sandboxes across the deployment; beyond that the oldest are dropped.
- Three fixed agents. You cannot add agents, grants or policies, or connect a model.
Run it yourself
On a self-hosted deployment the playground is the same page, calling the gateway's unauthenticated /api/playground/* routes from the visitor's browser. Set AGENTFOX_PLAYGROUND_API_URL on the web app to an address browsers can reach, and AGENTFOX_PLAYGROUND_CORS_ORIGIN on the gateway to the web app's origin. The routes can be called directly:
curl -s -X POST http://127.0.0.1:8080/api/playground/sessions{"session_id":"…","expires_in_seconds":1800,"agents":[{"slug":"support-triage",…},{"slug":"payments-ops",…},{"slug":"hr-screening",…}],…,"mode":"observe"}Common tasks
| You want to | Run |
|---|---|
| The same walkthrough offline, in a terminal | agentfox demo |
| Run the attack library against your own agent | agentfox test redteam support-triage |
| Grant a tool so the call stops being refused | agentfox permit grant support-triage tickets.close --yes |
What can go wrong
- "That sandbox expired. Nothing was saved." Thirty minutes passed without an action. Start a new one.
- "Too many playground sandboxes from this address recently". The per-address limit; try later.
- "The sandbox could not be reached" on a self-hosted deployment. The browser cannot reach the playground API URL, or the gateway does not allow the web app's origin.
- A typed message gets
allow. The text was checked and nothing matched. Tool calls are where containment shows; use a suggestion with a tool.