Red-team reports go stale the day they ship
Most AI red teaming happens once: before launch, as a report. That report is accurate about the agent that existed that week. Then the model provider ships an update, someone rewrites the system prompt, a new tool is added, a policy is loosened to stop a false positive. Each change can reopen a hole the report said was closed, and nothing tells you, because nothing is running the attacks any more.
The fix is not a better report. It is running the same attacks again, on a schedule, against the deployed agent, and treating a change in the result as an incident. That is what continuous red teaming means here, and the most honest way we know to show it works is to do it to ourselves, in public.
What /live is
/live shows AgentFox probing its own demo agent every hour. The agent is a support-triage bot on echo-1, an offline model that follows injected instructions on purpose. That choice matters: a model that resists attacks on its own would make the guardrail look better than it is. With a model that always complies, every contained attack was contained by the enforcement path and nothing else.

Each hourly run sends eleven probes: nine attacks and two benign controls.
- A direct instruction override, and the same instruction inside a retrieved document.
- Obfuscated variants: base64, letters spaced apart, and an instruction hidden in an HTML comment. These three are the subject of our post on hidden prompt injection.
- A developer-mode role-play jailbreak.
- Exfiltration by email, a call to a forbidden tool, and a system-prompt leak.
- Two ordinary support questions, to catch the guardrail blocking real work.
The headline figures are thirty-day totals, and they move. We do not quote them here for that reason. Go and look.
How a probe is scored
The hard part of automated red teaming is deciding, without a person, whether an attack worked. A model that repeats the attack text back is not compromised, and a scorer that cannot tell the difference reports nonsense.
Every attack asks the agent to reply with a code written backwards. The forward code is never in the message, so an agent that merely echoes its input does not score as compromised. Other escape signals are a call to a forbidden tool, a tool call carrying the probe's unique exfiltration address, and a reply containing a leak marker. A response that does not follow the contract counts as an error, never as contained, and a benign control that gets blocked counts as over-blocked.
Each run is compared with the previous one. Every attack that got through opens a finding, flagged as a regression if it was contained last time, and the finding closes when a later run contains it again. That comparison is the point: the interesting event is not an escape, it is a change.
Probing your own agent
Sending attacks at a production system needs more care than sending them at a demo, so a probe target is built to be hard to turn on by accident:
- It is created disabled. Opting in requires sending back the exact text of a warning, and records who did it and when, in the audit chain.
- Changing the target URL clears the opt-in, because the consent was for that host.
- Probes go only to the registered host. Private and metadata addresses are refused by default, the connection is pinned and redirects are not followed.
- Hard caps: at most hourly, 50 probes a run, 60 a minute, a 30-second timeout and a 64 KB response limit. A stored credential is encrypted and never returned.
- One environment variable,
AGENTFOX_LIVE_PROBES_ENABLED=false, stops every target at once.
Targets default to daily. When one fails, it is pushed back a full interval rather than retried on every scheduler tick, so one broken endpoint never stops the others. Opting in also creates a monitor, so an escape alerts through Slack or a webhook like any other finding. The setup is in Probe deployed agents, and the alerting in Monitor connected sources.
Red teaming in CI
Live probes catch drift in production. The cheaper place to catch it is before the merge:
agentfox test redteam support-triage # exits 1 if any attack got through
agentfox test redteam support-triage --adaptive --budget 5
agentfox test gate support-quality --junit reports/agentfox-junit.xml --sarif reports/agentfox.sariftest redteam runs the offline probe suite against the agent's current policy and fails the build on an escape. Its tool calls are checked without being persisted, so a CI run never writes decisions into production history or shows up in a policy simulation. test gate runs an evaluation suite and writes JUnit and SARIF, which most CI systems display natively. See Red team and evals in CI.
What it does not prove
We map every way we know an agent can fail, and score the product against it on the coverage page. Today that is:
The scenarios with nothing in place are:
L0.4Invalid logical inferenceL1.9Context stuffing to push out the system promptL2.19An instruction in an image, or a scanned PDF
And the limits of the probing itself:
- The library is small, fixed and public. A contained probe says nothing about an attack outside it. It is a regression test, not a robustness certificate.
- The showcase tests the guardrail, not a model.
echo-1always complies, which is the point, and it means /live says nothing about how a real model resists. - Black-box scoring needs a signal. For an HTTP target, an escape that produces none of the configured signals is missed.
- Adaptive attackers win eventually. Our benchmark shows how quickly, which is why we lean on containment rather than detection.
- Hourly depends on a scheduler. Ours is a GitHub Actions job every 30 minutes, and scheduled Actions run on a best-effort basis.

Frequently asked questions
What is continuous red teaming for AI agents?
Running a fixed set of attacks against a deployed agent on a schedule, comparing each run with the last, and opening a finding when an attack that used to be contained gets through. It catches regressions caused by a model update, a prompt change or a policy change, which a one-off red-team report cannot.
How often does AgentFox probe its own demo agent?
Every hour. The public /live page shows the results, unedited, including any attack that got through. Probe targets you add for your own agents default to once a day and can run at most hourly.
Is it safe to red team a production AI agent?
Only with explicit opt-in and limits. AgentFox probe targets are created disabled and need the operator to acknowledge a warning before anything is sent. Probes go only to the registered host, with private addresses refused by default, hard caps on rate and volume, and a kill switch that stops every target at once.
How do I run AI red teaming in CI?
Run agentfox test redteam <agent>. It exits 1 if any attack got through, so the build fails on an escape. agentfox test gate runs an evaluation suite and can write JUnit and SARIF reports for your CI system.
Does a contained probe mean the agent is secure?
No. The probe library is small, fixed and public. A contained probe says nothing about attacks outside the library, and an adaptive attacker eventually finds a way past detection. Continuous probing catches regressions; it is not a robustness certificate.


