Operate
Benchmarks
What containment and detection were measured to do, what containment costs in escalated legitimate work, and where each number stops applying.
When to use this
Before you quote a number, decide on taint_scope, or turn on a judgment tier. Every figure on this page is bound to a results file in benchmarks/claims.yaml, and CI fails if a page quotes one that no longer matches its source. The full methodology is on Benchmarks and in each benchmarks/<name>/README.md.
Containment, with every detector off
| Benchmark | Result |
|---|---|
Containment scenarios (benchmarks/containment) | 8/8 attacks contained with zero detector signal; 4/4 legitimate calls still allowed. |
| AgentDojo, 97 user tasks and 949 attack pairs, ground truth replayed with provenance inferred from the real tool outputs; no model run | Session-level taint (the default): 588 of 588 attack pairs contained, but only 24 of 97 benign tasks (24.7% [17.2, 34.2]) ran without escalating to a human. |
What containment costs: AgentDojo utility
Containment is cheap to make complete and expensive to make useful. When provenance has to be inferred, a legitimate call that copies a value out of a tool output looks exactly like an attack, and it is escalated to a person. The whole trade, from benchmarks/agentdojo/README.md:
| Provenance | Benign tasks run without escalation | Attack pairs contained | Attacker write calls contained |
|---|---|---|---|
| Session-level taint (shipped default) | 24/97 (24.7% [17.2, 34.2]) | 588/588 | 702/702 |
| Argument-level taint | 37/97 (38.1% [29.1, 48.1]) | 527/588 | 641/702 |
| Argument-level taint, read-only tools exempt | 62/97 (63.9% [54.0, 72.8]) | 527/588 | 641/702 |
| Taken from the benchmark's labels (upper bound) | 97/97 | 588/588 | 702/702 |
| None: grants and impact tiers only | 97/97 | 0/588 | 61/702 |
- Session-level taint is complete and expensive. Once the agent reads any tool output, every later irreversible call counts as untrusted. It escalates three benign tasks in four.
- Argument-level taint misses 61 of 702 attacker write calls: identifiers shorter than six characters are never matched, and attacker text inside a longer argument is not found.
- Most of the remaining cost is not fixable by provenance alone. With read-only tools exempt, 43 of 73 legitimate write and irreversible calls (59%) are still escalated, and every one copies a value out of a tool output: paying a bill to an account number read from the bill. A user/tool-output label cannot tell those from an attack. Declaring a tool's output trusted is the lever you have.
- Labels are an upper bound. 97/97 with 588/588 needs provenance someone labelled correctly, which a deployment does not get for free.
- It is not AgentDojo's "utility under attack" metric, which needs a live model. No model was run.
The setting that picks the row is taint_scope (Concepts). The Quickstart shows the same trade on four tickets.
Detection, the layer we trust least
| Benchmark | Result | Where it stops |
|---|---|---|
| Adaptive attacker vs. the default detectors | 73% attack success at 50 attempts on the attacks they catch | 38/38 of those bypasses were still contained at the action. |
| Adaptive red team vs. the capability and taint layer | provenance 0/30, tool_scope 0/29, structural 0/105 escapes | structural is after a nested-argument fix this benchmark found. A posture delta, not a pass rate. |
| Opt-in classifier ensemble, SPML | 85.6% recall | Needs the classifiers extra and weights; not the shipped default, and on long prompts it mostly times out. |
| Crescendo (multi-turn) | 10/13 detected; control false positives 0/9 | Small set; the signal is the trajectory, not any single turn. |
| Cross-lingual parity | 19/19 matched pairs reach the same verdict | Benign non-English text is still flagged more often. |
Judgment tiers, opt-in
Judgment tiers add a model to the decisions where measurement says a model wins, and are forbidden from the ones where it loses. They are off by default and need allow_egress for the hosted ones (Configuration).
- 160/165 of the injection payloads that defeated our own pattern detectors are caught once a tier is on, at the cost of a network round trip per guarded call.
- The deterministic answerability check abstains on 57/676 contested questions; with a judgment tier that becomes 572/676, with more over-refusal on answerable ones.
- SQL blast-radius analysis is unchanged with every tier on, because no model may decide it.
Where a standard score does not apply
Source authority is a declared registry, not a classifier, so labelled hallucination datasets test something else. Entitlement on PrivacyLens is a mechanical check of purpose limitation, expected by construction. Composed privilege escalation is implemented and tested, but no labelled dataset exists to score it. The benchmark index says so rather than forcing a number.
Check the numbers yourself
From a clone of the repository:
python scripts/check/claims.py containment.attacks_contained_with_detectors_off 8/8 quoted in 5 place(s)
containment.legitimate_calls_allowed_with_detectors_off 4/4 quoted in 3 place(s)
agentdojo.session_taint.attack_pairs_contained 588/588 quoted in 9 place(s)
agentdojo.session_taint.benign_tasks_allowed n=24, d=97, pct=24.7, ci=[17.2, 34.2] quoted in 12 place(s)
agentdojo.argument_taint bk=37, bn=97, bpct=38.1, bci=[29.1, 48.1], ak=527, an=588, ck=641, cn=702, missed=61 quoted in 7 place(s)
…python scripts/check/claims.py --check is what CI runs. It prints all 30 published claims match their sources, or the quote that drifted. Each directory under benchmarks/ has its own fetch and run scripts, committed data, and a README with the method and its caveats.