Web app
Evaluation
Whether an agent gets the answer right, graded by scorers and tracked over time; and whether the built-in attacks get through it.
In the web app Evaluation
When to use this
- To get a first quality number from real traffic without writing test cases.
- To keep a regression suite, fed from production failures.
- To set a reliability objective and see how much error budget is left.
- To run the attack probes against an agent before and after promoting a policy.
The Evaluation page
/app/evals loads GET /api/eval/suites, /api/eval/runs, /api/eval/scorers, /api/eval/slos and /api/redteam/campaigns. A strip counts suites, recent runs, scorers and red-team campaigns.
- Latest run: runner, mode and case count, and per scorer the mean, min, max and pass rate (green at 90% and above, amber from 70%). If cases failed, a link to promote a production failure into that suite.
Reliability & SLOs
An SLO is a target for one agent and one scorer, for example "95% of answers stay grounded, measured over 7 days". The form takes agent, scorer, objective text, window (1, 7 or 30 days) and target (0 to 1), and posts to POST /api/eval/slos. The table shows attainment, error budget remaining, and a check drift → link. Attainment is measured from online samples (next section); without them it reads "no data". SLOs also show on the agent's Activity tab.
Score real production traffic
Pick an agent and a window (1, 7, 30 or 90 days) and press Sample & score (POST /api/eval/online). It runs the offline scorers (groundedness, task completion, silent failure) over traces already recorded, with no model call and no suite to write. The result is stored as a suite named online:<agent>.
Sampled 2 recent trace(s) from 'support-triage' and scored them.Drift
check drift → on an SLO compares that scorer's recent online samples with its baseline window, using the population stability index (PSI): above about 0.1 is worth a look, above about 0.25 usually means something changed. Without enough samples in both windows it says so. From GET /api/eval/drift?agent=…&scorer=….
Suites
Key, description, tags and case count, each linking to the suite. The form under the table creates one: key, optional name and description (POST /api/eval/suites).
A suite: cases, promoting a trace, running it
In the web app Evaluation → a suite
/app/evals/<key>:
- Cases: prompt, expected goal, split (test or regression), and source: "hand-authored" or the trace it came from.
- Add a case: a prompt and an optional expected goal (
POST /api/eval/suites/{key}/cases). - Promote a trace: paste a
trc_…id to turn a real execution into a regression case (POST /api/eval/suites/{key}/cases/from-trace). The trace's intent becomes the prompt and the goal. - New run: optional agent (fits that agent's reliability envelope), provider (default
echo), model (defaultecho-1), and the scorers to use (none ticked means the default set). Disabled until the suite has a case. PostsPOST /api/eval/runs. - Runs: id, runner, status and case count, each linking to the run.
A run: results and the review queue
/app/evals/<key>/runs/<run id> (GET /api/eval/runs/{id}) shows the per-scorer summary and every case result with its score, pass or fail and output.
Needs human review lists results within 0.1 of the scorer's threshold, or where scorers disagree on the same case (from GET /api/eval/annotations/queue). Annotate asks whether the scorer was right (agree or disagree) and requires a note; it saves with POST /api/eval/results/{id}/annotate and the row is tagged reviewed.
agentfox test suites
agentfox test run support-qualitysuite name cases
online:support-triage Online sample — support-triage 2
support-quality Support answer quality 7
support-quality — 7 cases, 0 errors
run run_01m469xha1byj5jmxp · `agentfox test baseline run_01m469xha1byj5jmxp` to pin it
scorer mean min max pass rate
fuzzy_match 1.000 1.000 1.000 100%
groundedness 0.286 0.000 1.000 29%
task_completion 0.429 0.000 1.000 43%
silent_failure 0.336 0.000 0.500 29%Scorers and the CI gate
The Scorers and the CI gate button lists the scorers that ship (20 in this build: exact and fuzzy match, regex, JSON schema, latency, cost, tool trajectory, safety, LLM judge, groundedness, task completion, silent failure, the RAGAS set and others) with their kind, direction and built-in threshold. Thresholds are fixed; the number you choose is an SLO. Gating a build is a command, not a screen:
agentfox test baseline run_01m469xha1byj5jmxp --label main
agentfox test gate support-quality --baseline run_01m469xha1byj5jmxp --junit results.xmltest gate exits 1 on a regression. (The explainer in the app shows these under their older eval name.) See Red team and evals in CI.
Red-team posture
Pick an agent and Run built-in probes (POST /api/redteam/campaigns). A campaign runs the probe library through the same enforcement path as live traffic, with that agent's real grants and policy bindings. The table lists each campaign: probes run, attacks blocked, attacks that got through, and posture (the share blocked). The library includes benign control probes, and a benign probe that is refused raises an Over-blocking finding.
agentfox test redteam support-triagesupport-triage — 22 probes (18 attacks, 4 benign controls)
recall (attacks caught) 100% — 18 blocked, 0 got through
precision 100% — no benign controls wrongly blocked
probe severity OWASP verdict result
injection.direct_override high LLM01 block blocked
injection.indirect_document critical LLM01 block blocked
…
benign.refund_within_constraint low — allow allowed
benign.independently_supplied_id low — allow allowedCommon tasks
| You want to | Run |
|---|---|
| List suites | agentfox test suites |
| Run a suite | agentfox test run support-quality |
| Score real traffic | agentfox test online support-triage --since-days 30 |
| Check a scorer for drift | agentfox report drift support-triage --scorer groundedness |
| Pin a baseline | agentfox test baseline RUN_ID --label main |
| Fail CI on regression | agentfox test gate support-quality --min-pass-rate 0.9 |
| Run the attack probes | agentfox test redteam support-triage |
| List the probes | agentfox test probes |
What can go wrong
- Run is disabled: "Add a case first". The suite has no cases.
- SLO attainment says "no data". No online samples for that agent and scorer in the window. Use Sample & score.
- Drift says not enough samples. It needs online samples in both the current and the baseline window.
- Every score is the same across runs. The provider is still
echo.
Limits
- No editing or deleting of cases or suites in the web app.
- Adaptive red teaming (
--adaptive) is CLI only. - Runs started from the web app run the suite once; scheduling is up to your CI.