The Detector That Wins the Benchmark Was Trained On It
On 2 October 2026, Zhuowen Liu of the Cybersecurity Lab at the Japan Advanced Institute of Science and Technology posted Passing the Test You Trained On (arXiv:2610.03448, cs.CR, CC BY 4.0). It asks a question that sounds procedural and turns out not to be: when a team picks a prompt-injection detector by its benchmark score, does that score predict anything about how the detector behaves inside an actual agent?
The answer, across fifteen detectors, two agent benchmarks and BIPIA, is no. Not “imperfectly” — the Kendall rank correlation between detection on BIPIA and detection on AgentDojo is 0.01 (p = 1.00). A leaderboard with that correlation to deployment behaviour is not a weak signal. It is noise with error bars.
The method is the contribution
The methodological move is clean enough to steal. Rather than running an LLM agent and trying to separate the model’s mistakes from the detector’s, the paper replays the ground-truth tool calls of AgentDojo and τ-bench with no LLM in the loop. That yields tool outputs that are benign by construction — the false-positive set is not a judgement call — and injected outputs are labelled by differential replay against the clean run. The result is a measurement of the detector on exactly the text a deployed detector screens: serialised records, YAML, JSON confirmations. Not prose. Not standalone attack strings.
That distinction is the whole paper, and it has a motivating example that should end the discussion on its own. In an AgentDojo banking task, the agent calls a password-update tool and receives {'message': 'Password updated.'}. PIGuard scores that output 1.00. A deployment that drops flagged outputs has just failed a task in which nothing whatsoever happened.
Three numbers
Detection is reported as TPR at a 1% false-positive rate, which is the honest operating point for something that screens every tool output — and a stricter bar than most vendor numbers. The findings that matter:
- False positives on benign tool outputs range from 0% to over 90%. Horizon-Labs flags 0.0% of clean AgentDojo outputs; Prompt Guard 2 (22M) 0.0%; Wolf Defender 0.9%. ProtectAI v2 and PIGuard flag 30.1% and 29.5% of outputs — which translates to stopping 69.1% and 59.8% of tasks outright. Crucially, this property does not show up on ordinary text: ProtectAI v2 flags 0.5% of web pages and 1.2% of emails, then 66.9% of τ-bench outputs. If you benchmarked it on prose you learned nothing about what it will do to your agent.
- Rankings do not transfer. PIGuard ranks first on BIPIA at 95.1% and thirteenth of fifteen on AgentDojo at 2.1%. Prismor catches 72.2% on AgentDojo and 15.2% on τ-bench; Sheltron goes the other way, 20.1% to 95.2%. Even between the two agent benchmarks, τ is only 0.28, not significant.
- False-positive rank, unlike detection rank, is stable across the two agent benchmarks (τ = 0.67, p = 0.001). So the one thing a benchmark reliably tells you about a detector is the thing vendors rarely put on the slide.
The training-data audit
The paper then does what evaluation papers almost never do: it audits where the scores come from. Only two of the fifteen detectors publish enough to make that possible.
PIGuard releases its training split. Normalising whitespace and case and searching for the first 80 characters of every BIPIA attack string, the authors find 1,116 labelled BIPIA inputs in it — 65 of 75 text attacks and 6 of 50 code attacks from BIPIA’s training half appear verbatim, along with 31% of BIPIA’s clean emails. The paper is careful here, and the care strengthens the point: the headline results use only BIPIA’s test half, none of which appears in training. The 95.1% is not raw contamination. It is something subtler and worse — the detector learned the form of BIPIA inputs, and that form is what it recognises.
The controlled comparison nails it. PIGuard’s training split also contains 53 of 62 InjecAgent attack instructions — but as short prompts paired with a generic task sentence, median 178 characters, not as attacks embedded in tool responses. On τ-bench, PIGuard detects 52.1% of outputs carrying an attack it was trained on and 55.8% of those carrying one of the nine it was not. Having literally seen the attack text is worth nothing — slightly less than nothing — when the surrounding structure differs.
The counter-example confirms the mechanism. Horizon-Labs excludes evaluation texts from training; the authors audited its five public training sources and found no BIPIA attack string, no AgentDojo template, no InjecAgent instruction. It leads both agent benchmarks (82.2% on AgentDojo, 100.0% on τ-bench) with zero data overlap. What it shares with them is input form: structured tool output with planted instructions.
Detectors work on inputs whose form resembles their training inputs. Not on attacks whose text they have seen.
What does not rescue it
Two obvious escape hatches are tested and both fail. Handing a general-purpose model the user’s request makes a reasonable detector — a Llama judge catches 55.3% on AgentDojo and 78.3% on τ-bench at 1% FPR — but stays below the best dedicated detectors while costing a 7–8B forward pass per tool output. Giving the same task context to a small classifier actively hurts: Sheltron’s AUROC rises from 0.86 to 0.91 while its detection at 1% FPR falls from 20.1% to 9.0% and its false-positive rate on clean τ-bench outputs explodes from 4.4% to 69.7%. An AUROC improvement that coincides with a usability collapse is a good argument for never quoting AUROC alone.
Formatting helps, but only where it does not matter. Stripping markup from READMEs drops PIGuard’s false-positive rate from 24.6% to 5.6%; removing YAML keys so the detector sees only values drops its AgentDojo FPR from 20.0% to 6.7% at unchanged TPR. Free wins, worth taking. But at the 1%-FPR operating point, PIGuard with keys removed still catches 3.0% of AgentDojo injections. Presentation moves scores; it does not move an out-of-distribution detector into distribution.
And one category defeats nearly everything: task-irrelevant injections, which instruct the model to do something unrelated and harmless. Apart from BIPIA-trained PIGuard at 93%, the best approach is Prompt Guard 1 at 39%, and every other method catches at most 20%.
Where this sits
This is the second rigorous result in a week pointing the same direction, and the two are complementary rather than redundant. The buried-injections benchmark showed that Prompt Guard 2’s default threshold was the problem and a single number fixed it. This paper fixes the threshold by construction — everything is measured at 1% FPR — and finds a deeper failure underneath: even correctly tuned, a detector only works on input shaped like what it was trained on. Threshold was the surface. Distribution is the floor. Notably, Prompt Guard 2 does respectably here (69.4% and 60.9% at 1% FPR on AgentDojo) precisely because the threshold question has been removed.
Read alongside the finding that the model identifies the injection early and complies anyway, and MLCommons naming the connector layer’s missing controls, the shape of the problem is consistent: content classification is a weak control at the agent boundary, and the industry keeps selecting these classifiers with instruments that do not measure the deployment.
The paper’s own limitations are stated and should be honoured. Both agent benchmarks are simulations — synthetic YAML and JSON from synthetic databases — while real tool outputs are messier HTML and API responses, and the paper itself demonstrates that format moves false-positive rates, so absolute numbers will shift in production. Commercial API detectors are closed and were not tested. The training audit finds verbatim copies only, so the overlap figures are a lower bound, not an estimate.
What to do
- Re-measure on your own tool outputs before trusting any detector score. Replay your agent’s real tool calls, screen the outputs, and count how many tasks a flag would have killed. A 30% output-level false-positive rate became a 69% task failure rate here; that multiplier is the number your product team cares about.
- Demand TPR at a fixed low FPR, and refuse accuracy or AUROC on balanced sets. Both hide exactly the error that breaks deployments, and this paper shows AUROC moving opposite to usable detection.
- Audit the training data before the leaderboard. Ask whether the detector was trained on the benchmark it leads, or on data generated the same way. Where a vendor publishes nothing, treat the score as unverifiable — thirteen of fifteen detectors here could not be audited at all.
- Prefer detectors trained on agent-style inputs over detectors with higher public scores. That is the single predictive variable the paper isolates, and it beat data overlap outright.
- Strip serialisation markup before screening. Removing YAML keys and HTML scaffolding cuts default-threshold false positives by up to roughly twentyfold at unchanged detection. Cheap, immediate, and it changes nothing about the underlying limits.
- Do not budget a detector as the control. At the best measured operating points, roughly one in five AgentDojo injections still gets through, and task-irrelevant injections pass nearly everything. Scope the tools, constrain what the agent can do with what it reads, and treat the classifier as one filter among several.
Verification note: we read the full HTML version of arXiv:2610.03448v1 (submitted 2 October 2026, CC BY 4.0) on arXiv, and took the methodology, detector set, all reported percentages, Kendall correlations, training-overlap counts, findings 1–5 and the stated limitations from the paper’s own text, including Tables 2 and 3 and §5.1–5.5. Author affiliation is as listed on the paper. The work is a single-author preprint and, as of publication, we found no indication of peer review; figures are the authors’ and we have not independently reproduced any experiment. Comparisons to earlier coverage are our own editorial framing, not claims made by the paper.
Sources: