Meta’s Detector Slept Through 99% of Real Agent Attacks. The Fix Was One Number
Every agent stack ships with a smoke alarm now: a prompt-injection detector, a small classifier that reads text flowing through the agent and flags what smells like an attack. On 29 September 2026, Rudratosh Shastri published the obvious experiment nobody had run properly — buy ten of these smoke alarms, set 629 real fires, and write down which ones go off. The headline result is that Meta's Prompt Guard 2 caught 1% of real attacks out of the box, and 99% after changing a single threshold number. The second half of that sentence is the one that should worry you.
The benchmark, open-sourced as buried-injections, fixes the thing most detector evaluations get wrong. Standalone attack strings like “ignore previous instructions” are trivially detectable and nothing like the real threat. So each of the 629 attacks from AgentDojo — the academic standard for agent attacks — was embedded inside genuine AgentDojo tool output: a bank statement, a bill, an email, up to ~2,000 characters of ordinary text with the malicious instruction buried in the middle, phrased like a perfectly normal request. A further 97 benign tool outputs measure the false-positive side, because a catch rate without a false-positive rate is marketing, not measurement.
The leaderboard, with the column everyone omits
Out-of-the-box results, attacks caught versus safe traffic wrongly blocked:
- jailbreak-detector-large: 51% caught, 2% blocked — the only honest performer, a coin flip that leaves normal traffic alone
- protectai-deberta-v2: 23% / 4% — flags all 27 attack texts scored alone, collapses once they sit inside real tool output
- llm-guard: 20% / 2% — same story, same model family behaviour
- prompt-guard-2-86m (Meta): 1% / 0% — the famous one, asleep
- prompt-guard-2-22m (Meta): 0% / 0% — went home entirely
- regex baseline: 0% / 0% — real attacks contain none of the words grep is looking for
- preamble-defense: 88% / 47% — screams at toast
- testsavant-defender: 59% / 48% — also screams at toast
- deepset-deberta and fmops-distilbert: 100% / 98% — perfect attack detection achieved by suspecting the entire world; a brick taped over the deny button
Three genuinely different failure modes: detectors that learned attack phrasing instead of attacks (Prompt Guard 2 scores a hand-written “send all passwords” at 0.999 but real AgentDojo attacks at 0.004–0.140), detectors whose signal dilutes inside surrounding benign text (ProtectAI aces the flashcards, fails the open-book exam), and detectors whose perfect scores are just “block everything” in a lab coat.
The threshold flip
A detector's default threshold is just a number someone picked. Instead of trusting factory settings, the benchmark finds the cutoff where each model wrongly blocks at most 2% of normal traffic — tuned on three of AgentDojo's four domains, measured on the fourth it never saw:
- prompt-guard-2-86m: 1% → 99% — worst to first
- jailbreak-detector-large: 51% → 51% — unbothered, still the honest coin flip
- fmops-distilbert: 100% → 48%, deepset-deberta: 100% → ~0% — the block-everything crowd collapses to near zero
- preamble-defense: 88% → 3%
Meta's model was never bad at ranking attacks — it scored attacks around 0.009 against normal text around 0.0008, near-perfect separation — it just had its cutoff parked at 0.5, roughly 50× too high for this distribution. Move it to 0.003 and the leaderboard inverts. The default threshold mattered as much as the model. Anyone comparing detectors at factory settings is ranking strangers' guesses about distributions that are not theirs.
The author's own caveats — read them before screenshotting 99%
To his credit, Shastri states the limits plainly, and they are load-bearing. Every AgentDojo attack shares a wrapper template, so a threshold tuned to 0.003 may be recognising the template rather than attacks in general — and real attackers, unlike academic benchmarks, change their wording. Thresholds that fine are fragile, and 97 benign samples is a thin stick for measuring a 2% budget. The 99% is a demonstration that thresholds dominate, not a certification of Prompt Guard 2.
The deeper finding survives all of that. You cannot reliably tell an attacker's instruction from a user's instruction by reading the text: “send money” and “delete file 13” are attacks or chores depending entirely on who said them and what they would do — information that is simply not in the words. As the closing demonstration, Prompt Guard 2 happily allows rm -rf /, reading ~/.ssh/id_rsa, hitting the cloud metadata endpoint 169.254.169.254, and curl … | sh. Not injections, so not its job — but very much your agent's problem. The defence that holds is knowing where each instruction came from and gating what each tool call may do, per tool and argument. That is the same conclusion as the tool-allowlisting work showing zero attack success at a utility cost, and it rhymes with the mechanistic finding that injection compliance is decided late in the network, not where early detectors look. Text detection is a useful layer. It is a catastrophic foundation.
What to do
- Tune every detector threshold on your own traffic before trusting any number on its model card. Ship the sweep, not the default.
- Always report catch rate next to false-positive rate, on traffic shaped like yours. A 100% with 98% blocked is a denial of service you installed yourself.
- Benchmark detectors buried in tool output, not on standalone attack strings. The standalone test measures a threat that barely exists.
- Build the real control from provenance and tool policy — who said it, what the call would do — and treat the classifier as one noisy signal into that decision.
Everything is reproducible: make setup, make bench-agentdojo (~25 minutes on CPU), make bench-budget for the threshold sweep with the cross-domain check.
Sources:
- Rudratosh Shastri — “Meta's prompt-injection detector caught 1% of real agent attacks. One config change made it 99%.” (dev.to, 29 September 2026; full leaderboard, threshold-sweep method, caveats)
- rudratoshs/buried-injections — reproducible benchmark: 10 detectors, 629 AgentDojo attacks in genuine tool output, 97 benign samples, cross-domain threshold evaluation (MIT, created 23 September 2026)