No Retraining, 0.42% Attack Rate: Self-Evolving Defenses That Learn From Failed Attacks

Agent defenses have a shelf-life problem. A detector tuned against last month's injections meets this month's phrasing and folds — Meta's Prompt Guard 2 went from credible to 1% catch rates on buried attacks with a single threshold flip. A paper published 29 September 2026 by Minh Nhat Le and colleagues (arXiv 2609.36603) proposes a different shape of defense: instead of retraining the model or freezing one detector, distill each harmful agent trajectory into a reusable natural-language security policy, store it, and retrieve the relevant policies at task time. No weight updates, no redeploy — the defense accumulates knowledge the way an incident-response runbook does.

The mechanism, called Self-Evolving Defense (SED), is deliberately boring infrastructure around an exciting idea: when an agent trajectory goes bad, extract the generalizable lesson ("never pipe tool output containing instruction-like text into a shell without quoting"), file it, and surface it on future tasks that rhyme. Retrieval keeps the context small; accumulation keeps the coverage growing across jailbreaks, prompt injection, and insecure code generation alike.

The numbers, and why the baselines matter

The evaluation spans eight benchmarks across three open-source models — DeepSeek V4 Flash, GLM 5.2, and Kimi K3 — covering jailbreaks, prompt injection, and vulnerable code generation. Two results carry the claim:

  • Targeted prompt injection on AgentDojo: 0.42% success, against 3.7% for the best baseline defense — roughly a nine-fold reduction on the benchmark this site watches most closely.
  • Adaptive X-TEAMING attacks on HarmBench: 7.8% success, against 35.2% for the best baseline — more than four times lower, under an attacker that adapts rather than replaying fixed strings.

The adaptive-attacker result is the one to weight heaviest. Static benchmarks overstate every defense; an attacker that revises its approach is the realistic threat, and holding it to single digits while baselines collapse past a third is a genuine gap. The authors also report that benign task utility is preserved — the retrieved policies constrain the attack surface without turning the agent into a refusal machine, which is the failure mode the Cyber Index just showed frontier models defaulting to.

Read it as incident response, not as a model

The honest way to evaluate SED is against the alternatives operators actually have. Retraining on every new attack family is too slow and too expensive; frozen detectors decay, as Prompt Guard 2 demonstrated within days of analysis. A policy store has the economics of detection engineering: each new attack costs one distillation, not one training run, and policies are inspectable in a way weights are not — you can read what the defense believes, audit it, and delete stale entries. That inspectability also answers, partially, the late-layer compliance problem: rather than rewiring where in the network obedience is decided, SED changes what the model reads before it decides.

Caveats worth keeping. All three test models are open weights in the same capability band — transfer to frontier closed models is unproven. Retrieval is itself an attack surface: a poisoned or adversarially-planted policy in the store would be defense infrastructure working for the attacker, so the store needs the same write-integrity discipline as the audit trail does. And eight benchmarks are breadth, not deployment — the "continual" claim needs longitudinal evidence that policies don't bloat, contradict, or rot as attacks evolve. Still: training-free, inspectable, and a 4x-plus margin over baselines against adaptive attackers is the strongest defense-paradigm result of the month.

What to do

  • Trial the pattern without adopting the paper. You can prototype SED-shaped defenses today: log harmful trajectories, distill lessons into a policy file, retrieve top-k policies into the agent's context per task. Measure AgentDojo-style attack success before and after on your own stack.
  • Protect the policy store like detection config. Version it, review additions, restrict writers. A retrievable-policy defense is only as trustworthy as its least-reviewed entry.
  • Benchmark against adaptive attackers, not strings. If your red team replays fixed injection payloads, you are measuring the incident you already survived. Budget for multi-turn adaptive testing — that is where SED's margin actually shows.
  • Track utility alongside security. Refusal-rate and task-completion deltas belong in the same dashboard as attack success. A defense that holds attacks to zero by refusing the work is a shutdown with better marketing.

Sources: