A 23.85-Point Gap the Authors Refuse to Hide: AdvSim2Real Trains Web Agents Against an Adversary That Learns Back

Every indirect-prompt-injection defence that fine-tunes on a fixed set of attacks has the same expiry date: the attacker adapts and the defence does not. arXiv:2610.08773, posted 6 October 2026 by Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov, Praneeth Vepakomma and Nils Lukas (Mohamed bin Zayed University of Artificial Intelligence, with affiliations at Amazon and MIT), attacks that expiry date directly. AdvSim2Real co-evolves three policies against one another — a curriculum that writes tasks, an adversary that plants injections, and the web agent itself — inside a frozen web world model, so neither the tasks nor the attacks go stale as the agent improves.

What makes the paper worth your time is not only the method. It is that the authors state the limits of their own result more bluntly than most readers would.

The problem both halves of the literature get half-right

The paper frames the bind precisely. A web agent cannot simply ignore page content, “because the page also holds the values and controls the task requires.” Instruction and data arrive through the same channel by construction. Existing work splits into two incomplete answers: defences that fine-tune on injections fixed before training, which an adaptive attacker walks around; and adversarial training that lets the attacker adapt but keeps the task set fixed, so “a task stops teaching once the agent solves it.”

AdvSim2Real removes both ceilings by running everything inside a world model — a frozen pretrained transition model (WebWorld-14B) that predicts the next page for any goal, page and action, including a page an adversary has asked it to alter. No live website is required to invent a new task, and no live website is required to render a new attack.

Two reward designs carry the work:

  • The curriculum is paid for difficulty, not success. It is rewarded for proposing tasks the current agent solves about half the time, as graded by a fixed language-model judge. Reward is zero at judged rates of 0 and 1, so the generator cannot farm trivial or impossible tasks.
  • The adversary is paid only for a success flip. A clean run the judge already accepted is replayed to a chosen step, the injection is rendered there, and the adversary earns credit only if the continuation then fails. This is the sharp idea in the paper: it rules out the degenerate adversary that scores by breaking runs which were going to fail anyway.

Training runs in two stages. First the curriculum and the executor alternate updates so task difficulty tracks competence; then the curriculum freezes and the adversary and executor alternate, with the executor training on clean tasks, new attacks and the attacks of earlier rounds.

The numbers

The executor is a Qwen3.5-4B model evaluated on a 150-task benchmark inside the frozen WebWorld-14B world model (fifty of the 150 are close variants of eight parent tasks — a composition detail worth holding onto). Judging is by a fixed Qwen3.8-27B judge.

  • Clean completion: 74.89% → 81.33% (+6.44 points).
  • Under the three learned adversaries: 48.07% → 57.48% (+9.41 points).
  • Under Kimi-K3, a frontier model that took no part in training: 23.00% → 30.72% — the 33.6% relative gain quoted in the abstract.
  • In a real browser, no world-model call: strict success on the submitted form 25.56% → 44.44%, with correct-field score 52.50% → 73.24%.

Two of these deserve emphasis for opposite reasons. The transfer to an untrained frontier adversary is the result that matters most for a defender, because it is the only evidence that the agent learned something about injections rather than about these injections. And the real-browser number is the one that answers the obvious objection to simulator training: the capability gain survived contact with a browser that, as the paper notes, “rejects malformed actions that the world model tolerated” — the base model issued at least one invalid action in 132 of 450 episodes.

The honesty is the contribution

Papers in this space routinely report an attack-success-rate drop and stop. This one does not. Its own conclusion reads:

“A 23.85-point clean-to-attacked gap remains, and every robustness number is a model judgment inside one web world model.”

Both halves of that sentence are load-bearing. The first says the defence narrows the gap and does not close it: an agent at 81.33% clean and 57.48% attacked is still losing roughly a quarter of its completions to injection after adversarial training designed specifically against injection. The second says the robustness figures are a judge model’s opinion about trajectories rendered by a single simulator. Only the capability numbers were re-measured in a real browser; the attacked results were not. The paper says so in-line: “all attacked results are measured inside the world model.”

The ablation is reported with the same discipline. Removing the first stage leaves a pipeline whose clean completion trails the full method at every round (by 5.78, 2.00 and 4.67 points), but whose attacked performance does not consistently trail — the full pipeline leads by 0.89 points after round one, trails by 0.52 after round two, and leads by 0.44 after round three. The authors note the ablation does not match compute and removes two components jointly, and they decline to attribute the gains to any single mechanism: “we designed for three mechanisms, none of which we have isolated.” They also record that the first adversarial round lowers both metrics before later rounds recover them.

That is what a replicable claim looks like. It is also the standard we have been applying to this literature all year — the same lens we brought to threat-preserving representation sensitivity in agent benchmarks, where surface form rather than threat recognition drove the scores, and to branch steering against dual-LLM planners, where a defence evaluated only on its authors’ own benchmark reported a 0% attack success rate.

The ethics section names the obvious risk

Training an adversary that learns to divert agents produces, unavoidably, an adversary that can divert agents. The authors address it rather than eliding it: the released adversaries are low-rank adapters on a 4B model, trained and evaluated only on synthetic pages whose injections are rendered by a frozen world model, never tested against deployed agents or real websites, with no known deployments of the trained policies. They also disclose that language models are components of the method itself — curriculum, adversary, executor, the WebWorld-14B world model, the Qwen3.8-27B judge, and the Kimi-K3 evaluation adversary — and that generative assistants were used in drafting, with the authors taking responsibility for the content. The reported numbers are stated to be recomputable from per-seed counts in the appendix.

The release of code, benchmark and checkpoints is the right call for a defence paper. The explicit purpose — “so that injection defenses can be evaluated against adversaries trained on the defended agent” — is the correct evaluation standard, and almost nobody meets it today.

What a security team should take from this

  • Treat static injection corpora as expired. If your agent evaluation uses a fixed attack set, you are measuring robustness against attacks that predate your model. The paper’s central argument is that this number does not survive an attacker who adapts, and its Kimi-K3 result is the only transfer evidence it claims.
  • Do not read 57.48% as a safe deployment. Adversarially trained, judged by a cooperative judge, in a simulator, this agent still fails a quarter of its otherwise-successful tasks under attack. Adversarial training is a mitigation layer, not a boundary. Keep the architectural controls — capability constraints, destination allowlists, human confirmation on irreversible actions — underneath it.
  • Ask where robustness numbers were measured. This paper separates simulated robustness from browser-verified capability and labels each. Most do not. “Measured in a real browser” and “measured in the environment that generated the training data” are different claims.
  • Watch the benchmark composition. Fifty of 150 tasks are close variants of eight parents. That is disclosed, and it bounds how much task diversity the headline percentages represent.
  • Independent replication is still outstanding. Everything above is the authors’ own evaluation of their own method on their own benchmark, released four days ago. The artifacts are public; the third-party reproduction is not yet.

Verification note: we read the paper from the arXiv listing and full HTML text for arXiv:2610.08773v1 (announced 6 October 2026, cs.CL with cs.AI and cs.LG cross-lists, CC BY 4.0), and confirmed title, author list, submission date and abstract against the arXiv API metadata record. All percentages, point differences, model names, benchmark sizes and quoted sentences are taken verbatim from that paper — including the 23.85-point clean-to-attacked gap, the per-round ablation differences, the 132-of-450 invalid-action count, and the ethics and AI-use statements. Tables referenced by the authors (1, 2, 3 and 7) are cited as they number them. This is a preprint: we found no peer-review record, no venue acceptance and no third-party replication at the time of writing, and we do not imply any. We did not run the released code or reproduce any measurement.

Sources: