No Injected Instruction Required: arXiv:2610.03089 Steers Agents Down Branches You Already Approved

Almost every deployed prompt-injection defence is looking for an instruction. Classifiers hunt for imperative text in untrusted content; guardrails flag “ignore previous instructions” and its thousand paraphrases; the research literature benchmarks itself on attacks that, at bottom, consist of telling the model to do something else. arXiv:2610.03089 — “Securing Computer-Use Agents Against Branch Steering Attacks,” submitted 2 October 2026 by Giulio Zingrillo, Hanna Foerster, Ilia Shumailov, Yiren Zhao and Robert Mullins, and slated for the Agents in the Wild workshop at NeurIPS 2026 — describes an attack that needs none of that, and that is why it matters.

The guarantee that does not survive a GUI

The Dual-LLM pattern is the strongest system-level answer the field has to indirect prompt injection, and the only one offering a formal guarantee rather than a probability. The construction is simple: an isolated Planner LLM fixes the execution path before any untrusted data is read, and a Quarantined LLM processes the untrusted content without the authority to change the plan. If the plan is data-independent, injected text cannot redirect control flow. That is the whole argument, and in a text-tool setting it holds.

The paper's contribution is to show where it stops holding. Computer-use agents drive graphical interfaces, and a GUI is not something you can plan against blind. The agent cannot know in advance which dialog will appear, which list will be empty, which button will be disabled, which search will return three results or none. So the planner cannot emit a straight line. It must emit a branching plan that anticipates runtime content — covering, as the authors put it, all possible cases the agent may encounter.

The branches are the vulnerability. Every branch in that plan is pre-approved by the trusted planner. The attacker does not need to add a new path; they need only make the agent believe it is in the situation that selects an existing hazardous one. Branch steering is the manipulation of untrusted data to coerce a computer-use agent down a hazardous, pre-approved branch without injecting explicit instructions.

Read that condition carefully, because it is the operational finding. No imperative verb. No instruction-shaped string. Just page content arranged so that the branch the attacker wants is the branch the agent's own trusted plan says to take. The formal guarantee is not violated — the agent does exactly what the planner authorised. The guarantee was simply never about this.

The numbers

The authors build STEER-Bench, 101 tasks across 9 domains, and report attack success of 94.4% against standard computer-use agents and 89.5% against vanilla Dual-LLM agents.

The gap between those two figures is the headline. The architecture specifically designed to make indirect prompt injection structurally impossible absorbs under five percentage points of this attack. For a defence whose selling point is a guarantee rather than a filter, a five-point reduction is indistinguishable from no defence at all — and worse than none in practice, because it is the kind of result that gets deployed and then trusted.

This is a theme we have had cause to return to repeatedly: defences that score well on the benchmark they were designed against and poorly on anything shaped differently. arXiv:2610.03448 found detector rankings that do not transfer between benchmarks at all, and arXiv:2610.03585 found that renaming a tool moves the measured attack success rate more than the threat does. Branch steering is the same lesson arriving from the architecture side rather than the evaluation side: the Dual-LLM guarantee is real, precisely scoped, and scoped to a threat model that graphical agents have outgrown.

What the fix actually constrains

The authors propose COBRA, which pairs trusted branching plans with ahead-of-time capability constraints — strictly bounding the parameters and destinations each branch is permitted to execute. On STEER-Bench it reduces attack success to 0% while retaining 97% benign utility.

The design principle is worth extracting even if you never run their code, because it generalises past this paper. Branching is not the problem and cannot be removed; a GUI agent that cannot branch cannot function. What is removable is the assumption that a branch, once selected, inherits the agent's full authority. If each branch in a plan carries its own bound on which parameters and which destinations it may touch, then steering the agent into a branch stops being equivalent to owning the agent. The attacker still chooses the path; they no longer choose what the path can reach.

That is the same structural move as capability-bounded tool access in non-GUI agents, and it is noticeably more durable than detection, because it does not depend on recognising the attack. It is also where the honest caveats live. A 0% and a 97% are single-architecture results on a 101-task benchmark the same authors built, and the paper is a workshop submission in its first version. The attack result — 94.4% and 89.5% — is the finding that should change behaviour today; the defence result is a promising direction that needs independent replication on benchmarks its designers did not write.

What to take from it

  • Stop treating Dual-LLM as closing the injection question for GUI agents. The guarantee is sound in the setting it was proved for. Computer-use agents are not that setting, and the measured residual is under five points.
  • Audit your plans for hazardous pre-approved branches. If an agent's plan contains a path that deletes, sends, pays or escalates, that path is reachable by anyone who can influence what the agent sees — no injected instruction needed. Enumerate those branches before enumerating your filters.
  • Bound capability per branch, not per agent. Agent-level allowlists do nothing against an attack that works entirely inside the allowlist. Per-branch parameter and destination limits are the control that matches the threat.
  • Expect instruction-shaped detection to miss this entirely. Any defence keyed on imperative text, suspicious phrasing or instruction-like structure has nothing to fire on. Content that is wholly benign in isolation is the attack.
  • Treat the defence numbers as a hypothesis, the attack numbers as a finding. One is a result about systems you may already be running; the other is a first-version proposal evaluated on its authors' own benchmark.

Verification note: the arXiv identifier, title, author list, 2 October 2026 submission date, NeurIPS 2026 “Agents in the Wild” workshop designation and 16-page length were read directly from the arXiv listing and the arXiv API record for 2610.03089. The description of the Dual-LLM Planner/Quarantined split, the branch-steering definition, the STEER-Bench size (101 tasks, 9 domains), the 94.4% and 89.5% attack success rates, and COBRA's 0% attack success and 97% benign utility are the paper's own abstract figures as published, quoted and paraphrased from it. We have not reviewed the full text, reproduced the benchmark, or independently evaluated either the attack or the defence; all reported numbers are the authors' claims on their own benchmark, which at the time of writing has no third-party replication. Our reading that the sub-five-point Dual-LLM margin is the paper's most consequential result is our assessment, not a claim the authors make in those terms.

Sources: