Screening the Input Is Too Early, Grading the Run Is Too Late — arXiv:2610.05163 Audits at the Action Boundary

Most prompt-injection research assumes a shape of attack that is becoming obsolete: one poisoned document, one hijacked response, one measurable failure. arXiv:2610.05163 — "Blocking at the Boundary: Auditing Long-Horizon Agents against Staged Prompt Injection," submitted 4 October 2026 by Jingkai Liu, Yufei Han, Xiaoting Lyu, Wei Wang and Ting Yu — starts from the other end, with the agents people actually run, and the finding is less comfortable.

The authors name the threat staged prompt injection: an injection that exploits task-specific context, propagates across causally connected stages of a workflow, alters one consequential action, and lets the workflow carry on. Nothing crashes. The run completes. Somewhere in the middle, one tool call did something it should not have.

They ran it against the real products

This is not a simulated agent on a toy harness. The authors built an automated, feedback-guided attack-generation pipeline and pointed it at Claude Code and Codex in their native runtimes. The confirmed attacks span eight workflow scenarios, seven attack goals and six injection surfaces. The authors' conclusion, stated plainly in the abstract: production agents are vulnerable to context-aware, multi-step injection over long horizons.

That sentence is worth separating from the defence contribution, because it stands on its own and it is the part most readers should act on. Two coding agents that a very large number of engineering teams now run against untrusted repositories, issue trackers and web content were successfully attacked by an automated pipeline, repeatedly, across a wide matrix of surfaces. We have covered the single-shot version of this problem — opening a hostile repository as direct RCE. Staged injection is the quieter sibling: no escape, no shell, just one altered action inside a run you will later read as successful.

The argument about where to intervene

The paper's structural claim is sharper than its attack results, and it is aimed squarely at how most deployed defences are wired.

Input screening cannot work, because at ingestion time the injected content has not yet been composed with the context that makes it consequential. The text that will eventually steer a destructive tool call does not look like an attack when it arrives.

Completed-run evaluation cannot work, because by the time the trajectory is scorable, the effect has landed. You get an accurate post-mortem of a push you already made.

Neither, in the authors' phrasing, can locate the intervention point. And existing pre-action methods are not directly comparable to each other because they use incompatible units and labels — some judge messages, some judge tool calls, some judge steps. So the authors formulate the problem properly: boundary action auditing. Given the initial context, a trajectory prefix, and a fully specified pending message or tool call, an auditor returns Pass or Block before the effect occurs. One decision, one well-defined unit, immediately before each consequential action.

To evaluate it they pair attacked and benign executions, which yields a 479-pair, 3,112-unit benchmark. Pairing matters: it means the false-block rate is measured on the same workflows that produce the true blocks, rather than on a separate benign corpus that may differ in difficulty.

Path-Aligned Attribution

Their proposed auditor, PAA, is training-free — which, given what the last fortnight of evaluation research has established, is a design choice with real evidential value and not just an engineering convenience. PAA decomposes a pending action into its operative elements and traces, for each one, what supplied the value and what guided the decision. It blocks only when the model attributes an unwarranted, material effect on an element to an attacker-reachable source that either provides unqualified steering or conflicts with visible evidence.

That is a narrow and auditable condition. It is not "does this look suspicious." It is a provenance question about a specific argument of a specific call, with three conjunctive requirements — unwarranted, material, attacker-reachable — before a block fires.

The numbers

Under full-benchmark fail-open scoring with Claude Sonnet 5 as the backend:

  • PAA: 86% Block recall at a 6–8% false-block rate.
  • ARGUS: 44–47% recall at a 16–33% false-block rate.

Roughly double the catch rate at a fraction of the interruption cost. On the subset of tool calls that all three auditors natively support, and with the same backend, the authors report PAA has higher recall and lower false-block rate than both VIGIL and ARGUS, with all paired 95% confidence intervals excluding zero.

Two things deserve credit here. The first is fail-open scoring: an auditor that errors, times out or returns nothing is counted as having passed the action. That is the honest choice, and it is the opposite of the convention that quietly inflates numbers by treating a crash as a block. The second is reporting paired confidence intervals at all, which remains rare in this literature.

Reading it against the last two weeks

We have now covered three papers in eight days that each undermine a different assumption behind published agent-security scores. arXiv:2610.03448 showed injection-detector rankings collapsing once training-split overlap was controlled — 95.1 on one suite, 2.1 on another. arXiv:2610.03585 moved attack success by 11–13 points by renaming a tool, with a token-matched neutral control reproducing most of the shift. This paper adds a third: the unit of evaluation has been wrong, and pre-action defences have not been compared on a common one.

Those three findings interlock in a way that should make anyone cautious about the 86% figure too. A training-free auditor is structurally immune to the 03448 objection — there is no training split to overlap with. It is not immune to the 03585 objection, because PAA is a model-driven judgement over an action's surface representation, and a benchmark built by one team's attack pipeline carries that team's phrasing conventions. The right reading is that the comparison against VIGIL and ARGUS on this benchmark is credible and well-controlled, and that none of the absolute numbers should be carried into a different attack distribution without re-measuring.

It is also the third training-free or self-supervised approach this month to beat trained defences on their own ground, after Self-Evolving Defense's retrieval of distilled policies. A pattern is forming, and it is not subtle: training-time defences keep paying for robustness with capability, while inference-time structure keeps not doing that. Whether that holds under adaptive attack is the question none of these papers can answer yet.

What to take from it

  • If your agent defence sits at ingestion, it is in the wrong place for this threat class. That is the paper's central claim and the argument for it is strong: context-dependent injections are not identifiable before the context exists.
  • Instrument the action boundary, even without an auditor. The minimum viable version of this is logging the fully specified pending call with its argument provenance before execution. You cannot audit what you do not capture, and most agent logs today record the call after the fact.
  • Demand fail-open numbers from vendors. Ask explicitly how an auditor error is scored. If a timeout counts as a block, the published recall is not measuring what you need.
  • Watch the false-block rate as a deployment number, not a benchmark column. ARGUS's 16–33% is the interesting figure here: an auditor that interrupts a third of legitimate actions will be switched off within a week, which makes its recall irrelevant in practice.
  • Treat the Claude Code and Codex results as current, not historical. These are live products on fast release cycles; the attacks were confirmed against their native runtimes, and the paper does not claim vendor remediation.

Verification note: all figures — eight workflow scenarios, seven attack goals, six injection surfaces, the 479-pair / 3,112-unit benchmark, PAA at 86% block recall and 6–8% false-block rate, ARGUS at 44–47% recall and 16–33% false-block rate, the Claude Sonnet 5 backend, the fail-open scoring and the paired 95% confidence intervals — are taken from the abstract of arXiv:2610.05163 (v1, submitted 4 October 2026; 29 pages, 10 figures, 20 tables), retrieved by us directly from the arXiv API. We have not reproduced the experiments. This is a preprint with no peer review indicated at the time of writing, and the authors list an anonymised artifact repository rather than a named one. The comparison to arXiv:2610.03448 and arXiv:2610.03585 is our own analysis and is not a claim made by any of these authors. We have not independently confirmed the attack results against Claude Code or Codex, and we did not attempt to.

Sources: