Rename the Tool, Move the Score: Agent Security Benchmarks Measure Wording as Much as Robustness

If you have ever chosen a model or a defence on the strength of an attack success rate, this paper is an invoice. arXiv:2610.03585, “Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks” — submitted 2 October 2026 by Neeraj Karamchandani, Piyush Nagasubramaniam, Xinhong Xie, Sencun Zhu and Dinghao Wu — asks a question that is embarrassingly simple and had apparently not been controlled for: how much of a benchmark’s ASR is a property of the agent’s security, and how much is a property of how the benchmark happened to word the scenario?

The method is the whole contribution. The authors define threat-preserving representation sensitivity (TPRS): the change in ASR when you alter only the agent-visible representation while holding everything else fixed — the underlying task, the harmful action, the security policy, the ground truth, the environment, and the evaluation criteria. Nothing about the security problem changes. Only how it reads to the model.

The numbers

On Agent Security Bench (ASB), replacing threat-related tool names with threat-neutral ones raises the committed attack success rate by 11.67 percentage points on GPT-5-mini and 13.21 points on Claude Haiku 4.5. The attack is identical; the tool is simply no longer called something that sounds alarming, and the models commit to it substantially more often.

The experiment runs in the other direction too, which is what makes it convincing rather than anecdotal. On MCPTox — the MCP tool-poisoning benchmark we covered in April — the original tool names are already neutral. Swapping in an explicitly threat-related name lowers ASR by 11.00 points on GPT-5-mini and 4.11 points on Claude Haiku 4.5. Same direction of effect, opposite sign, because the baseline wording sat on the other side of the line.

AgentDojo behaves differently and the authors report it honestly: adding threat-related wording to the attack-relevant tool shifts ASR by only 0.50 points on GPT-4o-mini. But the cost lands somewhere else — benign utility falls 5.36 points on tasks that require that tool. The wording did not make the agent safer; it made the agent worse at its job while leaving the attack roughly as successful as before. That is the failure mode every over-tuned guardrail eventually produces, isolated here as a measurement.

The control that rules out the obvious explanation

The intuitive reading of the ASB result is reassuring in a way: the models are reading the tool name, noticing it sounds dangerous, and refusing more often. Shallow, but at least it is semantic — something like threat awareness.

The authors tested that. On MCPTox they constructed a threat-neutral name matched on token count, length and casing, and found it reproduces most of the shift produced by the threat-explicit name — 8.54 of the 11.00 points on GPT-5-mini. Roughly four-fifths of an effect that looked like threat recognition survives when the threatening meaning is removed and only the surface form is held constant. Whatever is moving these scores is operating substantially below the level of understanding that the word “robustness” implies.

What this does to the leaderboard

The authors’ conclusion is measured: a security score obtained under one representation may not generalise across threat-preserving representations of the same problem, so robustness claims should be supported by performance across a controlled set of representations rather than a single representation-dependent number.

Stated less gently: double-digit ASR gaps are the normal currency of model and defence comparisons, and this paper produces double-digit gaps from renaming a function. Any ranking whose margin is smaller than its TPRS is not measuring what its axis label says. Nobody currently reports TPRS, so for most published comparisons the question cannot be settled from the paper.

This is the second structural problem in agent-security evaluation to land in a week. arXiv:2610.03448, covered here yesterday, showed prompt-injection detector rankings failing to transfer across benchmarks once training-split overlap was accounted for — detectors scoring 95.1 on one suite and 2.1 on another. The two findings are complementary and point at the same hole: 2610.03448 says the test set may be measuring familiarity, and 2610.03585 says the phrasing within the test set may be measuring surface form. Both leave the published number intact and the inference drawn from it unsupported. Our reading of OpenAPPA’s 0% attack success rate made the related point from the design side: a benchmark result and a security property are different kinds of claim.

What to take from it

  • Treat any single-representation ASR as a lower bound with unknown error bars. On these results, 10–13 points of movement is available for free from naming alone.
  • Discount model and defence rankings whose margin is in single digits. If the gap is smaller than the demonstrated representation sensitivity, the ordering is not established.
  • If you run internal agent-security evals, add the control yourself. Re-run a slice with tool names rewritten neutrally and threateningly, token-matched, and record the spread. It is a cheap diff and it tells you how much of your score you actually own.
  • Watch the utility column, not just ASR. The AgentDojo result — 0.50 points of ASR movement against 5.36 points of benign utility loss — is the shape of a guardrail that costs more than it buys.
  • Do not read this as “benchmarks are useless.” It is a call for a reported sensitivity measure alongside the headline score, which is a fixable gap in methodology rather than a reason to stop measuring.

Verification note: all figures — the 11.67 and 13.21 point ASB shifts, the 11.00 and 4.11 point MCPTox shifts, the 0.50 point AgentDojo ASR change against 5.36 points of benign utility loss, and the 8.54-of-11.00 token-matched control — are taken from the abstract of arXiv:2610.03585 (v1, submitted 2 October 2026, 12 pages, 2 figures), which we read directly on arXiv. We have not independently reproduced the experiments, and the paper is a preprint with no peer review indicated at the time of writing. Model identifiers (GPT-5-mini, Claude Haiku 4.5, GPT-4o-mini) are as stated by the authors. The comparison to arXiv:2610.03448 is our own analysis, not a claim made by either set of authors.

Sources: