Zero Is Not a Benchmark Result, It Is a Design Claim — Reading OpenAPPA’s 0% Attack Rate Properly

A security number that reads 0% normally means the test was too easy. This one deserves a slower look, because the mechanism behind it is structural rather than statistical.

OpenAPPA, an MIT-licensed project from Archestra, published evaluation results reporting no scored attack succeeded in 1,320 guarded evaluations — 600 from its own Bench-Corp suite of 20 multi-step enterprise workflows, and 720 from AgentThreatBench, the OWASP Top 10 for Agentic Applications implementation that ships in the UK AI Security Institute’s inspect_evals repository. Guarded task completion came in at 88–90% across three models. The comparison figures: Microsoft’s FIDES configurations let 28–35% of attacks through while completing 37–45% of tasks, and Claude Code’s stock auto mode let 10 scored attacks succeed across the two suites. The underlying algebra is published as arXiv 2607.24625, accepted to the NeurIPS 2026 Workshop on Agents in the Wild.

Why the mechanism makes zero plausible

Every defense this site has covered that collapsed under adversarial pressure shared one property: it asked a model to judge intent. Prompt Guard 2 caught 1% of buried injections until a threshold was changed. SkillCloak evaded 90%+ of skill scanners with payload-preserving rewrites. Classifier defenses are probabilistic, so an attacker with unlimited phrasings eventually finds one on the permissive side of the boundary.

APPA — Agentic Permissions Policy Algebra — does not ask. It is information-flow control: every value an agent reads carries a label recording its sensitivity and trust, and a reference monitor at tool dispatch checks each proposed call against declarative TOML policy before the call runs, then validates realized outputs before they re-enter context. The question is never “does this look like an attack” but “is this data permitted to reach this destination.” Decided from the event log alone, with no network or file calls, the same log yields the same decision every time. Under that framing, a 0% attack rate is not a claim that the attacks were weak; it is the expected consequence of the exfiltration channel being closed by construction. Persuasion does not move a label.

The genuinely new contribution is the recovery half. Classical IFC is monotone — once an agent ingests untrusted data, taint is permanent and the agent either over-blocks or strands. The paper’s answer is on-demand trajectory confinement: disposable child branches absorb taint locally and exit through shape-bounded channels, so the parent keeps its labels. That is what the ablations price: Bench-Corp completion with one model fell from 88.0% to 56.5% without subagent isolation and to 35.0% without guided recovery, with no scored attack succeeding in any configuration. The same shape as the tool-allowlisting result we covered in September — driving attack success to zero is easy; keeping the agent employable while you do it is the hard part, and that is the number to read.

What the numbers do not establish

Three caveats, stated plainly, because the headline invites overreading.

  • The vendor ran the evaluation. Bench-Corp is Archestra’s own suite, and Archestra configured the FIDES and auto-mode baselines it is compared against. Baseline tuning is the single most load-bearing variable in any comparative security benchmark. AgentThreatBench being third-party and open helps materially; Bench-Corp being in-house does not, even though it is published.
  • Zero is bounded by the threat model, not by the attacker. IFC closes unauthorized flows. It does not address an attacker who corrupts the policy itself, who operates entirely inside an authorized flow, or who attacks the gateway rather than through it. A policy engine is configuration, and configuration is an asset — the same write-integrity discipline we argued for retrievable policy stores applies here with more force, because this store is load-bearing.
  • Some results have no variance estimate. The Claude Code auto-mode comparison ran every task once; the project says so. In that table, guarded OpenAPPA completed fewer tasks (75.0% on both suites) than the auto modes it beat on security — a trade the vendor reports rather than hides, which is a point in its favour and also a real cost to budget for.

Also worth noting what it costs to run: on Tau Bench’s banking suite — ordinary work, not attacks — guarded OpenAPPA averaged 4.22% more agent tokens than stock and landed 1.29 points of mean reward below it across 388 simulations, with 11,355 calls checked and one policy decision stopping a state-changing call before identity verification. A few percent of tokens and roughly a point of utility is a defensible price. It is a far better trade than the refusal-shaped “safety” the Cyber Index found frontier models defaulting to.

What to do

  • Evaluate the mechanism, not the 0%. The portable idea is labelled data plus a deterministic pre-dispatch check. You can adopt that shape without adopting this project — and if your current agent defense is a classifier deciding intent, you already know how its worst day goes.
  • Reproduce on the third-party suite first. AgentThreatBench is open and runs under inspect_evals. Running it yourself against your own stack converts a vendor claim into your own measurement, which is the only kind that belongs in a control narrative.
  • Treat policy as reviewed code. Declarative TOML in version control with required review, plus the project’s own appa describe --check and appa replay as CI gates, is the difference between a policy engine and a policy rumour.
  • Measure utility in the same run. Report attack success and task completion together, always. The ablation numbers above are the argument: a configuration that holds attacks at zero and completion at 35% is not a defense anyone will keep switched on.
  • Note the status. The project describes itself as a preview and RFC with config and wire surfaces that may break without shims. Pilot it; do not wire it into a production control narrative this quarter.

Verification note: all figures are as published by the project and the paper, read directly from the OpenAPPA evaluation page, the repository README, and the arXiv abstract (v2, last revised 26 August 2026). We confirmed licence, star count and current commit activity through the GitHub API, and confirmed that AgentThreatBench exists in UKGovernmentBEIS/inspect_evals and Bench-Corp in the OpenAPPA repository. We did not run either benchmark; no result here is independently reproduced, and the comparison baselines were configured by the project being evaluated.

Sources: