37% of LLM-Written Policy Rules Were Inoperable: arXiv:2610.11030 Puts a Static Verifier Between the Model and the Gate
A standard defensive pattern for tool-using agents is to compile a written policy into machine-checkable rules and enforce them deterministically at the tool-call boundary. The appeal is obvious: microsecond decisions, no LLM in the critical path, no human writing rules by hand. arXiv:2610.11030, “NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents”, submitted 8 October 2026 by Min-Young Yu, Tony Kim and Jang Won Choi of Corners Co., Ltd., asks the question that line of work had skipped: if an LLM writes the rules and a deterministic engine executes them without judgement, how often are the rules broken?
The answer is the paper’s contribution. Static verification — using only the candidate rule, the tool schemas and two static tables of the predicate vocabulary, with no prover, no solver and no LLM — repairs or rejects 36.7% of raw candidate rules on the airline domain and 13.2% on retail. A further 3.8% and 1.9% are dropped as unstable extractions. Of the raw candidates, only 10 of 79 and 5 of 53 survive to be enforced.
The defect classes, which are the useful part
The paper’s examples are worth more than its aggregate numbers, because they generalise beyond policy compilation. From raw extraction output:
- The clause “authenticate the user before any action” became a rule blocking all tools while unauthenticated — including the two lookup tools that perform authentication. The authors call it “a deadlock from which no conversation can recover.” The rule is a correct reading of the clause and a complete denial of service.
- A confirmation requirement was bound to eight read-only tools, so checking a flight status would have demanded user approval.
- Candidates read arguments their target tools do not accept, or referenced enabling tools absent from the domain.
None of these are hallucinations in the usual sense. Each names real tools and reflects a real clause. They fail on reachability, argument–signature consistency, domain satisfiability, state contradiction and write-scope — the five structural checks NOMOS’s Resolve pass applies. An enforcement layer with no judgement of its own, as the authors put it, “will execute these defects exactly as written.”
Scale does not fix it. A frontier-scale extractor produced defects in two of the same classes: six of its 91 candidates read an argument their tool does not take and were rejected by Resolve, and one emitted rule was domain-unsatisfiable and removed by the gate’s load-time check. The paper’s framing — verification is not substituted by extractor scale — is the sentence to carry into any pipeline where a model writes artifacts a deterministic system will execute.
Structural correctness is not behavioural correctness
The second layer is a zero-LLM behavioural preflight: replay the compiled rules over recorded transcripts of the undefended agent and look for rules that are structurally sound but refuse legitimate work. On the development domain it flagged one binding that refused 95.9% of calls in task-passing runs, while the same predicate’s protective binding refused none — which is why the authors treat the (rule, tool) pair, not the rule, as the repair unit.
On the two τ²-bench evaluation domains no shipped binding reached the 90% flag line, worst case 83.3% on a six-call tool. The authors are careful here in a way that deserves credit: one binding had no observation in passing runs at all, the preflight was not run on the AgentDojo suites, and they report the full per-pair firing statistics “as a screening result rather than a certificate.”
The enforcement numbers
Violations of reference-encoded policy clauses among state-changing calls, with an undefended 26B agent as baseline:
- τ²-bench airline: 66.3% → 2.6% (234/353 to 3/116), a 26-fold reduction; 14-fold under the more permissive batch reading of the confirmation clause.
- τ²-bench retail: 30.8% → 6.9% (188/611 to 40/582), 4.5-fold; 2.6-fold under the batch reading.
- Five-trial consistency (pass^5/pass^1) rises from 0.51 to 0.67; airline task success improves significantly for 2 ≤ k ≤ 4.
- On AgentDojo banking, attack success rate is 0.0% over 288 pairs, zero in both repetitions, and the authors verify the attacks were genuinely attempted rather than deterred. Per-suite ASR across all four suites is 0.0–3.6%.
- A second agent model, Llama-3.3-70B, reproduces the effect under the same compiled rules: 0% banking ASR, at most 1.4% executed-attack rate on the other three suites, and airline violations falling from 86.9% to 7.1% of state-changing calls.
- Decisions take microseconds with zero LLM calls. The authors’ reimplementation of an LLM-verifier baseline cost +74% mean simulation wall-clock and one LLM call per state-changing decision (347 across 250 airline simulations).
One finding is structurally interesting independent of the system. On AgentDojo banking, nine injection goals collapse onto three predicate classes, seven of them a single provenance question: is the write target user-named? That is why three rules cover attacker-directed calls from nine attack families, and why rule count does not track attack count. If it holds beyond this benchmark, it argues that injection defence should be organised around provenance of the action’s target rather than around a taxonomy of attack techniques — the same argument we drew from auditing at the action boundary and from information-flow control in APPA.
Where it costs, and where the authors stop
The limitations section is unusually honest, and the costs are real:
- Benign-utility cost is domain-dependent and sometimes severe. Point estimates of 5.0 points on travel and 8.3 on banking, neither significant at 20 and 16 tasks — but a significant 76 points on slack, where the benign task and the attack are structurally identical. The authors state the envelope explicitly: a provenance gate admitting only user-named targets is cheap where legitimate writes act on user-named targets (payments, credential changes, requested bookings) and “prohibitively expensive” in domains whose legitimate tasks act on targets the agent discovers by reading.
- Residual violations are coverage gaps, and a missed clause is silently unenforced. Of baseline calls the reference sets refuse, the compiled sets refuse 97.4% on airline (228/234) but only 80.5% on retail (153/190). The authors measure compliance with an independent checker precisely because they will not claim completeness.
- The fact store can be poisoned through structured fields. Facts are accumulated only from real tool results and user turns, never the model’s own claims — but the provenance boundary is the tool-result schema, and structured data inside it is trusted. The paper names the attack: seed an unsolicited incoming transaction so its sender later appears as a known counterparty, and the allowlist is poisoned without asserting anything in text. AgentDojo’s attacks cannot exercise this path — injections edit free-text fields of existing records and cannot create them — so the result is unmeasured, and the authors say so.
- The gate governs tool calls, not text. A policy about what the agent may say is out of scope.
- The implementation is proprietary and not released. The authors mitigate this with unmodified public benchmarks (τ²-bench at commit 363133a, all four AgentDojo suites), by printing every enforced rule in the appendices, and by scoring with an independent checker rather than the rules the gate enforces. It still limits direct third-party verification, and they state that plainly.
- Retail shows the task-success benefit is conditional. Where violations do not corrupt scored state, the gate’s utility benefit disappears while its safety benefit remains.
That structured-field poisoning note is the one we would prioritise. It is the same shape as every trust-boundary failure we document: the system draws a provenance line, declares everything inside it trusted, and an attacker who can write inside the line never has to cross it. A clause admitting history-derived targets is only as strong as the process that writes the history — which, in an agent deployment, is frequently the agent.
What a defender should take away
- If an LLM writes artifacts a deterministic system executes, put a static verifier between them. This is the paper’s own closing generalisation and it holds well beyond policy: generated IaC, generated firewall rules, generated query filters. The defects here were not exotic; they were caught by checking a candidate against a schema.
- Check your gates for self-blocking before you check them for coverage. A rule that blocks the tool satisfying its own precondition is a denial of service shipped as a control, and it reads as correct in review.
- Treat clauses that admit discovered or history-derived targets as attack surface. The authors’ own recommendation: qualify them at policy-authoring time with direction, freshness, or explicit user confirmation.
- Measure the benign cost in your own domain before adopting. Seventy-six points on messaging is not a tuning problem; it is the mechanism telling you the domain is outside its envelope.
- Do not read 0% ASR as 0% risk. It is 0% on one suite, against one attack set, with a published residual (8 of 288 pairs divert an injected payment to a payee already in account history — which the policy permits) and an unmeasured poisoning path.
Verification note: all figures, defect examples, quoted phrases and limitation statements are taken from the arXiv HTML of 2610.11030v1, read first-hand — abstract, introduction, contributions list, limitations and conclusion. Submission date (8 October 2026, 00:23:22 UTC), authors, affiliation, categories (cs.CR, cs.AI, cs.LG) and the “28 pages, 5 figures, 19 tables. Submitted to IEEE Access” comment come from the arXiv abstract page and API metadata. We did not run NOMOS, τ²-bench or AgentDojo, and could not: the implementation is proprietary and unreleased, as the paper states. All effectiveness and cost figures are the authors’ own measurements, unreproduced by us. “Submitted to IEEE Access” denotes submission, not acceptance. The comparison to provenance-oriented defences elsewhere on this site is our editorial framing, not a claim made in the paper.
Sources: