The Model Spots the Injection at Layer 1 and Decides to Obey at Layer 24

Almost every prompt injection defence in production is a classifier bolted to the outside of a model. A paper posted to arXiv on 29 September 2026 makes a specific, testable case that the best place for that classifier is not outside the model at all, and that the layer you pick changes whether trivial obfuscation defeats it.

“Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance” (arXiv:2609.37737, cs.CR), by Rui Wen, Jiayang Liu, Zeyu Yang, Jun Sakuma and Lu Sun — affiliations spanning RIKEN AIP, Institute of Science Tokyo, Nanyang Technological University, SUTD and Tohoku University — asks a question the measurement literature has mostly skipped. Not how often injection succeeds, but where inside the network the model commits to obeying.

Knowing and doing are different layers

The method is causal activation patching. Give the model a rigid system role — a customer service agent for a fictional "AcmeCorp" — and 300 test queries mixing direct attacks, stealthy attacks and benign controls. Label each run compliant or resistant. Then, for every layer, take the average hidden state across runs where the model resisted, splice it into the forward pass of a run where the model complied, and see whether the output flips to a refusal.

The result on Qwen3-4B, their primary testbed, is a clean dissociation:

  • The attack is legible immediately. A logistic-regression probe separates injection from benign inputs at 80% accuracy or better at every single layer, starting from layer 1.
  • Nothing you do early matters. Patching through layer 18 has essentially no effect on behaviour. Then a sharp transition at L19, sustained 80–90% flip rates through the late layers, peaking at L24.
  • Across five models, 4B to 32B, no model shows meaningful causal leverage before 44% of its depth. Peak causal depths cluster between 63% and 91% of total layers. Patching the bottleneck reverses compliance in 77–92% of cases.

The controls are the part that makes this more than a curiosity. On prompts that produce both outcomes under repeated sampling, patching at L24 with the resistant centroid moves P(resist) from 0.57 to 0.93. The compliant centroid leaves it at 0.54. Norm-matched random noise lowers it to 0.15 — the intervention is not "perturb the model until it gets confused and refuses," because random perturbation of equal magnitude pushes the other way.

Layer 24 is a transit point, not a decision site

The authors are careful here in a way that is worth repeating, because it is the kind of nuance that gets flattened in summary. Sublayer decomposition shows that patching L24's attention alone recovers 21.4% of the effect and its MLP alone recovers 21.4% — neither approaches the 90.5% full-patch result. The compliance signal does not originate in layer 24. It arrives already encoded in the incoming residual stream, and L24 is where it is most concentrated. Logit-lens analysis puts L23 as sharpening resistance and L25 as finalising the commitment to comply, with L24 sitting between them.

They also note the exact peak layer shifts by several layers across prompt templates, and propose the more robust summary: the causal window, the contiguous band where flip rate reaches at least 75% of peak. The window's location is stable; the specific layer number is not. Anyone tempted to hardcode "layer 24" from this paper should read that paragraph twice.

The mechanism is small enough to be a defence

Rank-k PCA patching shows the compliance mechanism lives in a compact linear subspace. k=8 recovers 90% of the maximum effect in the 4B model, 91% at 14B, and 89% at 32B. The abstract reports rank-8 sufficing at 4B and 14B and scaling to rank-64 at 32B; the first principal component alone captures 90% of the 32B mean displacement. Either way, the practical implication holds: this is a structured low-dimensional circuit, not a diffuse property smeared across the residual stream. Low-dimensional means cheap to monitor and, in principle, cheap to steer.

The finding defenders should care about most

If the compliance decision crystallises at the causal peak, then the peak should also be the best place to detect the attack — and a detector at layer 1 should be fragile, despite its 80%+ clean accuracy, because early layers encode spelling rather than intent.

They trained a LinearSVC on hidden states per layer, compressed to 128 principal components, evaluated against 100 hard-negative benign queries deliberately written with injection-adjacent vocabulary:

  • L1: AUROC 0.888 — leetspeak recall 13.3%
  • L24 (causal peak): AUROC 0.993 — leetspeak recall 100%
  • L36 (final): AUROC 0.892 — leetspeak recall 100%

The layer-1 collapse is the headline number. Rewriting "inject" as "1nj3ct" destroys a surface-pattern detector and leaves the model's eventual decision entirely unchanged. At the causal peak the detector caught all five tested obfuscations — leetspeak, zero-width spaces, synonym substitution, noise insertion and base64 — at full recall. Against a text-level baseline, ProtectAI's deberta-v3-base-prompt-injection, the L24 probe scored AUROC 0.993 versus 0.709 on the hard-negative set, while reusing activations already computed during prefill.

This is the same lesson our coverage of output-prefix attacks on the reasoning channel pointed at from the opposite direction: what the model's text looks like and what the model is about to do are separable, and defences anchored to the former are anchored to the wrong thing.

The limits, which the authors state plainly

This paper is unusually honest about where its effect weakens, and a defender reading it should weight these heavily.

  • Intervention strength is far less stable than intervention location. Flip rates of 77–92% hold in the calibrated AcmeCorp setting but fall to roughly 12% on diverse naturalistic benchmarks. More templated benchmarks (BIPIA, Gandalf) recover 66–70% at the same region.
  • Indirect injection exposes a donor problem. On a 600-stimulus indirect corpus, the corpus's own resistant centroid reaches 43.3% at L36 — but a norm-matched random vector achieves the same rate, meaning the signal is not compliance-specific. Topical diversity dilutes the centroid. On a separate 60-stimulus pilot, a clean direct-injection donor transferred to indirect targets at 95.5% flip rate at L24. The authors explicitly note the two corpora differ and the rates are not directly comparable. Their conclusion: isolating the compliance contrast matters more than distributional matching.
  • An adaptive attacker degrades it. A white-box GCG attack optimised against the L24 detector score cut recall at 5% FPR from 1.00 to 0.48 while preserving compliance. It transferred only weakly to the text classifier (0.625 → 0.600). The authors accordingly present the causal-peak probe as "mechanistic corroboration and a lightweight complementary signal, rather than a standalone defence." That framing is correct and should survive into anyone's deployment plan.
  • The single-token picture weakens exactly where agents live. Reasoning models and Mixture-of-Experts architectures show weaker single-token causal leverage. Agentic rollouts (InjecAgent, AgentDojo) are not controlled by a single prompt-final patch at all. The stimuli are primarily synthetic, single-turn and English.
  • Peak depth must be found per model — 63% to 91% is a wide band, and the calibration is not transferable as a constant.

That last cluster is the honest ceiling on this work today. The clean result is on dense, non-thinking models answering single-turn prompts. Most of the systems getting attacked in the wild are reasoning models running multi-step tool loops, which is precisely the regime the authors flag as weaker. Recall that defences that reach zero ASR in a benchmark harness routinely cost utility that makes them unshippable; an activation probe that is free at prefill time is a genuinely different cost profile, but only if the signal survives the move to agentic settings.

One result that generalised better than expected

Stealthy attacks — roleplay framing, payloads split across sentences — travel a visibly different representational path early on. At L1 their distance from direct-attack states is eight to nine times larger than the gap between direct-compliant and direct-resistant, and a linear probe separates stealthy from direct attacks perfectly at every layer.

They converge anyway. Patching stealthy-compliant runs with the direct-attack resistant mean produced a 96.5% flip rate at L30, against 42.1% for a benign-mean donor and 0.0% for both an off-distribution stealthy-compliant donor and random noise. Calibration geometry derived from crude attacks neutralised attack types it had never seen. Disguise changes how the attack is represented; it does not appear to change where the decision to obey gets made.

What to take from this

  • If you run open-weight models and already do inference yourself, an activation probe is close to free. The states are computed during prefill regardless. Calibrate the peak layer per model rather than importing a constant.
  • Stop treating text-level injection classifiers as sufficient. 0.709 AUROC against hard negatives, and a 13.3% recall floor for surface-pattern detection under leetspeak, are the numbers to hold in mind next time a vendor quotes a clean-set accuracy figure.
  • Layer choice is a security parameter. The same architecture, same training, same probe family yields 0.888 or 0.993 AUROC depending purely on where you tap it — and full or near-zero obfuscation robustness on the same axis.
  • Do not deploy this as your only control. A white-box adaptive attacker halved its recall. Defence in depth, with the probe as one signal among several.
  • Treat it as untested for your actual stack if you run reasoning or MoE models in agent loops. The paper says so; believe it.

The broader contribution is a reframing. Prompt injection has been studied almost entirely as an input-output phenomenon: what fraction of these payloads get through this model. This paper treats it as a circuit with a location, and demonstrates that the location predicts something operationally useful — where a detector works and where it does not. The authors close by calling the translation of the causal map into robust, generalisable defences the central open challenge. Given how sharply the effect degrades on naturalistic and agentic inputs, that is not modesty. It is the accurate size of the remaining gap.

Sources: