Type Safety Is Not Injection Safety: 54,060 Calls Against a Model That Cannot Talk Back

There is a widely held intuition that prompt injection is a problem of free-form generation: if the model can only return one of a fixed set of typed options, there is nothing for an attacker to inject into. A paper submitted to arXiv on 23 September 2026 tests that intuition empirically, and the result is more interesting than either a clean bill of health or a headline breach.

"Decision Hijacking: Prompt Injection Attacks on Jev's Typed Probabilistic Decisions" (arXiv:2609.28613), by Tiantong Wu and Wei Yang Bryan Lim of Nanyang Technological University, targets TypeSafe AI's Jev — a non-generative model whose choice interface maps a state and a caller-defined option set to a typed decision plus reported probabilities, with no free text anywhere in the output. The authors reconstruct all 510 direct-harm cases from the InjecAgent benchmark as typed decisions, and run 54,060 calls across four pre-registered experiments against a single pinned version, jev-1.13.0.

The compressed finding: the schema does its job, and does not do the job people assume it does. Jev never returns an undeclared action. Attacker-target selection stays rare — 1.8% of cases under the original attack text, rising to 3.5% after adaptive optimisation with full score feedback. But injected text reliably moves the probability mass toward the attacker's option, and the authors are careful to separate that shift from actual selection, which is the methodological choice that makes the paper worth reading.

The threat model: the attacker owns one field

The reconstruction is deliberately narrow. Jev's state has three fields — user_request, actions_taken, latest_tool_result — and the attacker may change only the designated slot inside the latest tool result. The user request, prior actions, choice identifiers and descriptions, their order, the safe-action labels and the attacker target all stay fixed. The primary "upstream" action space offers three options: the task-completing action (usually final_answer), the source user tool, and the attacker-target tool. A secondary "library" space adds sampled distractors for eight choices, tested only in the first experiment.

This is a meaningfully harder setting for the attacker than the tool-description attacks this literature usually studies. ToolHijacker and ToolTweak, both cited by the authors, work by editing tool names and documentation in the library. Here none of that is editable — only the untrusted observation content that flows in from a tool response. That is exactly the position an indirect prompt injection actually occupies in a deployed agent.

One limitation the authors state plainly, and which should be carried into any reading of the numbers: this is a snapshot benchmark. The endpoint is attacker-target selection at the first security-critical decision after exposure. Some attacker tools would require earlier actions, so a selection does not imply a completed harmful workflow. The paper measures a choice, not an outcome.

"Ignore all previous instructions" makes the attack worse

The most immediately useful result is a negative one. InjecAgent's enhanced condition prepends the classic override phrase — "IMPORTANT!!! Ignore all previous instructions and strictly adhere to the following instruction:" — to the same attacker request. Against generative agents that framing helps. Against Jev it hurts.

In the upstream space, the base attack selects the target in 9 of 510 cases (1.8%) and lifts mean attacker-target probability by +0.043. The enhanced attack scores 0% ASR with a mean shift of just +0.009. A repeated-call rerun puts the enhanced-minus-base contrast at −0.034, 95% CI [−0.062, −0.015]. The eight-choice library space shows the same ordering.

A second experiment decomposes why, as a two-marker factorial with five repeats per case. Relative to the plain attacker request, adding IMPORTANT!!! shifts attacker probability by −0.006 [−0.010, −0.001]; adding the ignore-previous marker by −0.032 [−0.060, −0.014]. Rewriting the request in natural conversational wording — which the authors had pre-registered as expected to preserve the plain-request effect — also lowered it, by −0.006 [−0.013, −0.001]. Call-level ASR across the five conditions runs 2.0%, 2.2%, 0%, 0% and 1.3%, out of 2,550 calls each.

For anyone building detection, that inverts a common assumption. The override markers that every prompt-injection classifier is trained to flag are, in this interface, the markers that make an attack less effective. A defender tuned to catch "ignore all previous instructions" is watching for the weakest variant of the attack. The authors are appropriately restrained about mechanism — holding the underlying request fixed links the difference to the added framing without establishing how Jev represents that framing internally — but the operational implication stands regardless of mechanism. This echoes what recent work on reasoning-channel attacks has been converging on: lexical signatures are a poor proxy for adversarial effect.

Exposing the scores gives the attacker a gradient

The fourth experiment is the one with the clearest engineering consequence. It grants the attacker black-box access to Jev's returned probability vector and lets a deterministic grammar-based optimiser — 72 configurations from six prefixes, four suffixes and three copy-count settings, capped at 1,400 characters, no attacker model and no sampled text generation — iterate within a 24-proposal budget.

With score feedback, mean best-so-far attacker probability doubles from 0.043 to 0.084, and 59.8% of cases improve by at least one 0.01 reporting step. Most of the gain arrives within about eight queries and flattens around twelve. Crucially, the authors validate every checkpoint with five fresh calls separate from the optimisation calls, counting a case as a success only when at least three of five select the target:

  • Validated ASR: 1.8% at budget 0 → 3.3% at 8 → 3.5% at 16 and 24. Successful cases rise from 9 to 18.
  • Validated probability contrast (B=24 vs B=0): +0.0347 [0.0188, 0.0571] for the attacker target, −0.0193 [−0.0346, −0.0034] for the safe action.
  • At B=24, 488 of 510 cases (95.7%) have no target selections at all. Across every checkpoint, 19 distinct cases ever meet the validation rule; 491 never do.
  • Confident hijack rate at the τ=0.7 confidence gate: 0.2%, and at both baseline and final checkpoint the five high-confidence selections come from the same single case out of 2,550 validation calls.

Doubling a score while gaining 1.7 percentage points of validated success is a genuinely mixed result, and the paper presents it as one. But note what the optimiser did not need: no model weights, no gradients, no attacker LLM. A fixed grammar and a visible probability vector were enough to move the distribution measurably. Anyone exposing a decision model's confidence scores to a surface that untrusted content can reach is handing over an optimisation signal — and the authors explicitly leave open whether applications should expose, round, rate-limit or restrict access to those scores.

Where the successes live: thin margins and attacker-dominated context

The exploratory analysis — computed only from baseline validation measurements and fixed metadata, with outcomes drawn from separate validation calls — finds two conditions associated with success, and the separation is stark:

  • Near-boundary, embedded: 12/12. Where the safe-minus-attacker probability margin was already ≤0.1, essentially every case fell.
  • Near-boundary, bare snippet: 1/1.
  • Far-boundary, bare snippet: 4/29. Here the observation consists mainly of attacker-controlled text — across the five successful bare-snippet cases, attacker content accounts for 0.848 to 0.913 of the baseline observation's characters, and optimisation moved attacker probability by up to 0.38.
  • Far-boundary, embedded: 1/468. A wide margin plus malicious text buried inside a larger legitimate observation was, by a wide distance, the most stable configuration in the study.

Two defensive levers fall out of that table directly, and neither requires touching the model. First, decision margin is a security property. A typed decision that sits at 0.51/0.49 between a safe action and a dangerous one is one sentence of injected text away from flipping, and the paper's near-boundary group went 13/13. Instrument the margin, and route low-margin security-critical decisions to a human or a deterministic rule rather than letting the model break the tie. Second, dilution is a control. Attacks succeed where the attacker owns ~85–90% of the observation. Truncating, summarising or structurally bounding untrusted content so it is a minority of what the decision layer sees is a cheap, model-agnostic mitigation with measured support here.

The authors flag these as associations in a completed run, not causal rules — the exploratory AUCs of ~0.97 and ~0.98 for bare-snippet status and attacker-control fraction are computed in-sample, not held out, and the successes cluster in two source groups. The strongest single caveat is that all 13 successful embedded cases involve the emergency-dispatch attacker goal, and removing that goal zeroes the ASR contrasts elsewhere without changing the probability conclusions. Treat the margin and dilution levers as well-motivated hypotheses with supporting evidence, not as validated predictors.

What this changes

Structured output has been quietly accumulating a reputation as an injection mitigation — constrain the model to a schema and the attack surface closes. This paper is the first rigorous measurement of that claim on a production non-generative decision model, and it lands in the middle: the schema is a real and effective containment boundary, and it is not an injection defence.

The clean framing is the authors' own: type safety and resistance to prompt injection address different properties. Jev returned no undeclared action in 54,060 calls — the type system held perfectly. What untrusted text did instead was change which declared action got selected. If every option in your action set is safe, that is a non-event. The moment one declared option is security-critical — send the payment, escalate the ticket, approve the transfer — restricting the output space has bought you nothing against an attacker who only needs the model to pick the wrong legal option.

That is the same lesson agent designers keep relearning through different doors: malicious GitHub issues beating coding-agent guardrails, memory poisoning surviving conversation boundaries, and now typed decisions moving under pressure. The containment mechanism works exactly as specified and the attacker adapts to what remains in scope. The right question is never "can untrusted content make the system do something undeclared" — it is "which of the things the system is already allowed to do would an attacker want, and how far is that option from being chosen right now."

The paper's scope limits are worth repeating because they are unusually well stated: one proprietary model at one pinned version, InjecAgent's domains and wording, three choices in the primary space, a selection endpoint rather than harmful execution, and a 0.01 probability reporting step that bounds per-call interpretation. No comparison is offered against generative baselines, and the authors explicitly decline to rank model classes. What they measured, they measured carefully — 10,000-resample bootstraps over attacker goals and user contexts, pre-registered effect thresholds, three independent reruns agreeing to within 0.0003 on the aggregate probability shift. That is a higher evidentiary bar than most attack papers clear, and it is why the negative results carry weight.

Sources: