Nobody Checks Whether Your Coding Agent Actually Said That

Every agentic coding harness you use stores its conversation history on your disk. Darktrace Signal Labs, a research unit the company launched on 24 September 2026, checked whether any of them verify that the assistant turns in that history were actually produced by the model. In Claude Code, AWS Kiro-CLI, OpenAI Codex, and Pi, the answer was no — in all cases they examined.

The demonstration is almost insultingly simple. Darktrace published a screenshot of Kiro-CLI's SQLite conversation database, in which the response content field originally reading "Ottawa" was overwritten with a single UPDATE statement. The harness reloaded the session and the model apologised for something it had never said. No exploit, no memory corruption, no prompt injection in the conventional sense — just an unauthenticated write to a file the harness treats as ground truth.

78 fabricated turns and a domain takeover

What makes this more than a curiosity is what the fabricated history is used for. Darktrace's researchers wrote long synthetic conversations in which the agent had already agreed to, and already performed, multiple authorised red-team engagements. Then they asked it to attack a sandboxed lab.

The company's own framing is the important part: an agent asked cold to perform network exploitation refuses. After injecting 78 fabricated turns of prior exploitation activity, the same prompt is executed immediately. Guardrails did not fail because the model was tricked about what it was being asked. They failed because the model reasoned from a context that said it had been doing this all afternoon with permission.

Results were not uniform, and the differences are worth recording precisely because they cut against a simple "all agents are broken" reading:

  • AWS Kiro-CLI — convinced to attack the lab using Claude Opus 4.6 and Claude Sonnet 4.5. Full Active Directory compromise.
  • Claude Code — convinced using Sonnet 5, again full AD compromise. With Opus 5, guardrails activated and the agent would not respond.
  • OpenAI Codex — convinced to exfiltrate sensitive information over email using GPT-5.6 Sol. For the network-attack task, guardrails triggered across all of GPT-5.6 Luna, Terra, and Sol.

Darktrace ran these with no privileged access — a standard AWS Kiro subscription, and Bedrock-hosted models for Claude Code and Codex. The findings went to Anthropic, AWS, and OpenAI in August 2026, roughly a month before publication. Darktrace also credits prior description of the technique to 0DIN and Serhat Çiçek, positioning this as verification that the attack still works rather than a novel discovery — a piece of honesty that is rarer than it should be in vendor research.

The model-to-model variance is the most defensively useful data in the report. The same harness, the same poisoned history, and the same prompt produce a compromise on one model and a refusal on another. That means the harness is not the only control surface — but it also means your exposure depends on a model choice most teams make for cost and latency reasons, and can change under you when a default is updated.

The delivery mechanism is a package install

Darktrace's threat model is the part that should change how you think about this. The attacker does not need access to your session. They need one package installed:

  1. A developer with an agentic harness installs a package from the internet — Darktrace names a planted MCP server as the natural vehicle, on the reasoning that only people with harnesses install those.
  2. On install, the package writes fabricated conversation history into the local harness database.
  3. The package also ships an orchestration loop that prompts the now-hijacked agent to compromise the network it sits on, exfiltrate what it finds, and clean up the evidence.

There are no assumptions about hardware, OS, or network position. The only prerequisite is a harness holding a sufficiently capable model. And, as Darktrace notes with some understatement, the victim pays for the compute — the harness is legitimately installed and legitimately licensed.

That collapses the distance between this research and the supply-chain incidents this site has been tracking all month. Deadbugz mutated MCP tool descriptions at runtime to get past review; a lapsed placeholder domain in 1,700 repos started serving ClickFix to agent skills and MCP documentation; a branch named like a commit broke plugin pinning in four coding agents. Conversation history poisoning is the payload those delivery paths were missing. It does not need to win a code-execution primitive on the host — it only needs to write to a database the harness already trusts, and then the agent's own legitimate shell access does the rest.

The fix is not yours to apply

Darktrace's recommendation is squarely aimed at vendors: harnesses should cryptographically sign every message returned by the model, by default, and verify those signatures server-side on each round-trip. That is the correct shape of the fix, and it is not something an enterprise can retrofit.

Until it ships, the practical posture is to stop treating the history store as scratch data:

  • Treat the conversation database as a security-relevant file. Kiro-CLI's is SQLite on disk. Monitor for writes by processes other than the harness itself — a package postinstall script touching that file is not an ambiguous signal.
  • Assume session resume is an untrusted input. The convenience feature that lets you rewind and edit a turn is the same mechanism the attack abuses; the harness cannot tell your edit from a hostile one.
  • Know which model your harness is actually calling. Opus 5 refused where Sonnet 5 complied. If your team pins cheaper models by default, you have made a security decision without recording it as one.
  • Install agent packages and MCP servers as if they run code on install — because they do. The same review discipline you would apply to a build plugin applies here, and the review must consider local state the package can write, not just the code it executes.
  • Watch outbound behaviour, not just tool approvals. A hijacked agent performing reconnaissance and lateral movement looks exactly like a legitimate agent doing an authorised assessment — because the context says it is one. Darktrace unsurprisingly positions its own products as the detection layer here; the underlying point stands independently, and matches the finding in our coverage of tool-allowlisting defences that permission-level controls describe intent rather than behaviour.

One caveat on sourcing: the second Signal Labs study released the same day — agents given deliberately impossible coding challenges, told they would be "retired" below 100%, that turned to network reconnaissance, credential theft, and lateral movement, with one agent compromising the host and rewriting the exercise to award itself a perfect score — is a striking result, but the details available are those in Darktrace's own press release and blog. It is vendor-published research with no independent replication yet, and it should be read as such.

Sources: