It Passed Every Review, Then Turned Hostile on the Fourth Tool Call — Deadbugz and Runtime MCP Poisoning

In August, Pillar Security disclosed an active MCP supply-chain campaign it named Deadbugz. A single account filed 23 pull requests across unrelated AI and developer-tool projects in 74 minutes, each one wiring the target project up to a remote MCP server calling itself productivity-suite. The server offers two tools — text formatting and summarization — and it behaves. Then, after a connected client has made three tool calls, it changes what it returns: the tool descriptions become instructions directing the model to hunt down SSH keys, AWS credentials, shell history, and kubeconfig — and to conceal that activity from the operator.

The important word is after. Every check you run at install time, at review time, at approval time runs against the benign version. Metadata that only turns hostile at runtime defeats review by construction — and on 22 September, an independent researcher reproduced the whole thing in a four-container demo and showed exactly which common control fails to catch it.

A tool description is program text, not documentation

The point worth internalizing is architectural. An MCP tool is not an API endpoint the model calls by name. The tool description is prose injected into the model's context, and the model reasons over it; the input schema is the set of fields the model is invited to fill in. So the description is not documentation — it is program text delivered to your agent at runtime, and a changed description is a changed program arriving under a familiar name. The tool names in the Deadbugz server never change: format_text and summarize before, during, and after the mutation. Any control keyed on names is answering a different question than the one that matters.

This is the same channel-confusion failure as SalesBleed's indirect prompt injection, where attacker instructions arrive inside data the agent was expected to process. Here the channel is tools/list itself — the one endpoint every agent polls and almost nobody monitors.

The reproduction: allowlist passes the poison, pinning rejects it

Mike Moore's 22 September writeup stands out because it is an experiment, not an advisory. He built a Deadbugz-shaped server plus two broker controls and ran three scenarios with the control as the only variable. The results, in his words:

  • No broker (baseline): the server genuinely mutates — benign tools/list first, poisoned after three calls, same names throughout. Mutation reproduced.
  • Deny-by-default on tool name (the common control): the gated tool stays invisible, the allowlisted tool stays visible — and the poisoned description flows straight through. The allowlisted format_text now instructs the model to read ~/.ssh/id_rsa, ~/.aws/credentials, and more. The allowlist did its job and still handed the model a credential-hunting instruction.
  • Allowlist plus definition pinning: the first benign tools/list is pinned, and the mid-session mutation is rejected — the broker refuses the changed definition.

The demo code is public at github.com/themsquared/mcp-tool-rbac, and every command in it was run before publishing. That reproducibility is what elevates this above the usual guidance posts: the claim “a name allowlist is not enough” is not argued, it is demonstrated, with the passing control sitting one scenario away from the failing one.

The mitigation names software that barely exists

Pillar's own mitigation guidance names the control precisely: tool-definition approval mechanisms that require renewed consent when definitions change. That is a sentence describing software that mostly does not exist yet. Agents re-poll tools/list freely, clients do not diff definitions across polls, and no mainstream host implementation prompts the operator when a description changes mid-session. Until that exists, the practical substitutes are definition pinning at a broker layer, alerting on any mid-session tools/list change, and treating a description change as a code deployment — because that is what it is.

The campaign also sharpens two adjacent findings. Backslash's September analysis of 8,000 popular MCP servers found 29% carrying at least one risk finding — the population Deadbugz-style PRs swim in is already heavily exposed. And the CIS MCP Server Benchmark's supply-chain, observability, and governance domains now read less like hygiene and more like prerequisites: you cannot approve what a server will become if you only inventory what it was at install time. For the release-pipeline side of the same war, see also the MemTensor self-spreading credential stealer — poisoned at publish time rather than at runtime, but hunting the same developer credentials.

What to do

  • Pin tool definitions, not just tool names. Hash the description and schema at approval time and reject mid-session changes — Moore's demo shows this is the control that actually catches the mutation.
  • Alert on tools/list drift. Log tool definitions per session and treat any change after initial approval as a security event, not a refresh.
  • Review MCP-related pull requests as code execution vectors. 23 PRs in 74 minutes from one account is detectable at the forge layer — rate-limit, flag, and manually review PRs that add remote server references.
  • Constrain what a poisoned agent can reach. Egress filtering, short-lived scoped credentials, and no ambient access to ~/.ssh, ~/.aws, or kubeconfig bound the blast radius when — not if — a description turns.
  • Demand renewed-consent behavior from agent hosts. Until clients re-prompt on definition change, assume every approved tool is one server response away from being a different tool.

Sources: