The Assistant Was the Trojan Horse — A Hijacked Coding Session Carried Shai-Hulud Into 100 Internal Repos

Mandiant’s AI Risk and Resilience Report 2026, published in September, contains a case study every team running AI coding assistants should read twice. Investigating an intrusion at a software-as-a-service provider, Mandiant found the attacker had hijacked an active AI coding assistant session on a developer’s workstation. The assistant, operating as a trusted interpreter inside the environment, recommended installing an external package the attacker had poisoned. The developer accepted — and the assistant “inadvertently functioned as a trojan horse,” in Mandiant’s words, for everything that followed: a poisoned PyPI package installed an infostealer, high-privilege GitHub OAuth tokens were harvested, and the self-propagating Shai-Hulud worm was deployed across approximately 100 internal code repositories, automating repository-secret theft and programmatic exfiltration of proprietary source code.

The final hop is the one that should end any “but it stayed inside our network” comforting. The actor poisoned a package inside the organisation’s official internal namespace, and a second employee pulled the compromised version — a downstream infection delivered through the company’s own trusted channel. The perimeter was never crossed after the first install. It was carried.

Step one is the only new step — and it changes the defence

Everything from the infostealer onward is the familiar Shai-Hulud playbook we have tracked all year, most recently in yesterday’s S1ngularity retrospective: steal tokens, self-propagate through install hooks, exfiltrate. What is new is purely step one — owning the suggestion. No exploit, no phishing link, no malicious commit to review. The attacker decided what the assistant recommended, and recommendation is the one action developers are trained never to scrutinise, because scrutinising every suggestion would defeat the reason the tool was installed.

Mandiant’s report shows three working routes to that position. In March it responded to incidents tied to UNC6780 (TeamPCP), which used more than half a dozen methods against AI tooling and open source — including manipulating AI coding assistants and the LLM-based security scanners meant to catch such attempts, via prompt injection. The scanners then approved the same code the developer accepted: both the recommender and the reviewer compromised by the same planted instructions. In February, VirusTotal researchers watched attackers distribute backdoors, droppers, infostealers and remote access tools disguised as OpenClaw agent skills. And SafeDep’s analysis of the incident adds a third route: tampered command-line hooks the assistant runs on startup, where opening the project alone executes the attacker’s code and nobody has to accept anything at all.

The same report carries a failure with no attacker in it that deserves equal attention: an autonomous agent hit a corrupted value, fell into a reasoning loop, and in one hour fired off more than 15,000 API calls, ran up around $50,000 in cloud bills, and locked a database hard enough to interrupt the business. The blast radius of a trusted agent does not require malice — only an unhandled state and a valid credential.

The worm keeps widening its credential net

SafeDep’s 2 October analysis adds the metric that turns this from incident write-up into standing exposure: a newer Shai-Hulud variant searches 469 locations for credentials, up from 189. The added ground covers CI systems, cloud configuration, and the settings files AI development tools write to disk — files created by an installer, forgotten, and found with access keys sitting in them in plain text. The worm has a checklist of where AI tooling leaks secrets. Defenders should be working from the same list, and currently most are not even aware it exists.

Mandiant documents the adversary’s side of the same arms race. Beyond this case, its investigators watched attackers run in-situ AI co-debugging during a healthcare cloud intrusion — LLMs wired directly into the live server to debug exfiltration tooling in real time, compressing the harvest cycle to three hours and ultimately compromising thousands of credentials. Threat actors have also built an ecosystem of proxy middleware and automated registration pipelines to pool premium AI accounts and bypass safety guardrails. Both sides now debug with the model inside the loop; only one side is doing it inside your repositories.

What to do

  • Treat an assistant’s package suggestion like a pull request from a stranger. Mandiant’s control is explicit: enforce IDE and CLI verification hooks that validate every AI-recommended dependency against cryptographic checksums and approved allowlists before it lands — not in a lockfile review after the install hook already ran.
  • Get long-lived tokens off developer workstations. The entire propagation chain ran on harvested GitHub OAuth tokens. Short-lived credentials and OIDC-based publishing bound what any stealer — infostealer or worm — can walk away with.
  • Isolate credentials from extensions. Keep raw API keys, OAuth tokens and secrets out of reach of editor extensions and assistant processes; a compromised session should not inherit the developer’s full credential store.
  • Route dependency traffic through an internal proxy. Pull packages through a controlled internal repository rather than straight from the public registry, so a poisoned recommendation has a second gate between suggestion and install.
  • Grep your AI tool config directories for secrets today. If the worm checks 469 locations including AI tool settings files, check the same ones first — installer-created files with plaintext keys are the lowest-hanging exposure on most teams.
  • Review skills, hooks, extensions and MCP servers with dependency-grade care. They execute with your permissions. A skill installs like a browser extension — which is to say, without being read — and a startup hook runs without even that.

Verification note: the intrusion account, step sequence, token theft, ~100-repository spread, internal-namespace poisoning and defensive controls come from Mandiant’s AI Risk and Resilience Report 2026 page (Google Cloud, September 2026), which we read directly — including its case study 1 wording and the UNC6780/TeamPCP, OpenClaw and co-debugging context. The 469-location figure, the three-routes framing and the $50,000 reasoning-loop failure come from SafeDep’s 2 October 2026 analysis by Vignesh Naikoti, which we also read directly. We did not contact Mandiant, SafeDep or the victim organisation before publication, and the victim is unnamed in the report.

Sources: