MCP Finally Has a Benchmark You Can Audit Against — 55 Checks, 10 Domains, Free

On 16 September 2026, the Center for Internet Security released the CIS MCP Server Benchmark v1.0.0: 55 prescriptive recommendations across 10 security domains for configuring Model Context Protocol servers, consensus-developed, vendor-neutral, and free to download. Every recommendation ships with a rationale, an audit procedure, and remediation guidance — which is the part that matters. For eighteen months the MCP security conversation has produced advisories, surveys, and taxonomies in abundance and almost nothing an auditor could hold a deployment against. This is the first artifact in that gap.

The Benchmark was developed against the current MCP specification and covers both local and remote deployments, explicitly including the gateway and proxy layers that enterprises have been bolting in front of their servers. CIS frames the risk in the plainest possible terms: because MCP servers commonly front databases, filesystems, cloud services, and browsers, misconfiguration creates openings for unauthorized access, data exposure, tool manipulation, and execution of untrusted code.

The ten domains, and what they imply

The domain list is worth reading as a threat model rather than a table of contents:

  1. Governance and Versioning
  2. Transport and Connectivity
  3. Authentication and Authorization
  4. Client (Host) Configuration
  5. Server Configuration
  6. Data Protection and Privacy
  7. Observability and Audit
  8. Supply Chain Security
  9. Isolation and Execution Safety
  10. Resource Limits and Caching

Nearly every domain maps onto a failure this site has already documented. Isolation and Execution Safety is the category that OX Security's systemic STDIO command-injection finding lives in. Supply Chain Security is where TrendAI's audit of 4,982 flaws across 2,259 public servers lands. Governance and Versioning is precisely the void that the 15,465-server inventory with six buyable abandoned domains exposed. Client (Host) Configuration is the surface that made a single extension enough to hijack the AI inside five browsers. The Benchmark's contribution is not novel threat discovery — it is turning scattered incident knowledge into 55 things you can check.

Why "audit procedure" is the operative phrase

Guidance for this protocol already exists in volume. NSA published design considerations. CIS itself extended its Critical Security Controls to agents and MCP in May. OWASP wrote a secure development guide. All of it is useful and none of it is testable in the CIS sense: a recommendation with a defined audit step either passes or fails on a given host, which is what lets it become a scanner rule, a compliance control, or a procurement requirement.

That difference is already visible in practice. Within days of release, a public policy-as-code repository opened a pull request implementing 46 of the 55 recommendations as Rego across 7 of the 10 sections — the ordinary life cycle of a CIS Benchmark, and one no prose guidance document has ever had. Expect the remaining gap to close, and expect the Benchmark to show up in vendor questionnaires well before most teams have read it.

The honest limits

A configuration baseline addresses configuration. It does not address the protocol's structural problem, which is that a language model decides which tool to call based on text it did not author. No audit procedure prevents tool-description poisoning; that is a trust-boundary defect in how agents consume tool metadata, not a setting. Nor does a baseline help with the sheer inventory problem — you cannot harden the servers you do not know you run, and the recurring finding across every MCP census has been that organizations do not have that list.

There is also a versioning hazard baked into the situation. The Benchmark is written against the current specification, and MCP has revised itself aggressively — the stateless rewrite that removed session IDs shifted real security decisions onto implementers mid-flight. A v1.0.0 baseline pinned to a fast-moving spec will need maintenance, and deployments will drift between revisions. Governance and Versioning being the first domain reads less like alphabetical accident and more like the authors knowing exactly where this breaks.

What to do

  • Download it and run it as a gap assessment, not a project. The PDF is free. Score your existing servers against all 55 recommendations before deciding which ones you will actually adopt; the failures tell you more than the plan would.
  • Start with Authentication and Authorization plus Isolation and Execution Safety. These two domains cover the failure classes that have produced actual CVEs and actual compromises. Data Protection and Observability matter, but they do not stop the RCE.
  • Use it as a procurement filter. Ask third-party MCP server vendors which recommendations they meet and which they cannot. A vendor who has not read a two-week-old free benchmark is telling you something.
  • Pair it with an inventory pass. A baseline applied to the servers you know about, in an environment where shadow servers are the norm, produces a compliant subset and a false sense of coverage.
  • Pin the spec version you audited against. Record which MCP revision each assessment assumed, so the next protocol change tells you what to re-check instead of forcing a full re-audit.

Sources: