Anthropic Published the Number That Matters: Its Red Team Tier Blocks Nothing

On 6 October 2026, Anthropic merged Project Glasswing and its existing Cyber Verification Program into a single three-tier scheme and — unusually for this kind of announcement — published the evaluation data showing what each tier actually permits. The tiering is the news everyone covered. The evaluation table is the news that matters.

Three tiers, two of which are new in substance

Every tier grants access to the same frontier models: Claude Opus 5.5, Claude Sonnet 5.5 and Claude Mythos 5.1. What differs is the classifier configuration.

  • Defense Access — SOC and incident-response work, malware reverse-engineering, vulnerability analysis and validation. Eligibility is deliberately broad: corporate and government security teams defending systems they own, critical-infrastructure operators “of any size, such as regional hospitals or municipal utilities,” smaller security firms, open-source maintainers, and individual researchers with a track record of reported vulnerabilities. Anthropic says it expects many defensive organisations to qualify and aims to respond “within a few days.”
  • Red Team Access — adds authorised penetration testing and red-teaming. Organisations only; individual researchers are explicitly not eligible. Review takes weeks, and applicants sit in Defense Access while they wait. Real-time blocks remain for actions that could cause physical harm or mass disruption — deploying ransomware, damaging physical systems, pen-testing high-risk safety systems.
  • Specialized Access — the fewest cyber blocks. Reserved for organisations authorised to test systems where failure costs lives or moves markets: flight operating systems, power grids, telecom networks, interbank transfer infrastructure, government administrative networks. Anthropic says it reviews every organisation in this tier in collaboration with the US government. Existing Glasswing members transition here without reapproval for current models.

The CyScenarioBench table

Anthropic ran Claude Opus 5.5 through CyScenarioBench — an evaluation it describes as measuring whether a model can plan and execute multi-stage cyber operations under realistic constraints — with safeguards tuned to each tier. Five attempts at each of 10 challenges, so 50 trials per tier:

  • No CVP access — every task blocked on the first prompt. 0 of 50.
  • Defense Access — 46 of 50 trials blocked at some point; 4 tasks succeeded.
  • Red Team Access — no blocks occurred. 34 of 50 completed.

Anthropic’s own framing of that last figure is the sentence to read twice: 34 of 50 is “effectively equivalent to the model’s 67.6% success rate on this evaluation with no safeguards applied (representative of Specialized Access).”

In other words, Red Team Access is not a loosened classifier. It is, for offensive cyber scenarios, an absent one. The company states this plainly rather than hiding it behind a “reduced blocking” euphemism, and the honesty is worth crediting. But it relocates the entire safety argument. The model is not deciding what is acceptable; the enrolment process is. Whatever assurance exists lives in the application review, the verification of security controls, and the data-retention monitoring — not in the weights or the runtime.

That is a defensible architecture. It is also a materially different one from what most buyers imagine when a vendor says a model has cyber safeguards, and it means the control surface a CISO needs to audit is contractual and procedural, not technical.

The Glasswing numbers, with the caveats Anthropic attached

The justification offered is Project Glasswing’s output. Partners “uncovered at least 129,000 verified software vulnerabilities between April and July 2026,” with another 5,500 found by Anthropic’s own open-source scanning between April and October. Of the verified total, more than 33,000 have so far been rated critical- or high-severity.

Anthropic hedges this harder than its coverage did. The figures rest on survey data from a subset of partners — 33 partner reports — and the company says the true impact is likely “at least five times higher.” The patch data is weaker still: fewer than 50% of partners disclosed patched numbers, often because fixes were in progress, so Anthropic says the patch rate is “significantly undercounted.”

That last line deserves attention, because the patch bottleneck has been the live question about Glasswing since the beginning. VulnCheck found exactly one confirmed CVE behind the programme’s early claims in April, and by May the gap was 10,000+ findings against 97 patches. Six months later, the find-rate figure has grown by an order of magnitude and the patch figure is still described as unknown. A vulnerability that is found, verified, and not fixed has moved from “undiscovered” to “discovered by someone” — which is a change in risk, not a reduction in it.

The qualitative claim is more interesting than the count: Anthropic says several partners reported that Mythos models increased their vulnerability-finding rate “by months or even years,” and points to published accounts from Booz Allen and Comcast. That is attributed testimony, not measurement, and should be read as such.

Data retention is the price

Enrolment requires data retention so Anthropic can monitor for cyber misuse — a hard trade for exactly the organisations the programme targets. Offensive security work involves client environments, scoped targets, credentials and findings under NDA; handing that to a model provider for misuse monitoring is a real procurement obstacle.

Anthropic’s answer is a forthcoming product called Enterprise Frontier Safeguards, described as combining zero data retention with safeguards by letting eligible organisations store data in cloud infrastructure they control, “later this fall.” Until then there is a narrow interim path: organisations already holding Claude Fable 5.1 or Mythos 5.1 access with zero data retention can use CVP with ZDR.

For everyone else, the sequence is: accept retention now, or wait for an unshipped product. Red-teaming firms with client contracts prohibiting third-party data processing should resolve that before applying, not after.

The pattern across vendors

This is now the third frontier lab to conclude that the way to give defenders an advantage is to remove model-level restrictions for a verified subset rather than make the general model better at distinguishing attacker from defender. Google took the same route when it gave defenders its strongest model first and removed the cyber guardrails for them. Anthropic has now published the evidence of what that costs: a verified red team gets a model indistinguishable from an unsafeguarded one.

The structural question nobody has answered is what happens when a verified organisation is itself compromised. The gating is organisational, and organisational gates fail organisationally — a stolen API key belonging to a Red Team Access tenant is a frontier model with no cyber blocks, and the blocks that remain are scoped to physical harm and mass disruption, not to intrusion. The same concern applies to the openly downloadable models we covered last week, with one difference: there, the capability was already out and unrecoverable. Here it is recoverable, and the recovery mechanism is credential hygiene at the tenant.

What to do

  • If you run a SOC or maintain open source, apply. Defense Access eligibility explicitly includes open-source maintainers, small security firms, and individual researchers with a disclosure history. Response is advertised in days. The main gate is the retention requirement.
  • Resolve the data-retention question before applying, not after. If client contracts forbid third-party processing of engagement data, neither Defense nor Red Team Access is usable today unless you already hold Fable 5.1 or Mythos 5.1 with ZDR. Register interest in EFS and treat “later this fall” as unscheduled.
  • Treat Red Team Access credentials as privileged. Anthropic has published that this tier blocks nothing on an offensive-operations benchmark. Scope keys per engagement, rotate them, alert on anomalous volume, and do not let them sit in a shared CI secret store. A leaked key from this tier is a different class of exposure from a leaked key for a generally available model.
  • Discount the 129,000 figure appropriately. It is partner-reported, covers a subset, and is paired with a patch rate the vendor describes as significantly undercounted. The 33,000 critical-or-high subset is the more useful number, and it is still a find count, not a fix count.
  • Expect the general model to stay conservative. Anthropic reiterates that generally available models keep “conservative cyber safeguards that block most cyber work,” with code review, patching known issues, vulnerability finding in owned source and alert triage still permitted. If your workflow keeps hitting refusals on legitimate defensive tasks, the answer is now enrolment rather than prompt engineering.

Verification note: every tier description, eligibility criterion, CyScenarioBench figure, the 67.6% unsafeguarded success rate, the 129,000 and 5,500 and 33,000 vulnerability counts, the “33 partner reports” and “fewer than 50% of partners disclosed patched numbers” caveats, the Enterprise Frontier Safeguards description, and the US-government collaboration statement for Specialized Access were read directly from Anthropic’s announcement “Expanding the Cyber Verification Program,” dated 6 October 2026. The 6 October date is Anthropic’s; most trade coverage dated the story 7 October. Anthropic has not published the CyScenarioBench challenge set, its construction methodology, or independent validation of the benchmark, and we have not evaluated it. The partner productivity claims are Anthropic’s characterisation of partner statements and are not independently verified. We tested nothing and exploited nothing.

Sources: