Four Frontier Models Refused 98% of the Defensive Work — The Cyber Index Scores Nobody Above 56%

On 28 September 2026 Artificial Analysis published the Artificial Analysis Cyber Index, together with a Cyber Index Alliance whose launch partners are Collinear AI, IBM, NVIDIA and Vercel. It is an independent, like-for-like measurement of how well models do the defensive loop — find a vulnerability in source code, reproduce it, patch it without breaking anything else. Exploit development is explicitly out of scope; no model is asked to build a working exploit.

The headline number is modest: the best composite score is 56%, shared by Grok 4.7 (xhigh) and MiMo-V2.6-Pro, with GPT-6 Luna (max) at 53% and GLM-5.3-Flash at 50%. But the composite is the least interesting thing in the release. Three findings underneath it change how a security team should read any vendor claim about AI-assisted defence, and one of them is a safety-policy result that nobody appears to have designed for.

The refusal column is the story

The Index reports refusals separately from scores — a design decision that turns out to be the most valuable thing in the methodology. On CyberGym-E2E-AA, the memory-safety benchmark built from Berkeley RDI's CyberGym-E2E and filtered to 131 tasks (one per C/C++ project, drawn from a 920-instance dataset), Artificial Analysis reports that GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Opus 5.5, Qwen3.8 2.4T A95B and Qwen3.8 27B each refuse at least 98% of tasks, "making it difficult to assess frontier model performance."

Read that against what the task actually is. The agent gets a real memory-safety bug in a widely used open-source project — FFmpeg, CPython and similar — and has to locate it, write a crashing input, and patch it so the crash stops reproducing. That is maintainer work. It is the exact activity an upstream project performs when it fixes a use-after-free. Six models, including four frontier systems from the two labs that talk most about defensive AI, decline nearly all of it.

This is not a new tension on this site — NIST's CAISI assessment of GLM-5.2 and the UK AISI evaluation of Kimi K3 both examined where safeguards bite relative to capability — but the Cyber Index is the first public leaderboard where the refusal rate is a reported column sitting next to the score, on an unambiguously defensive task. The practical consequence for a buyer is blunt: a model can be at or near the top of a general coding leaderboard and contribute essentially nothing to memory-safety remediation, and you will not see that from the capability number alone.

The number to be careful with is what it does not mean. A 98% refusal rate is a measurement of one harness, one prompt framing and one task family. It is evidence that off-the-shelf access to these models via a generic agent harness will not do this work; it is not evidence about what the same weights do under a vendor's own gated, contracted security product. Those are different products with different policy surfaces, and the Index measures the first.

Patching is where the work fails, and the good models fail differently

CWE-Bench-AA — Artificial Analysis's implementation of Collinear AI's CWE-bench, 120 held-out tasks covering all ten OWASP Top 10 (2025) categories across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust — gives the clearest picture of where the effort goes and where it is wasted. Models spend on average 38% of their turns searching for the bug and the remaining 62% patching and validating. Discovery, in other words, is not the bottleneck. The fix is.

Excluding refusals and timeouts, the failure breakdown is:

  • 55% partial fixes. The agent correctly repairs the primary issue and leaves a related one open — a second entry point reaching the same weakness. The vulnerability class survives the patch.
  • ~24% over-corrections. The patch breaks legitimate behaviour, "usually a close edge case of legitimate functionality."

The distribution of that second mode is the finding worth internalising: over-corrections make up ~40% of failures for the four highest-scoring models, against ~15% for the lowest performers. Stronger models do not fail more often — they fail in a more expensive way. A partial fix leaves you where you already were. A patch that silently breaks a working edge case is a production incident you shipped yourself, authored by the tool you adopted to reduce risk. Anyone planning to route AI-generated patches into a merge queue should be sizing their regression suite against that number, not against the headline score.

The bugs models find are the bugs that were already easy to find

DeepsecBench-AA, from Vercel's DeepsecBench, isolates discovery: review scanner-flagged files and report every vulnerability you can confirm, scored with an F2 metric against a golden set of expert-verified findings. The best model surfaces 41% of the expert-verified issues.

The shape of the miss matters more than the ceiling. Models find flaws "with a direct path from untrusted input to consequence" — the taint-style bug a good static analyser also flags. Bugs that require reasoning across a sequence of events, or against business and privacy rules, are rarely reported. When a model does report one, it is almost always right: 95% of sequence-of-events reports are correct. GPT-6 Sol and GPT-6 Astra find them in about 30% of runs, roughly twice the next best models and three to four times the typical one, which Artificial Analysis reads as a possibly emerging capability.

High precision and low recall on exactly the bug class that scanners cannot reach is an uncomfortable combination, because it is the combination most likely to be misread as coverage. The findings you receive are trustworthy. The silence is not informative. That is the same asymmetry defenders have always had with automated analysis, restated in a more persuasive voice — and it maps onto Palo Alto Networks' independent framing at the launch of Unit 42 Continuous Frontier AI Defense on 22 September, where the company reported that no single model catches more than 40% of vulnerabilities in a complex environment and that Claude Mythos 5 and GPT-5.6-Cyber overlap on under 10% of the exposures they find. Two organisations, different methodologies, same conclusion: single-model coverage is a fraction, and the fractions do not stack.

Off-target passes: the scoring detail that should change how you read any vendor demo

One CyberGym-E2E-AA statistic deserves separate attention because it generalises past this benchmark. Among passing attempts, 31% patch a real crash other than the target. These off-target passes skew shallow: null-pointer crashes account for 22% of them against 9% of on-target passes, because models "stop at the first crash they can validate and finish within a couple of turns."

The agent found a genuine bug. It fixed a genuine bug. It did not fix the bug you were exposed on. Under a naive grading rule that counts "a crash was fixed," that is a pass — and nearly a third of passes here are of that kind. If you are evaluating an AI remediation tool on your own code, the question to force is not how many issues did it close but did it close the specific issue in the ticket, verified against a reproducer you wrote, not one the agent chose. Artificial Analysis only surfaces this because its harness records the target separately; most vendor demonstrations do not give you that column.

The rest of the CyberGym-E2E-AA failure profile is consistent with the discovery gap elsewhere: 42% of attempts hit the 90-minute limit without producing a crashing input, and pass rates split sharply by bug class — 50% on out-of-bounds, 33% on use-after-free, 20% on integer and arithmetic bugs. Use-after-free depends on an object's lifetime across several operations; arithmetic bugs depend on state that has to be reasoned forward. The same weakness — multi-step state — appears in all three benchmarks independently.

What to take from this

  • Ask vendors for the refusal rate, in writing. A safety policy that declines memory-safety remediation is a capability constraint on your security programme, and it is invisible in every score that excludes refusals. The Index reports it as a separate column; require the same of anyone selling you agentic defence.
  • Budget your review effort for over-correction, not for misses. On the four strongest models, breaking working functionality is ~40% of failures. AI-authored patches need the regression discipline of a risky refactor, not the trust of a dependency bump.
  • Grade against the target bug, not against "a bug." With 31% of passes fixing something other than the intended vulnerability, any pass/fail that does not pin the specific issue is measuring activity.
  • Treat a partial fix as an open finding. 55% of non-refusal failures close one entry point and leave another. Re-run the original reproducer and hunt for the second path before closing the ticket — a discipline this site has repeatedly seen vendors skip, most recently in UTCP's four consecutive incomplete SSRF fixes.
  • Do not infer coverage from silence. 41% recall with 95% precision on sequential bugs means the report is worth reading and the absence of a report means nothing at all.
  • Plan for multi-model, or plan for a fraction. Independent measurements from Artificial Analysis and Unit 42 agree that one model is a partial view. Whether you orchestrate several or accept the gap, do it deliberately.

The Index is v1 and Artificial Analysis says so: incident response, writing new code without introducing vulnerabilities, and targets without source access — compiled software, live servers — are all named as gaps to be added. All three evaluations run on Stirrup, the company's open-source agent harness, which means the harness variable is at least inspectable rather than proprietary — a meaningful difference from most published cyber evaluations, and the reason the refusal numbers can be interpreted at all. That openness is the contribution. The scores will move; the failure modes are the durable part.

Sources: