Half of What AI Finds Is RCE — Inside Google’s Doubled Disclosure Year

Google Threat Intelligence Group published Vulnerability Discovery and Exploitation Trends in the AI Era — written by Robin Grunewald, Supriya Mazumdar and Kelli Vanderlee — covering disclosures from 1 January 2025 through 31 August 2026. The headline is easy to misread: monthly CVE disclosures doubled, from 5,045 in January 2026 to 10,477 in July and 10,740 in August. That number is close to meaningless on its own, and GTIG says so in the same breath.

The useful findings are in the ratios underneath it, and two of them should change how a security team reads its own backlog.

The disclosure count is inflated and GTIG tells you by how much

Automated CNA assignment across open-source ecosystems manufactures volume. GTIG's own example: CVEs whose description mentions the Linux kernel generated roughly 5,000 records between January and August 2026, with zero observed in-the-wild zero-day exploitation. That is on the order of half the growth in the raw curve, attributable to one project's assignment policy and carrying no observed adversary interest.

Strip the noise and the exploitation picture is proportionate rather than explosive. GTIG recorded 141 distinct disclosed-and-exploited vulnerabilities in the first eight months of 2026 — already past the 127 for all of 2025 — lifting the monthly average from 10.5 to 18. But only 0.23% of 2026 disclosures, about one in 431, were ever seen under active exploitation. Since May 2026 exploitation growth (+127% indexed) has tracked disclosure growth (+128%) almost exactly. Exploitation is scaling with the landscape, not outrunning it.

A vulnerability-management programme that responds to a doubling of CVE volume by doubling patching effort is responding to CNA policy, not to threat. GTIG's explicit recommendation is the transition "from unprioritized mass-patching to threat-intelligence-driven triage" — which reads as vendor boilerplate until you put the 1-in-431 figure next to it.

Zero-days are flat. N-days are the growth market.

Zero-day exploitation rose only marginally, from an average of 8 per month in 2025 to 11 in 2026, holding between 8 and 12 through mid-year before jumping to 22 in August. Zero-days still make up 62% of all observed exploited vulnerabilities, but they are not where the growth came from. Exploitation of High-Risk vulnerabilities — GTIG's own rating, not CVSS — more than doubled, from 28 in 2025 to 75 in the first eight months of 2026.

GTIG's reading, offered as possibility rather than conclusion: threat actors may be finding it "more accessible or efficient" to point LLMs at version diffs, patches, advisories and public PoC code to weaponise n-days quickly, rather than to hunt new zero-days. That is a materially cheaper use of the technology than zero-day discovery, and it attacks the one window defenders actually control — the gap between a patch shipping and a patch landing. This site has tracked that window narrowing repeatedly, most bluntly when attackers needed 11 hours against a WordPress core fix.

Where that pressure lands is unchanged: 14% of 2026 exploitation hit edge and security appliances, 11% hit enterprise directory and collaboration hubs, and over 65% of exploited edge flaws carried High or Critical risk ratings. Unauthenticated public management interfaces on perimeter devices remain the premier initial-access vector, in part because they sit in enterprise EDR blind spots.

The finding with the sharpest edge: AI finds different bugs

GTIG separated vulnerabilities it could identify as likely AI-discovered — via lab and vendor ledgers, and by parsing CISA advisories, MITRE records and vendor bulletins for explicit attribution to autonomous agents such as Hacktron AI and AISLE — and compared their profile against everything else.

The distribution inverts. By GTIG risk rating, comparing vulnerabilities not discovered by AI against those that were:

  • Low risk: 69% of non-AI findings, but only 39% of AI findings.
  • Medium risk: 28% of non-AI findings, against 58% of AI findings — more than double the baseline.
  • High risk: 3% versus 4% — essentially unchanged, and worth noting that the shift is concentrated in the middle band, not the top one.

And on consequence: exactly 50% of AI-discovered vulnerabilities result in remote code execution, against 26% across the broader CVE ecosystem. AI-found flaws under-index on the cheap categories — information disclosure 8% versus 18%, data manipulation 5% versus 9%.

The honest caveat is that this is substantially a tasking artefact, and GTIG says as much: researchers point agents at critical infrastructure and privilege boundaries rather than running them as broad scanners for cosmetic findings. But GTIG also offers a capability explanation — agents synthesise fuzzing harnesses, model memory states and chain edge-case logic across C/C++ libraries, runtimes and hypervisors, surfacing memory corruption and logic bypasses that static analysers miss. Both can be true, and for a defender the distinction does not change the intake problem: the AI-attributed stream arriving in your triage queue is enriched for code execution by roughly a factor of two.

GTIG is equally clear that the count is an undercount, for two structural reasons worth remembering whenever anyone quotes an "AI found N bugs" statistic: public CVE repositories have no standardised metadata tag for AI attribution, and major cloud and SaaS providers silently remediate AI-surfaced flaws in production without requesting CVE IDs at all.

The collision case: CVE-2026-1731

The report's one confirmed instance of an AI-discovered vulnerability being exploited in the wild is CVE-2026-1731, an unauthenticated OS command injection in BeyondTrust Remote Support and older Privileged Remote Access versions, rated CVSS 9.8 and published 6 February 2026. It was found autonomously by the Hacktron AI research agent.

The timeline after disclosure is the part to sit with. GTIG observed one threat cluster exploiting it within four days of public disclosure, and five more within seven days — followed by privilege escalation, data exfiltration, and deployment of SNOWLIGHT, SPARKRAT and cryptominers. A defensive research agent surfaced a pre-auth RCE in a privileged-access product, and within a week six distinct clusters were using it.

GTIG frames this as an early indicator, not a trend, and that framing is right — it is a single case. What it demonstrates is that the two curves now intersect: the flaws defensive agents are good at finding are precisely the flaws adversaries most want.

AI as target: 2,076 CVEs, and half of them are orchestration

GTIG tracked 2,076 AI-related CVE disclosures across the 20-month window, with over 1,500 in 2026 alone, broken into eight architectural layers. The distribution is lopsided in a way that should surprise nobody reading this site:

  • AI orchestration & agent frameworks — 782 (2026). Flowise, Langflow, LangChain, Dify, LlamaIndex, AutoGen, CrewAI, Semantic Kernel, Letta, MCP, Pydantic-AI. RCE and command injection via untrusted workflow serialisation, insecure Python tool calling, SSTI. 50% of all AI-related flaws, up 347% in 2026.
  • AI web apps & portals — 230. SSRF via chat proxying, stored XSS in markdown rendering, LFI through document upload handlers.
  • Inference & serving infrastructure — 212. vLLM, Ollama, LiteLLM, Triton, Ray, SGLang. Nearly a quarter stem from unauthenticated API endpoints or SSRF.
  • Model security advisories — 106. Prompt injection, system-prompt exfiltration, guardrail bypass, training-data poisoning.
  • ML frameworks & hubs — 99; frontier models — 97; MLOps & experiment tracking — 39; vector databases — 19.

The frontier-model row is the most quietly damning, because it is a list of things this site has covered as individual incidents and GTIG has now aggregated into a category: RCE via unvalidated CLI shell interpolation and implicit execution of untrusted workspace configs, sandbox escape via Git worktree directory confusion and memory-tool symlink traversal, and covert exfiltration via prompt-injection-induced Markdown image rendering and permissive fetch allowlists. Ninety-seven CVEs in the vendors' own harnesses.

Against 2,076 disclosures, only a handful are confirmed exploited — GTIG names CVE-2026-42271 (LiteLLM command injection in MCP preview endpoints), CVE-2026-5027 (Langflow path-traversal file write) and CVE-2025-3248 (Langflow unauthenticated code injection via exec()). Both Langflow flaws have appeared here before, in our coverage of CVE-2026-5027 under active exploitation. Critically: GTIG has not observed zero-day exploitation of AI infrastructure. Everything confirmed so far is n-day against exposed middleware.

What to take from this

  • Stop reading the disclosure count as a threat signal. One project's CNA policy contributed ~5,000 records with zero observed zero-day exploitation. The meaningful denominators are 141 exploited and 0.23%.
  • Budget patch velocity for n-days, not zero-days. Zero-day exploitation is roughly flat; High-Risk exploitation more than doubled. Your exposure is the interval between a public patch and your deployed patch, and CVE-2026-1731 says that interval is now four days.
  • Treat AI-attributed advisories as a higher-severity intake stream. A 50%-versus-26% RCE rate means the provenance line in an advisory carries triage information, even though no CVE field records it.
  • If you run agent orchestration, you own the largest AI attack surface there is. 782 CVEs and +347% growth in one layer. Dynamic code-execution nodes reachable from workflow JSON or prompt content are the recurring root cause.
  • Do not treat the absence of AI-stack zero-days as comfort. It means adversaries have not needed them. Unauthenticated endpoints on inference and orchestration services are being taken with published n-days.

The report closes on a claim worth testing rather than accepting: that if pre-release AI code review becomes standard practice, public disclosure growth could eventually slow. That is a reasonable hypothesis about supply. It says nothing about the demand side — the four-day weaponisation window — which is the number that will hurt first.

Sources: