Nine CVEs, Three Dead Fix PRs — vLLM’s October Batch and the Bookkeeping Problem Behind AI-Found Bugs

On 5 October 2026, nine consecutive CVE identifiers landed against vLLM: CVE-2026-105752 through CVE-2026-105760. They are a single batch in every sense. All nine were published to NVD on the same day and still sit at status Received. All nine carry the same credit line: “Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)”, with the explicit note that the issue “was discovered using GPT-5.5-Cyber.”

The flaws themselves are mostly unglamorous — eight denial-of-service and isolation bugs in the range CVSS 3.1 to 6.5, no remote code execution, nothing exploited. The interesting part is the paperwork. We pulled every advisory from the GitHub API and then read vLLM’s source at seven release tags to check the claims. The fixes are real and already shipped. But three of the nine advisories point the reader at a pull request that was never merged, and the machine-readable version ranges contain typos that would defeat any scanner parsing them literally.

The batch, and when it was actually fixed

vLLM cut v0.30.0 on 22 September 2026 and v0.31.0 on 5 October. Seven of the nine advisories name 0.30.0 as the patched version; one names 0.28.0 (26 August). In other words, by the time these CVE records appeared, every issue in the batch had been fixed for roughly two weeks to six weeks. The repository advisories themselves went public earlier — 23 and 28 September — and were edited on 5 October, minutes before the CVEs issued. This is the healthy ordering: ship, then publish.

  • CVE-2026-105752 (CVSS 3.1, CWE-524/200) — on the GPT-OSS “Harmony” path, a tool continuation re-submitted the next turn without the caller’s cache_salt, so continuation prefixes landed in the global unsalted namespace and a co-tenant could read exact cached_tokens_per_turn counts. Fixed in 0.30.0.
  • CVE-2026-105753 (6.5, CWE-617) — the mirrored multimodal cache desyncs when a request is rejected after the frontend hashes the media but before the engine receives it; a later request with the same media hash trips a receiver assertion. Fixed in 0.28.0.
  • CVE-2026-105754 (6.5) — scale-out disaggregated multimodal transport trusted caller-supplied features across five sites. Fixed in 0.30.0.
  • CVE-2026-105755 (4.2, CWE-639) — late-interaction scoring cached query embeddings under a caller-controlled request id, breaking cross-request integrity on /score and /rerank. Fixed in 0.30.0.
  • CVE-2026-105756 (6.5) — loose cache_salt validation let one request kill EngineCore on LMCache-MP deployments. Fixed in 0.30.0.
  • CVE-2026-105757 (6.5, CWE-755) — structured-output request errors escaped the request boundary and terminated the shared engine. Fixed in 0.30.0.
  • CVE-2026-105758 (5.3, CWE-770) — Qwen2-VL and Qwen3-VL video samplers bound on a request-controlled max_frames. Fixed in 0.30.0.
  • CVE-2026-105759 (5.9, CWE-400) — attacker-controlled HTTP method tokens produced unbounded Prometheus label cardinality in the Rust frontend’s metrics middleware, unauthenticated. Fixed in 0.30.0.
  • CVE-2026-105760 (5.3, CWE-400) — GLMGA video sampling permitted request-driven CPU and memory exhaustion. Fixed in 0.30.0.

Three advisories cite a fix that does not exist

Each advisory ends with a line of the form “a fix for this issue is proposed in a public pull request” and a link. We checked all seven linked PRs against the GitHub pull-request API. Four were merged in early September — #51445 (3 September), #51444 (5 September), #51898 (6 September), #51450 (9 September) — and reached users in 0.30.0 as the records say.

The other three did not:

  • PR #51818, “Preserve cache isolation across Harmony tool turns” (the fix cited by CVE-2026-105752) — closed, never merged.
  • PR #51897, “Roll back mirrored multimodal cache entries after rejection” (CVE-2026-105753) — closed, never merged.
  • PR #51969, “[Security] Enforce server-side num_frames ceiling in VideoMediaIO merge” (referenced by CVE-2026-105758) — still open as of this writing. To the advisory’s credit, its text says so plainly and spends several paragraphs explaining that the open PR’s clamp does not reach the Qwen samplers at all.

So are those three fixed or not? We went to the source rather than the metadata.

What the code says

We fetched the relevant files from the vLLM repository at tags v0.25.1, v0.26.0, v0.27.0, v0.28.0, v0.29.0, v0.30.0 and v0.31.0.

CVE-2026-105752 (Harmony salt). Through v0.29.0, vllm/entrypoints/openai/responses/serving.py contains the bare call engine_input = tokens_input(token_ids) on the tool-continuation path, exactly as the advisory describes. In v0.30.0 that code is gone. The continuation now reads a salt once at the top of _generate_with_builtin_tools — cache_salt = cast(str | None, engine_input.get("cache_salt")) — and threads it through render_responses_harmony_messages(context.messages, cache_salt=cache_salt, …) on every turn. The leak is closed in 0.30.0 and remains closed in 0.31.0. The advisory’s fixed version is right; the pull request it credits is not the one that fixed it. The call site was restructured by a renderer refactor rather than by the proposed patch.

CVE-2026-105753 (multimodal cache desync). Through v0.27.0, MultiModalReceiverCache.get_and_update_item in vllm/multimodal/cache.py ends with the advisory’s exact line: assert mm_item is not None, f"Expected a cached item for {mm_hash=}". In v0.28.0 that assertion is replaced with a typed, recoverable error:

# vllm/multimodal/cache.py (v0.28.0)
        # No data and not cached here: P0 sent data=None trusting its shadow, but
        # the P0/P1 caches have drifted. Raise a retryable error (not assert) so the
        # engine can have P0 drop the stale entry and the client resend the data.
        if mm_item is None:
            raise MultiModalCacheMissError([mm_hash])

That change landed in commit 3962042304 on 12 August, titled “[Bugfix][V1][Multimodal] Recover from P0/P1 processor cache drift” — a bugfix PR, not the security PR the advisory links. Again: correct fixed version, wrong provenance. Note also that the fix is not the rollback the reporter proposed. The maintainers chose recovery over prevention: the desync can still occur, it just no longer kills the engine.

CVE-2026-105758 (Qwen video frames). In v0.29.0, vllm/multimodal/video.py contains no _MAX_FRAMES class constant on the Qwen backends and no clamp on the request-supplied max_frames. In v0.30.0 there are three clamped call sites of the form min(kwargs.get("max_frames", cls._MAX_FRAMES), cls._MAX_FRAMES), with _MAX_FRAMES = 768 on the two Qwen samplers and 640 on GLMGA. The fix arrived in commit ea723c81c3, “[Security] Cap Qwen-VL video sampling knobs”, on 14 September — while PR #51969 sat open and still sits open. The ceiling the advisory warns is unreachable was separately implemented, elsewhere, under a different number.

The verdict is reassuring for operators and unflattering for the data: all nine are genuinely fixed at the versions named, and in three cases the advisory’s own fix pointer is a dead end that a reader checking the citation would mistake for an unpatched bug.

The version strings are broken

Pulled straight from GitHub’s repository advisory API, the affected and patched ranges for this batch include:

  • CVE-2026-105755 — affected < 030.0
  • CVE-2026-105756 — patched >= 30.0.0
  • CVE-2026-105759 — affected < v0.30.0
  • CVE-2026-105758 — affected >= 0.24.0, with no upper bound at all in the repository record

None of these is a version vLLM has ever published. 030.0 and 30.0.0 are typos for 0.30.0; the stray v prefix breaks PEP 440 comparison; and an open-ended “≥ 0.24.0” marks every future release vulnerable forever.

The saving grace is normalisation downstream. OSV ingested all nine and rewrote every one of them to the same clean event pair — {"introduced": "0"}, {"fixed": "0.30.0"} (or 0.28.0 for CVE-2026-105753) — and GitHub’s global advisory endpoint also displays < 0.30.0 where the repository record says < 030.0. Anything consuming OSV or the global database gets usable data. Anything consuming the repository advisory API directly, or a feed that rehosts those raw strings, does not. We saw the same class of failure last week in Langflow’s duplicate 9.9 pair that disagreed with itself about the fixed version; the mode here is less dangerous but the root cause — hand-typed version strings in a security record — is identical.

Eight of eighty-six

The attribution is the part worth sitting with. vLLM’s repository currently carries 86 published advisories. Eight of them — the seven CVE-bearing records above plus two with no CVE assigned (GHSA-p6g9-7v3x-m8mv on remote media materialised before size limits, GHSA-m52c-39rh-f3gp on an IAMF audio upload reaching a PyAV/FFmpeg native heap overflow) — carry the Patch the Planet credit, all published within a 16-day window in September. That is a meaningful fraction of one project’s lifetime advisory count, produced by one AI-assisted program in under three weeks.

Read the write-ups and the quality claim holds up on its own terms. The CVE-2026-105758 advisory does something most human reports skip: it applies the existing open fix PR locally, re-runs both paths through the real merge layer, and presents a four-row table showing num_frames clamped to 32 while max_frames: 1e9 survives and drives peak RSS from 706 MiB to 2,994 MiB. The CVE-2026-105752 write-up distinguishes itself from the prior CVE-2025-46570 prefix-cache oracle by naming the specific sink that the earlier fix did not reach. These are differential-analysis reports — “the control exists, here is the path that bypasses it” — which is the shape of finding that regression tests miss and that GTIG’s disclosure-trend data says is now arriving faster than projects can triage.

That volume is also the likeliest explanation for the bookkeeping drift. When a maintainer receives nine reports with nine attached patches in three weeks, fixes the underlying problems their own way, and then publishes the advisories a month later from the original report text, the citation decays silently. Nobody lied. The record simply stopped tracking the repository.

What to do

  • Upgrade to 0.30.0 or later. One release clears eight of the nine; 0.31.0 shipped 5 October and is the current stable. If you are on 0.29.x you are exposed to seven of them.
  • Do not match on the raw repository version strings. If your pipeline reads GitHub repository advisories directly, add a sanity check that every bound parses as a release that exists. >= 30.0.0 will silently mark you unaffected; >= 0.24.0 with no ceiling will mark you affected forever. Prefer OSV or the global advisory database, both of which normalised this batch correctly.
  • Treat “proposed fix: PR #N” as a lead, not evidence. Three of the seven cited PRs in this batch are closed-unmerged or still open. Verify against the release tag, not the pull request.
  • Authenticate /tokenize and /invocations at the proxy. The CVE-2026-105758 write-up notes that even with --api-key set, vLLM’s guarded prefixes cover /v1, /v2, /inference and /cohere — so /tokenize, which performs full media ingestion, stays open. A default vllm serve has no authentication at all.
  • If you rely on cache_salt for tenant isolation, audit the Harmony path specifically. The documented control worked on turn one and silently lapsed on every tool continuation for at least five releases. Salting is only an isolation boundary if every code path carries it.
  • Expect more of these. Eight advisories from one AI-assisted program in sixteen days against one project is the new baseline, and the ones that follow will keep being medium-severity resource and isolation bugs rather than headline RCEs.

Verification note: advisory metadata, CVSS vectors, CWE assignments, version ranges, credits and full descriptions come from the GitHub repository security-advisory API for vllm-project/vllm and from GitHub’s global advisory endpoint; normalised ranges come from the OSV API. NVD status (Received, published 5 October 2026) comes from the NVD CVE 2.0 API. Pull-request states and merge timestamps for #51444, #51445, #51450, #51818, #51897, #51898 and #51969, release dates for v0.22.1 through v0.31.0, and commit dates for 3962042304 and ea723c81c3 come from the GitHub REST API. The source analysis is ours: we fetched vllm/entrypoints/openai/responses/serving.py, vllm/multimodal/cache.py and vllm/multimodal/video.py at tags v0.25.1 through v0.31.0 and confirmed the presence of the bare tokens_input(token_ids) continuation through v0.29.0 and its replacement by a salt-carrying render call in v0.30.0; the assert mm_item is not None line through v0.27.0 and its replacement by MultiModalCacheMissError in v0.28.0; and the absence of any Qwen _MAX_FRAMES clamp through v0.29.0 with three clamped call sites appearing in v0.30.0. We ran no exploit, tested no deployment, and did not contact the maintainers or the reporting program before publication.

Sources: