The Monitoring Port Was the Attack Surface: NVIDIA CVE-2026-47483 and 12,000 Exposed GPUs

Research published on 9 October 2026 by Lava Security pairs a two-and-a-half-month-old NVIDIA bulletin with an internet-wide census, and the combination is worse than either half alone. CVE-2026-47483 is an unauthenticated resource-exhaustion flaw in NVIDIA DCGM Exporter: concurrent profiling requests against the /debug/pprof endpoints trigger uncontrolled resource consumption, crashing the GPU monitoring service and potentially disclosing information. NVIDIA rates it CVSS 3.1 8.2 High (CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:H), CWE-770. Lava researcher Michael Katchinskiy and his team found more than 2,000 servers exposing the exporter to the internet, reporting more than 12,000 unique GPUs with an estimated hardware value around $100 million — and reporting describes hundreds of those hosts as sitting on builds vulnerable to the flaw.

The bulletin is not new. NVIDIA published it as Security Bulletin 5857, “NVIDIA DCGM Exporter — July 2026,” on 28 July 2026, and the CVE record has carried that date since. What is new is the exposure data and the operational reading of it: the telemetry endpoint that makes the fleet observable is itself reachable, unauthenticated, over plaintext HTTP — and on vulnerable builds it is also a remotely triggerable kill switch for visibility into the training and inference fleet.

What the exporter leaks, and what the flaw does

DCGM Exporter reads data directly from the GPUs on a host — temperature, utilisation, memory usage, power draw, error events — and publishes it over HTTP, typically on port 9400, for Prometheus to scrape. The feed includes the exact GPU model and a unique UUID per GPU. As Katchinskiy put it, that is enough for an attacker to see what hardware a server runs, how heavily it is used, and whether it is throwing errors — a hardware inventory, a utilisation profile, and a health signal, served to anyone who connects.

The vulnerability sits on the same port. DCGM Exporter ships Go pprof debug endpoints, and on affected builds those endpoints accept concurrent profiling requests with no authentication. Profiling is expensive by design — it samples the running process — so an unauthenticated remote party can simply ask for enough of it at once to exhaust the exporter's resources. NVIDIA's own description names the two consequences: denial of service and information disclosure. The C:L in the vector is doing quiet work here: pprof output is process introspection, so the same endpoint that crashes the service can also leak runtime detail on the way down.

Affected range, per the CSAF bulletin: DCGM Exporter 0.0 through 4.8.2, fixed in 4.8.2; the bundled DCGM 0.0 through 4.5.2, fixed in 4.5.3. Remediation is a pull from the NVIDIA/dcgm-exporter repository. NVD still carries the record as Awaiting Analysis with only NVIDIA's vendor score — no independent NVD analysis — and CISA's SSVC entry marks exploitation as none but the flaw as automatable: yes. That pairing is the honest summary: no exploitation reported, and nothing about the bug would slow an attacker down.

The census, and its date problem

Lava's numbers come from four scans run between March and May 2026 — before the 28 July bulletin existed. That timing cuts both ways and should be stated plainly. On the one hand, every host in the census was necessarily unpatched at scan time, so the “hundreds vulnerable” figure describes the population Lava could fingerprint, not a post-patch residual. On the other, the exposure finding does not expire with the patch: the census measures how many GPU operators publish unauthenticated telemetry to the internet at all, and a patched exporter on port 9400 still answers the inventory question for anyone who asks.

The figures, as reported: more than 2,000 servers exposing the exporter, more than 12,000 unique GPUs, roughly $100 million in hardware value. One outlet's counting puts it at 2,100 servers and 12,096 GPUs. The spread between outlets is a fingerprinting-methodology difference, not a contradiction — and either way it is a four-to-five-digit population of AI compute with its monitoring plane on the public internet.

Katchinskiy's conclusion is aimed at exactly this layer: as AI infrastructure scales, organisations need to protect every layer of the stack and know precisely what they are responsible for versus what their provider is. The shared-responsibility line for GPU cloud capacity is genuinely murky — the provider owns the building, the customer owns the exporter config — and port 9400 falls squarely in the gap.

Why a monitoring crash matters to the workload

It is tempting to file this as “just a monitoring DoS” and move on. Three reasons not to. First, a blind fleet is a fleet you cannot schedule: autoscalers, thermal throttles, and failure detectors all consume the telemetry this service publishes, and a crashed exporter during a long training run removes the signal you would use to notice the run going wrong. Second, the bulletin explicitly warns the flaw can affect workloads running on the same host — resource exhaustion does not respect the boundary between the monitoring sidecar and the job it watches. Third, the recon value persists after patching: model, UUID, utilisation, and error state are exactly the inputs an attacker wants before a targeted campaign against AI infrastructure, and they remain free for the asking wherever the exporter stays internet-facing.

Current status: the CVE is not in the CISA Known Exploited Vulnerabilities catalog — we checked the published CSV on 10 October (1,739 entries) and it is absent. NVIDIA's bulletin is final at version 1.0.0 with no revision history, and Katchinskiy is the sole credited reporter.

What to do

  • Upgrade DCGM Exporter to 4.8.2 (DCGM to 4.5.3). The bulletin names both fixed builds and they are not interchangeable — check which component each host runs before scheduling.
  • Take the exporter off the internet regardless of version. Bind it to localhost, scrape through an agent or sidecar, and confirm port 9400 is not reachable externally. The patch closes the crash; only network posture closes the inventory leak.
  • Treat /debug/pprof as a finding wherever it appears. Go debug endpoints on unauthenticated HTTP are a recurring shape — cheap to leave on, expensive when discovered. Audit other exporters and sidecars in the fleet for the same pattern.
  • Watch the fleet's blind spots, not just its versions. If the exporter can be crashed remotely, monitoring absence is itself a signal: alert on scrape failures from Prometheus as a security event, not just an operational one, until the fleet is patched and unexposed.

Verification note: the CVE description, CWE-770 classification, CVSS 3.1 8.2 score and vector, affected and fixed version ranges (Exporter 0.0–4.8.2 → 4.8.2; DCGM 0.0–4.5.2 → 4.5.3), 28 July 2026 release date, and the Michael Katchinskiy acknowledgement were read directly from NVIDIA Security Bulletin 5857 in CSAF JSON from NVIDIA's product-security repository. The NVD record (published 2026-07-28, status Awaiting Analysis, NVIDIA vendor score only, SSVC exploitation:none / automatable:yes) was retrieved from the NVD 2.0 API on 10 October 2026. The exposure figures (2,000+ servers, 12,000+ GPUs, ~$100M, four scans March–May 2026, port 9400, telemetry contents, researcher quotes) are from Help Net Security's 9 October report and are attributed accordingly; the 2,100-server / 12,096-GPU counting is from TechNadu's coverage of the same research. The observation that the scan window predates the bulletin is our reading of the published dates. KEV absence was checked against CISA's published CSV (1,739 entries) on 10 October 2026. We tested nothing and exploited nothing.

Sources: