Ten CVEs, Zero Patches: LightLLM Keeps Shipping allow_pickle on 0.0.0.0

On 29 and 30 September 2026 VulnCheck published four more CVEs against ModelTC's LightLLM, a Python LLM inference and serving framework with roughly 4,300 GitHub stars. Three of the four are rated CVSS 4.0 9.3 / CVSS 3.1 9.8 unauthenticated remote code execution. That brings the public total to ten LightLLM CVEs in NVD, six of them published in the seventeen days between 14 and 30 September.

The number that matters is a different one. Every corresponding bug report on the LightLLM issue tracker is still open — including issue #1213 for CVE-2026-26220, filed 15 February 2026. There is no 1.2.1. The last tagged release is v1.2.0 from 10 August 2026, and the vulnerable code is present on main as of this writing.

One root cause, six ports

Strip the CVE numbers away and the September batch is a single architectural decision replicated across subsystems: internal control planes built on RPyC with allow_pickle: True, bound to all interfaces, with no authentication. RPyC with pickle enabled is a remote eval by design; the only thing standing between a network peer and the service account is the assumption that nobody can reach the port.

  • CVE-2026-103395 (9.3, CWE-502) — --run_mode visual_only. The visual server's ThreadedServer is constructed with no hostname argument, so it binds 0.0.0.0; exposed_remote_infer_images passes its argument through obtain(), which unpickles it. An object with a __reduce__ method executes before the code ever notices the type is wrong.
  • CVE-2026-103041 (9.3) — the multimodal embed-cache RPyC service, same configuration, different process and port.
  • CVE-2026-103040 (9.3) — the router profiler service, reachable when started with --enable_profiling.
  • CVE-2026-103042 (8.7, CWE-770) — the NCCL KV-transfer control channel's exposed_set_value stores unbounded key/value pairs, so an unauthenticated peer exhausts worker memory and takes the node down.
  • CVE-2026-103270 (8.7, CWE-306) — the reinforcement-learning control router is mounted unconditionally on the public HTTP API. /pause_generation, /abort_request, /flush_cache and /init_weights_update_group take no auth on --enable_rl deployments.
  • CVE-2026-103243 (6.9, CWE-918) — unvalidated image_url and audio_url parameters give unauthenticated SSRF through the multimodal endpoints.

Earlier entries in the same family cover the Config Server's /visual_register WebSocket feeding client frames straight to pickle.loads() (CVE-2026-90919), the unauthenticated /pd_register endpoint that lets an attacker register a node and receive the user prompts routed to it (CVE-2026-93839), and the NCCL PD RPyC control channel (CVE-2026-96560). The oldest, CVE-2026-26220, described the same pickle-over-WebSocket pattern in PD mode back in February.

The reporters did the hard part

These are not drive-by scanner findings. The GitHub issues read like vendor advisories: exact file and line references, the rpyc_config dict quoted verbatim, an explanation of why ThreadedServer without a hostname binds every interface, reproduction at a named commit, a proof-of-concept, and an explicit note of which other CVEs share the weakness class and why fixing those does not fix this one. VulnCheck credits Mingkai Yu and Jiajia Liu on the visual_only advisory.

The project's response to that work has been silence. Nine of the eleven security issues we reviewed have zero comments. The three that drew replies drew them from other community members offering to pick the work up — on issue #1563 a contributor volunteered on 10 September and was still scoping the fix on 13 September; the issue remains open. The repository has no SECURITY.md and no GitHub Security Advisories, which is why none of this reaches you through pip-audit or a dependency bot: there is no package-level advisory to match against, and the PyPI name lightllm is an unrelated 0.0.1 placeholder from 2024. Teams install LightLLM from source or a container image, so the standard supply-chain tripwires never fire. We saw the same blind spot in two MCP command injections that will never show up in npm audit.

Why "it's on an internal network" is not the answer

The predictable defence of every one of these bugs is that the ports are internal. That argument has aged badly. Inference clusters are multi-tenant, their control planes span nodes, and the RPyC ports here are bound to 0.0.0.0 rather than a loopback or a cluster interface — so the blast radius is whatever your Kubernetes network policy happens to be, which in most clusters is nothing. A single compromised pod, a misconfigured NodePort, a sidecar with egress, or an SSRF anywhere in the estate turns a "trusted" port into a 9.8.

And the consequences are not just code execution. The /pd_register flaw discloses the prompts of real users to an attacker-registered node — the SolarWinds ARM pattern, where an internal management channel assumed trust it never verified. Pickle on an inference control plane has been a recurring theme this year, from Unit 42's cross-tenant Vertex AI RCE to Hugging Face LeRobot. The difference here is the patch gap: those got fixed.

What to do

  • Treat every LightLLM deployment as unpatched. There is no fixed version. v1.2.0 is the latest tag and the vulnerable configurations are on main. If you are waiting for an upgrade path, you are waiting indefinitely.
  • Inventory the ports, not the version. Enumerate what each node actually listens on — visual RPyC, embed cache, profiler, NCCL transfer, config server, PD master WebSocket, HTTP API. Several only appear under specific flags (--run_mode visual_only, --enable_profiling, --pd_trans_mode nccl, --enable_rl), so the attack surface varies per deployment mode.
  • Enforce the isolation the code does not. Default-deny NetworkPolicy or security groups around every LightLLM pod, with an explicit allowlist for the peers that genuinely need the control plane. Bind the HTTP API behind an authenticating proxy; do not expose the RPyC ports at all.
  • Confine the service account. Non-root, read-only root filesystem, seccomp, no cloud-credential access from inference workers. Unauthenticated pickle means RCE as whatever the process is; the only lever you still control is what that identity can reach. This is the same containment argument as the AF_UNIX container escape — assume the boundary fails and cost the attacker the next step.
  • Watch the patch gap, not the CVE count. A project accumulating critical CVEs with an unresponsive maintainer is a supply-chain decision, not a patching one. If LightLLM is load-bearing for you, budget for either carrying your own patches or migrating. We flagged a related failure mode in 72 OpenClaw CVEs landing in one day for bugs patched months earlier; this is the mirror image — advisories landing for bugs never patched at all.

Sources: