LMCache — Remote Code Execution in Multiprocess Mode via Insecure Deserialization (CVE-2026-105192)

Publication date: October 7, 2026
Category: Vulnerability / Artificial Intelligence

Introduction

A critical vulnerability discovered in LMCache, an open-source software designed to accelerate large language model (LLM) serving frameworks such as vLLM, allows an unauthenticated attacker to execute arbitrary commands on the cache server without logging in. The flaw resides within the component’s multiprocess mode, where the ZeroMQ messaging transport handles inter-process messages insecurely. Security researchers at JFrog disclosed the flaw, noting that official container images run the vulnerable process as root, severely amplifying the risk when exposed to routable networks.

What is LMCache and Multiprocess Mode? (General Analysis)

LMCache serves as a distributed storage and key-value (KV) cache system tailored to optimize inference workloads for large language models. In advanced setups or multi-node environments, LMCache operates in a multiprocess mode (also known as distributed mode), running as a standalone server where worker processes connect over ZeroMQ to share cache blocks efficiently.

The critical security issues are tracked under the following officially verified NVD identifiers:

  • CVE-2026-105192: Official CVSS v3.1 score of 9.8 (CRITICAL) with vector CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H. Associated CWE categories are CWE-306 (Missing Authentication for Critical Function) and CWE-502 (Deserialization of Untrusted Data).
  • CVE-2026-105756: Official CVSS v3.1 score of 6.5 (MEDIUM) with vector CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H. Associated CWE categories are CWE-20 (Improper Input Validation) and CWE-248 (Uncaught Exception).

How Does It Work? (Technical Analysis)

A breakdown of the exploitation mechanism reveals a classic architectural design flaw involving network transport protocols and Python object serialization:

  • Initial Entry and ZeroMQ Transport Flow: The LMCache multiprocess server opens an unauthenticated ZeroMQ ROUTER socket (defaulting to TCP port 5555) to allow worker processes to register and share cached data blocks.
  • Insecure Pickle Deserialization: Messages transmitted across this socket use msgpack, but extension code 1 passes data directly into DeviceIPCWrapper.Deserialize. This function invokes pickle.loads while the server is still decoding request arguments, prior to any handler or validation checks.
  • Code Execution as Root: A single crafted message sent by an unauthenticated attacker through a ZeroMQ DEALER socket triggers the deserialization of malicious Python objects, executing code with the privileges of the user running the LMCache process. Because official container images run this process as root, complete node compromise occurs instantly.
  • Related vLLM Flaw (CVE-2026-105756): In vLLM versions prior to 0.30.0, OpenAI-compatible request models accepted malformed cache_salt values without enforcing length or character restrictions. When interacting with LMCache, an invalid salt triggered an uncaught ValueError during scheduler cache lookup, causing EngineCore to crash and creating a denial-of-service (DoS) condition.

Affected Systems / Environments

The impact is concentrated across AI deployment infrastructures integrating LMCache into distributed environments:

CVE Category (CWE) Impact CVSS Vector (summary)
CVE-2026-105192 CWE-306, CWE-502 RCE (Remote Code Execution) 9.8 (Critical) Network / No Auth / No Privileges
CVE-2026-105756 CWE-20, CWE-248 Denial of Service (DoS) 6.5 (Medium) Network / Low Auth / High Availability Impact
  • Affected Software (CVE-2026-105192): LMCache from version 0.3.9 (released October 2025) through 0.5.5, including 0.5.6 release candidates and the active development branch.
  • Affected Software (CVE-2026-105756): vLLM prior to version 0.30.0 (patched in v0.30.0).
  • Exposure Condition: By default, the multiprocess server listens exclusively on localhost. However, multi-node deployments or standard deployment examples (such as Kubernetes DaemonSets) explicitly configure the --host parameter to bind on a routable network address, exposing the port externally.

Mitigation and Detection

Remediation

  • Network Restriction and Binding: Until an official patched release becomes available, operators must avoid assigning routable or public IP addresses to the LMCache multiprocess server. Ensure the service binds strictly to localhost or trusted internal cluster networks.
  • Perimeter Firewall Segmentation: Apply strict packet filtering rules (firewalls / Security Groups) on TCP port 5555 to limit access to the transport socket exclusively to legitimate internal nodes within the inference cluster.
  • vLLM Upgrade: Update the vLLM serving engine to version 0.30.0 or later to remediate the related denial-of-service vulnerability tied to cache_salt.

Detection

  • Network Connection Auditing: Audit network interfaces across AI clusters to detect unauthorized inbound connections targeting ZeroMQ transport ports (default 5555).
  • Container Log Analysis: Monitor logs for unexpected EngineCore process terminations or uncaught ValueError exceptions linked to cache parameters.

Defensive Intelligence Note: Exposing internal asynchronous messaging services (such as ZeroMQ) to untrusted networks combined with implicit object deserialization routines introduces a critical intrusion vector that bypasses traditional security perimeters in AI applications.

Wrapping Up

The critical vulnerability CVE-2026-105192 in LMCache highlights the inherent risks of employing insecure serialization mechanisms (pickle) within internal transport protocols when accidentally exposed via routable network configurations. With a CVSS score of 9.8 and code execution under root privileges in default containers, the temporary absence of an official patch demands that operations teams enforce strict network isolation and perimeter segmentation to safeguard LLM inference environments.

References