Skip to main content
CacheBlend is a technique introduced by the Tensormesh team that reuses precomputed KV caches beyond prefix boundaries. Rather than requiring an exact prefix match, it reuses any contiguous block of previously computed tokens regardless of position, then selectively recomputes a small subset of tokens to partially update each reused block. CacheBlend is most useful when:
  • prompt reuse is non-prefix — shared blocks sit after a per-request preamble, system prompt, or other documents
  • you serve RAG workloads where the same documents are retrieved in varying order across requests
Two kinds of CacheBlend set up exists today:
  • Use Aggregated mode when your workload is prefill heavy
  • Use PD disaggregation when your workload has more decoding workload
Either mode can add an L2 external storage tier on top of engine’s CPU offloading.

Prerequisites

  • cert-manager — any recent release; the mutating webhook’s serving certificate is issued through it:
  • Access Token - The cacheblend-plugin image is distributed through a private registry. To pull it, request an access token from https://www.tensormesh.ai/contact.

Supported Versions

Supported Models

Aggregated mode

Install Chart

Install the chart with CacheBlend enabled. This creates the CacheBlendEngine — a per-node DaemonSet — in the tensormesh-platform namespace:
my-values.yaml

Opting a vLLM pod in

CacheBlend plugin attaches to vLLM pods you deploy — the operator does not create them. First create the workload namespace (PSS-privileged) and the plugin pull secret — the private init container is pulled in the pod’s namespace, so the secret must live here:
Then add a label + annotation to your pod template and launch vLLM args-only (a command: override makes the webhook skip injection):
Everything else (image, model, resources, replicas) stays your normal vLLM config. At admission the webhook injects the plugin init container and PYTHONPATH, the IPC path to the node-local engine (a /dev/shm mount — or hostIPC when the engine’s spec.hostIPC is true), the CacheBlend vLLM flags, and the --kv-transfer-config (CBKVConnector) pointing at that engine; blend tunables (cb.check_layer, cb.recomp_ratio) come from the engine config.

Verification

Engine reconciled
The <engine>-connection ConfigMap is the reconcile proof — and the gate the webhook reads. If it is missing, injection silently fail-opens (the pod runs without CacheBlend). Webhook wired
A non-empty caBundle means cert-manager issued the serving cert and the webhook can be called. Injection happened After creating an opt-in vLLM pod in the workload namespace:
Non-prefix blend actually engaged Send two requests that share a block after different-length preambles, then check the engine logs. A shifted retrieve means a chunk was reused at a different position than it was cached — i.e. a non-prefix blend (select the engine by instance to skip any regular LMCacheEngine in the same ns):

PD disaggregation

Install Chart

my-values.yaml

Namespace and pull secret

Create the cacheblend-workload namespace and the cacheblend-plugin-pull secret first, exactly as in Opting a vLLM pod in.

Deploy prefiller, decoder, router

Deploy your vLLM prefiller, decoder, and router together. The lmcache.ai/pd-role annotation declares each vLLM Deployment’s role — prefiller or decoder — and the webhook injects the following.
  • Prefiller — the CacheBlend plugin, the CacheBlend vLLM flags, and --kv-transfer-config set to the MultiConnector (NixlConnector + CBKVConnector) — so it blends locally and produces KV for the decoder.
  • Decoder — --kv-transfer-config only, set to a bare NixlConnector (kv_consumer). No plugin, no init container, no IPC wiring: it receives finished KV from the prefiller over NIXL.
  • Both — VLLM_NIXL_SIDE_CHANNEL_HOST (the pod IP) and VLLM_NIXL_SIDE_CHANNEL_PORT (spec.pd.nixlSideChannelPort, unless the pod already sets one), and the lmcache.ai/cacheblend-injected annotation.

Verification

PD ConfigMap shape With spec.pd set, the engine’s connection ConfigMap carries three role keys. The decoder key must not contain CBKVConnector — the decoder never blends:
Each role got its own injection
Expected: injected present on both; init=cb-plugin-stage on the prefiller only (empty on the decoder); nixl_port=5557 on the prefiller (its pre-set value) and 5558 on the decoder (injected from spec.pd.nixlSideChannelPort). A lmcache.ai/cacheblend-skip-reason annotation on either pod means the webhook skipped it — see Common mistakes. Requests flow through both roles
A completion back means the router split the request, the prefiller produced KV, and the decoder received it over NIXL and generated from it. Non-prefix blend engaged on the prefiller Same check as aggregated mode — blending happens only on the prefill leg, so only the prefiller’s traffic can produce shifted retrieves:

L2 external storage offloading

The blend engine’s L1 is CPU DRAM, bounded by l1.sizeGB. Adding an L2 adapter lets chunks that fall out of L1 be served from external storage instead of recomputed; the CacheBlend lookup reads L2 through the engine’s L2→L1 prefetch, so nothing changes on the vLLM side. For example for file system offloading, add the following l2.yaml to cacheBlend.spec in your helm values:
l2.yaml

Verification

Adapter engaged
Chunks written to and read back from L2 The engine exposes its counters on the same port as /status (8080 by default):
lmcache_mp_l2_store_completed_objects_chunks_total > 0 means chunks were stored to L2; lmcache_mp_l2_prefetch_load_completed_chunks_total > 0 means chunks were read back from L2 into L1 for a lookup. With storePolicy: default the second only rises after L1 evictions — send more distinct long prompts than L1 holds, then repeat an early one.

Helm values reference

Common mistakes

  • Running the opt-in vLLM pod in the operator (release) namespace — the webhook excludes it, so injection never happens
  • A bare engine ref (tensormesh-cacheblend) instead of <release-ns>/tensormesh-cacheblend — the webhook looks in the pod’s own namespace and finds no engine
  • Setting a command: on the vLLM container — the webhook skips injection (appended args would never reach vllm serve)
  • Workload namespace not PSS-privileged — Pod Security rejects the injected hostIPC
  • The cacheblend-plugin-pull secret missing from the workload namespace — the injected init container ImagePullBackOffs (the secret must live where the pod runs, not in the operator namespace)
  • Mismatched cacheblend-plugin and vllm-openai images — vLLM crashes at startup with an ImportError; pin a matched pair
  • webhook.enabled: false or no cert-manager — the engine deploys but nothing injects the plugin