- prompt reuse is non-prefix — shared blocks sit after a per-request preamble, system prompt, or other documents
- you serve RAG workloads where the same documents are retrieved in varying order across requests
- Use Aggregated mode when your workload is prefill heavy
- Use PD disaggregation when your workload has more decoding workload
Prerequisites
- cert-manager — any recent release; the
mutating webhook’s serving certificate is issued through it:
- Access Token - The
cacheblend-pluginimage is distributed through a private registry. To pull it, request an access token from https://www.tensormesh.ai/contact.
Supported Versions
Supported Models
Aggregated mode
Install Chart
Install the chart with CacheBlend enabled. This creates theCacheBlendEngine — a per-node DaemonSet —
in the tensormesh-platform namespace:
my-values.yaml
Opting a vLLM pod in
CacheBlend plugin attaches to vLLM pods you deploy — the operator does not create them. First create the workload namespace (PSS-privileged) and the plugin pull secret — the private init container is pulled in the pod’s namespace, so the secret must live here:command: override
makes the webhook skip injection):
PYTHONPATH, the IPC path to the node-local engine (a
/dev/shm mount — or hostIPC when the engine’s spec.hostIPC is true), the CacheBlend vLLM flags, and
the --kv-transfer-config (CBKVConnector) pointing at that engine; blend tunables (cb.check_layer,
cb.recomp_ratio) come from the engine config.
Verification
Engine reconciled<engine>-connection ConfigMap is the reconcile proof — and the gate the webhook reads. If it is
missing, injection silently fail-opens (the pod runs without CacheBlend).
Webhook wired
caBundle means cert-manager issued the serving cert and the webhook can be called.
Injection happened
After creating an opt-in vLLM pod in the workload namespace:
shifted retrieve means a chunk was reused at a different position than it was cached — i.e. a
non-prefix blend (select the engine by instance to skip any regular LMCacheEngine in the same ns):
PD disaggregation
Install Chart
my-values.yaml
Namespace and pull secret
Create thecacheblend-workload namespace and the cacheblend-plugin-pull secret first, exactly as in
Opting a vLLM pod in.
Deploy prefiller, decoder, router
Deploy your vLLM prefiller, decoder, and router together. The lmcache.ai/pd-role annotation declares each vLLM Deployment’s role — prefiller or decoder — and the webhook injects the following.- Prefiller — the CacheBlend plugin, the CacheBlend vLLM flags, and
--kv-transfer-configset to theMultiConnector(NixlConnector+CBKVConnector) — so it blends locally and produces KV for the decoder. - Decoder —
--kv-transfer-configonly, set to a bareNixlConnector(kv_consumer). No plugin, no init container, no IPC wiring: it receives finished KV from the prefiller over NIXL. - Both —
VLLM_NIXL_SIDE_CHANNEL_HOST(the pod IP) andVLLM_NIXL_SIDE_CHANNEL_PORT(spec.pd.nixlSideChannelPort, unless the pod already sets one), and thelmcache.ai/cacheblend-injectedannotation.
vllm-cacheblend-pd.yaml — prefiller, decoder, router
vllm-cacheblend-pd.yaml — prefiller, decoder, router
Verification
PD ConfigMap shape Withspec.pd set, the engine’s connection ConfigMap carries three role keys. The decoder key must not
contain CBKVConnector — the decoder never blends:
injected present on both; init=cb-plugin-stage on the prefiller only (empty on the
decoder); nixl_port=5557 on the prefiller (its pre-set value) and 5558 on the decoder (injected from
spec.pd.nixlSideChannelPort). A lmcache.ai/cacheblend-skip-reason annotation on either pod means the
webhook skipped it — see Common mistakes.
Requests flow through both roles
shifted retrieves:
L2 external storage offloading
The blend engine’s L1 is CPU DRAM, bounded byl1.sizeGB. Adding an L2 adapter lets chunks that fall
out of L1 be served from external storage instead of recomputed; the CacheBlend lookup reads L2 through the
engine’s L2→L1 prefetch, so nothing changes on the vLLM side. For example for file system offloading, add the
following l2.yaml to cacheBlend.spec in your helm values:
l2.yaml
Verification
Adapter engaged/status (8080 by default):
lmcache_mp_l2_store_completed_objects_chunks_total > 0 means chunks were stored to L2;
lmcache_mp_l2_prefetch_load_completed_chunks_total > 0 means chunks were read back from L2 into L1
for a lookup. With storePolicy: default the second only rises after L1 evictions — send more distinct
long prompts than L1 holds, then repeat an early one.
Helm values reference
Common mistakes
- Running the opt-in vLLM pod in the operator (release) namespace — the webhook excludes it, so injection never happens
- A bare engine ref (
tensormesh-cacheblend) instead of<release-ns>/tensormesh-cacheblend— the webhook looks in the pod’s own namespace and finds no engine - Setting a
command:on the vLLM container — the webhook skips injection (appended args would never reachvllm serve) - Workload namespace not PSS-privileged — Pod Security rejects the injected
hostIPC - The
cacheblend-plugin-pullsecret missing from the workload namespace — the injected init containerImagePullBackOffs (the secret must live where the pod runs, not in the operator namespace) - Mismatched
cacheblend-pluginandvllm-openaiimages — vLLM crashes at startup with anImportError; pin a matched pair webhook.enabled: falseor no cert-manager — the engine deploys but nothing injects the plugin

