> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tensormesh.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Non-Prefix KV Caching

> Reuse KV cache beyond prefix boundaries with CacheBlend

[CacheBlend](https://arxiv.org/abs/2405.16444) is a technique introduced by the Tensormesh team that
reuses precomputed KV caches beyond prefix boundaries. Rather than requiring an exact prefix match, it
reuses any contiguous block of previously computed tokens regardless of position, then selectively
recomputes a small subset of tokens to partially update each reused block.

CacheBlend is most useful when:

* prompt reuse is **non-prefix** — shared blocks sit after a per-request preamble, system prompt, or other documents
* you serve **RAG** workloads where the same documents are retrieved in varying order across requests

Two kinds of CacheBlend set up exists today:

* Use **Aggregated mode** when your workload is prefill heavy
* Use **PD disaggregation** when your workload has more decoding workload

Either mode can add an [L2 external storage tier](#l2-filesystem-offloading) on top of engine's CPU offloading.

## Prerequisites

* [cert-manager](https://cert-manager.io/docs/installation/) — any recent release; the
  mutating webhook's serving certificate is issued through it:
  ```bash theme={null}
  kubectl apply -f https://github.com/cert-manager/cert-manager/releases/latest/download/cert-manager.yaml
  kubectl -n cert-manager wait --for=condition=Available deploy --all --timeout=180s
  ```
* **Access Token** - The `cacheblend-plugin` image is distributed through a private registry. To pull it, request an access
  token from [https://www.tensormesh.ai/contact](https://www.tensormesh.ai/contact).

### Supported Versions

| TMO helm chart | LMCache Operator | LMCache vLLM | cacheblend-plugin |
| -------------- | ---------------- | ------------ | ----------------- |
| `1.0.0`        | `v0.5.5`         | `v0.5.5`     | `v0.5.5`          |

### Supported Models

| Model family | Models                    |
| ------------ | ------------------------- |
| GLM          | GLM-5.2                   |
| MiniMax      | MiniMax M3, MiniMax M2.5  |
| Gemma        | Gemma3                    |
| gpt-oss      | gpt-oss-20b, gpt-oss-120b |
| llama        | llama 3.1                 |

## Aggregated mode

### Install Chart

Install the chart with CacheBlend enabled. This creates the `CacheBlendEngine` — a per-node DaemonSet —
in the `tensormesh-platform` namespace:

```yaml my-values.yaml theme={null}
# REQUIRED — the mutating webhook injects the plugin (needs cert-manager).
webhook:
  enabled: true

engine:
  enabled: false                           # no LMCacheEngine — CacheBlend-only install
coordinator:
  enabled: false                           

cacheBlend:
  enabled: true
  name: tensormesh-cacheblend        # engine created in the release namespace
  spec:
    l1:
      sizeGB: 200
    image:
      repository: lmcache/vllm-openai
      tag: v0.5.5
      pullPolicy: IfNotPresent
    injection:
      payloadImage:
        repository: artifacts.tensormesh.ai/tensormesh-production/images/cacheblend-plugin
        tag: v0.5.5
        pullPolicy: IfNotPresent
      imagePullSecrets:
        - name: cacheblend-plugin-pull             # must exist in each workload ns
```

```bash theme={null}
helm install tensormesh-platform \
  oci://artifacts.tensormesh.ai/tensormesh-production/charts/tensormesh-platform \
  --version 1.0.0 \
  --namespace tensormesh-platform \
  --create-namespace \
  -f my-values.yaml \
  --wait
```

### Opting a vLLM pod in

CacheBlend plugin attaches to vLLM pods **you deploy** — the operator does not create them. First create the
workload namespace (PSS-privileged) and the plugin pull secret — the private init container is pulled in
the **pod's** namespace, so the secret must live here:

```bash theme={null}
kubectl create ns cacheblend-workload
kubectl label ns cacheblend-workload pod-security.kubernetes.io/enforce=privileged
kubectl -n cacheblend-workload create secret docker-registry cacheblend-plugin-pull \
  --docker-server=artifacts.tensormesh.ai --docker-username=x-access-token --docker-password=<token>
```

Then add a label + annotation to your pod template and launch vLLM **args-only** (a `command:` override
makes the webhook skip injection):

```yaml theme={null}
metadata:
  labels:
    lmcache.ai/cacheblend-inject: "true"                                        # opt in
  annotations:
    lmcache.ai/cacheblend-engine: "tensormesh-platform/tensormesh-cacheblend"   # <release-ns>/<name>
```

Everything else (image, model, resources, replicas) stays your normal vLLM config. At admission the
webhook injects the plugin init container and `PYTHONPATH`, the IPC path to the node-local engine (a
`/dev/shm` mount — or `hostIPC` when the engine's `spec.hostIPC` is `true`), the CacheBlend vLLM flags, and
the `--kv-transfer-config` (`CBKVConnector`) pointing at that engine; blend tunables (`cb.check_layer`,
`cb.recomp_ratio`) come from the engine config.

### Verification

**Engine reconciled**

```bash theme={null}
kubectl -n tensormesh-platform get cacheblendengine tensormesh-cacheblend
kubectl -n tensormesh-platform get configmap tensormesh-cacheblend-connection
```

The `<engine>-connection` ConfigMap is the reconcile proof — and the gate the webhook reads. If it is
missing, injection silently fail-opens (the pod runs without CacheBlend).

**Webhook wired**

```bash theme={null}
kubectl get mutatingwebhookconfiguration tensormesh-platform-mutating-webhook \
  -o jsonpath='{.webhooks[0].clientConfig.caBundle}' | head -c 20; echo
```

A non-empty `caBundle` means cert-manager issued the serving cert and the webhook can be called.

**Injection happened**

After creating an opt-in vLLM pod in the workload namespace:

```bash theme={null}
P=$(kubectl -n cacheblend-workload get pod -l lmcache.ai/cacheblend-inject=true -o name | head -1)
kubectl -n cacheblend-workload get $P -o jsonpath='{.metadata.annotations.lmcache\.ai/cacheblend-injected}{"\n"}'  # present = injected
kubectl -n cacheblend-workload get $P -o jsonpath='{.spec.initContainers[*].name}{"\n"}'   # cb-plugin-stage
kubectl -n cacheblend-workload get $P -o jsonpath='{.spec.containers[0].volumeMounts[?(@.mountPath=="/dev/shm")].name}{"\n"}'   # /dev/shm mount (hostIPC=true instead if the engine sets spec.hostIPC)
```

**Non-prefix blend actually engaged**

Send two requests that share a block **after different-length preambles**, then check the engine logs.
A `shifted` retrieve means a chunk was reused at a different position than it was cached — i.e. a
non-prefix blend (select the engine by `instance` to skip any regular LMCacheEngine in the same ns):

```bash theme={null}
kubectl -n tensormesh-platform logs -l app.kubernetes.io/instance=tensormesh-cacheblend --tail=500 \
  | grep -E "\[match_probe\] .* matches=[1-9]|Retrieved pre-computed for [1-9].* shifted=[1-9]"
```

## PD disaggregation

### Install Chart

```yaml my-values.yaml theme={null}
webhook:
  enabled: true                            # REQUIRED  

engine:
  enabled: false                           # no LMCacheEngine — CacheBlend-only install
coordinator:
  enabled: false                           

cacheBlend:
  enabled: true
  name: tensormesh-cacheblend
  spec:
    privileged: true
    l1:
      sizeGB: 200
    image:
      repository: lmcache/vllm-openai
      tag: v0.5.5
      pullPolicy: IfNotPresent
    injection:
      payloadImage:                        # pulled by PREFILLER pods, in their own namespace
        repository: artifacts.tensormesh.ai/tensormesh-production/images/cacheblend-plugin
        tag: v0.5.5
        pullPolicy: IfNotPresent
      imagePullSecrets:
        - name: cacheblend-plugin-pull
    pd:
      nixlSideChannelPort: 5558            # VLLM_NIXL_SIDE_CHANNEL_PORT on both roles unless the pod pre-sets it
      nixlLoadFailurePolicy: fail          # or `ignore`: fall back to local prefill
```

```bash theme={null}
helm upgrade --install tensormesh-platform \
  oci://artifacts.tensormesh.ai/tensormesh-production/charts/tensormesh-platform \
  --version 1.0.0 \
  --namespace tensormesh-platform \
  -f my-values.yaml \
  --wait
```

### Namespace and pull secret

Create the `cacheblend-workload` namespace and the `cacheblend-plugin-pull` secret first, exactly as in
[Opting a vLLM pod in](#opting-a-vllm-pod-in).

### Deploy prefiller, decoder, router

Deploy your vLLM prefiller, decoder, and router together. The **lmcache.ai/pd-role** annotation declares each vLLM Deployment's role — prefiller or decoder — and the webhook injects the following.

* Prefiller — the CacheBlend plugin, the CacheBlend vLLM flags, and `--kv-transfer-config` set to the `MultiConnector` (`NixlConnector` + `CBKVConnector`) — so it blends locally and produces KV for the decoder.
* Decoder — **`--kv-transfer-config` only**, set to a bare `NixlConnector` (`kv_consumer`). No plugin, no init container, no IPC wiring: it receives finished KV from the prefiller over NIXL.
* Both — `VLLM_NIXL_SIDE_CHANNEL_HOST` (the pod IP) and `VLLM_NIXL_SIDE_CHANNEL_PORT` (`spec.pd.nixlSideChannelPort`, unless the pod already sets one), and the `lmcache.ai/cacheblend-injected` annotation.

<Accordion title="vllm-cacheblend-pd.yaml — prefiller, decoder, router">
  ```yaml theme={null}
  # Prefiller — blends, then hands KV to the decoder over NIXL
  apiVersion: apps/v1
  kind: Deployment
  metadata:
    name: cb-pd-prefiller
    namespace: cacheblend-workload
    labels:
      app: cb-pd-prefiller
  spec:
    replicas: 1
    selector:
      matchLabels:
        app: cb-pd-prefiller
    template:
      metadata:
        labels:
          app: cb-pd-prefiller
          lmcache.ai/cacheblend-inject: "true"
        annotations:
          lmcache.ai/cacheblend-engine: "tensormesh-platform/tensormesh-cacheblend"
          lmcache.ai/pd-role: "prefiller"
      spec:
        runtimeClassName: nvidia
        hostNetwork: true
        dnsPolicy: ClusterFirstWithHostNet
        # No IPC wiring, no --kv-transfer-config: the webhook injects them for this role
        # (a /dev/shm mount, or hostIPC if the engine sets spec.hostIPC).
        containers:
          - name: vllm
            image: lmcache/vllm-openai:v0.5.5
            imagePullPolicy: IfNotPresent
            env:
              - name: HF_HUB_DISABLE_TELEMETRY
                value: "1"
              - name: UCX_NET_DEVICES
                value: "all"
              - name: NCCL_CUMEM_ENABLE
                value: "1"
              - name: VLLM_NIXL_SIDE_CHANNEL_PORT   # distinct from the decoder's: both share the host IP
                value: "5557"
            args:
              - "Qwen/Qwen3-0.6B"
              - "--served-model-name"
              - "Qwen3-0.6B"
              - "--port"
              - "8001"
              - "--gpu-memory-utilization"
              - "0.4"
              - "--max-model-len"
              - "8192"
              - "--attention-backend"             # prefiller only
              - "CUSTOM"
            ports:
              - name: http
                containerPort: 8001
            volumeMounts:
              - name: hf-cache
                mountPath: /root/.cache/huggingface
            resources:
              limits:
                nvidia.com/gpu: "1"
                memory: 32Gi
              requests:
                cpu: "2"
                memory: 16Gi
            readinessProbe:
              httpGet:
                path: /health
                port: http
              initialDelaySeconds: 60
              periodSeconds: 15
              failureThreshold: 40
        volumes:
          - name: hf-cache
            emptyDir:
              sizeLimit: 20Gi
  ---
  apiVersion: v1
  kind: Service
  metadata:
    name: cb-pd-prefiller
    namespace: cacheblend-workload
  spec:
    selector:
      app: cb-pd-prefiller
    ports:
      - name: http
        port: 8001
        targetPort: http
  ---
  # Decoder — receives KV from the prefiller; never loads the CacheBlend plugin
  apiVersion: apps/v1
  kind: Deployment
  metadata:
    name: cb-pd-decoder
    namespace: cacheblend-workload
    labels:
      app: cb-pd-decoder
  spec:
    replicas: 1
    selector:
      matchLabels:
        app: cb-pd-decoder
    template:
      metadata:
        labels:
          app: cb-pd-decoder
          lmcache.ai/cacheblend-inject: "true"
        annotations:
          lmcache.ai/cacheblend-engine: "tensormesh-platform/tensormesh-cacheblend"
          lmcache.ai/pd-role: "decoder"
      spec:
        runtimeClassName: nvidia
        hostNetwork: true
        dnsPolicy: ClusterFirstWithHostNet
        containers:
          - name: vllm
            image: lmcache/vllm-openai:v0.5.5
            imagePullPolicy: IfNotPresent
            env:
              - name: HF_HUB_DISABLE_TELEMETRY
                value: "1"
              - name: UCX_NET_DEVICES
                value: "all"
              - name: NCCL_CUMEM_ENABLE
                value: "1"
              # VLLM_NIXL_SIDE_CHANNEL_PORT left unset: the webhook injects spec.pd.nixlSideChannelPort (5558)
            args:
              - "Qwen/Qwen3-0.6B"
              - "--served-model-name"
              - "Qwen3-0.6B"
              - "--port"
              - "8002"
              - "--gpu-memory-utilization"
              - "0.4"
              - "--max-model-len"
              - "8192"
            ports:
              - name: http
                containerPort: 8002
            volumeMounts:
              - name: hf-cache
                mountPath: /root/.cache/huggingface
            resources:
              limits:
                nvidia.com/gpu: "1"
                memory: 32Gi
              requests:
                cpu: "2"
                memory: 16Gi
            readinessProbe:
              httpGet:
                path: /health
                port: http
              initialDelaySeconds: 60
              periodSeconds: 15
              failureThreshold: 40
        volumes:
          - name: hf-cache
            emptyDir:
              sizeLimit: 20Gi
  ---
  apiVersion: v1
  kind: Service
  metadata:
    name: cb-pd-decoder
    namespace: cacheblend-workload
  spec:
    selector:
      app: cb-pd-decoder
    ports:
      - name: http
        port: 8002
        targetPort: http
  ---
  # Router — splits each request into a prefill leg and a decode leg
  apiVersion: apps/v1
  kind: Deployment
  metadata:
    name: cb-pd-router
    namespace: cacheblend-workload
  spec:
    replicas: 1
    selector:
      matchLabels:
        app: cb-pd-router
    template:
      metadata:
        labels:
          app: cb-pd-router
      spec:
        containers:
          - name: router
            image: vllm/vllm-router:nightly
            imagePullPolicy: IfNotPresent
            command: ["/bin/sh", "-c"]      # not an opt-in pod, so a command wrapper is fine here
            args:
              - |
                exec vllm-router \
                  --policy round_robin \
                  --vllm-pd-disaggregation \
                  --prefill http://cb-pd-prefiller.cacheblend-workload.svc.cluster.local:8001 \
                  --decode  http://cb-pd-decoder.cacheblend-workload.svc.cluster.local:8002 \
                  --host 0.0.0.0 \
                  --port 30000 \
                  --intra-node-data-parallel-size 1
            ports:
              - name: http
                containerPort: 30000
            readinessProbe:
              httpGet:
                path: /health
                port: http
              initialDelaySeconds: 5
              periodSeconds: 10
  ---
  apiVersion: v1
  kind: Service
  metadata:
    name: cb-pd-router
    namespace: cacheblend-workload
  spec:
    selector:
      app: cb-pd-router
    ports:
      - name: http
        port: 30000
        targetPort: http
  ```
</Accordion>

```bash theme={null}
kubectl apply -f vllm-cacheblend-pd.yaml
```

### Verification

**PD ConfigMap shape**

With `spec.pd` set, the engine's connection ConfigMap carries three role keys. The decoder key must not
contain `CBKVConnector` — the decoder never blends:

```bash theme={null}
kubectl -n tensormesh-platform get configmap tensormesh-cacheblend-connection \
  -o go-template='{{range $k, $_ := .data}}{{$k}}{{"\n"}}{{end}}'
# kv-transfer-config.json  kv-transfer-config-prefiller.json  kv-transfer-config-decoder.json
kubectl -n tensormesh-platform get configmap tensormesh-cacheblend-connection \
  -o jsonpath='{.data.kv-transfer-config-decoder\.json}' | grep -c CBKVConnector   # 0
```

**Each role got its own injection**

```bash theme={null}
for r in prefiller decoder; do
  P=$(kubectl -n cacheblend-workload get pod -l lmcache.ai/pd-role=$r -o name | head -1)
  echo "$r  injected=$(kubectl -n cacheblend-workload get $P -o jsonpath='{.metadata.annotations.lmcache\.ai/cacheblend-injected}')" \
       "init=$(kubectl -n cacheblend-workload get $P -o jsonpath='{.spec.initContainers[*].name}')" \
       "nixl_port=$(kubectl -n cacheblend-workload get $P -o jsonpath='{.spec.containers[0].env[?(@.name=="VLLM_NIXL_SIDE_CHANNEL_PORT")].value}')"
done
```

Expected: `injected` present on **both**; `init=cb-plugin-stage` on the prefiller **only** (empty on the
decoder); `nixl_port=5557` on the prefiller (its pre-set value) and `5558` on the decoder (injected from
`spec.pd.nixlSideChannelPort`). A `lmcache.ai/cacheblend-skip-reason` annotation on either pod means the
webhook skipped it — see [Common mistakes](#common-mistakes).

**Requests flow through both roles**

```bash theme={null}
kubectl -n cacheblend-workload port-forward svc/cb-pd-router 30000:30000 &
curl -s localhost:30000/v1/completions -H 'Content-Type: application/json' \
  -d '{"model":"Qwen3-0.6B","prompt":"Hello from PD","max_tokens":16}' | jq -r '.choices[0].text'
```

A completion back means the router split the request, the prefiller produced KV, and the decoder received
it over NIXL and generated from it.

**Non-prefix blend engaged on the prefiller**

Same check as aggregated mode — blending happens only on the prefill leg, so only the prefiller's traffic
can produce `shifted` retrieves:

```bash theme={null}
kubectl -n tensormesh-platform logs -l app.kubernetes.io/instance=tensormesh-cacheblend --tail=500 \
  | grep -E "\[match_probe\] .* matches=[1-9]|Retrieved pre-computed for [1-9].* shifted=[1-9]"
```

## L2 external storage offloading

The blend engine's L1 is CPU DRAM, bounded by `l1.sizeGB`. Adding an L2 adapter lets chunks that fall
out of L1 be served from external storage instead of recomputed; the CacheBlend lookup reads L2 through the
engine's L2→L1 prefetch, so nothing changes on the vLLM side. For example for file system offloading, add the
following l2.yaml to `cacheBlend.spec` in your helm values:

```yaml l2.yaml theme={null}
    l2Backend:
      raw:
        type: fs_native                    # C++ filesystem adapter
        config:
          base_path: /data/lmcache/l2      # = the volumeMount below
          max_capacity_gb: 500             # cap on the adapter's disk use
      storePolicy: default                 # write-through: L1 keeps the hot set, L2 extends it
    volumes:
      - name: lmcache-l2
        hostPath:
          path: /mnt/nvme0/lmcache-l2      # per-node local NVMe (the engine is a DaemonSet)
          type: DirectoryOrCreate
    volumeMounts:
      - name: lmcache-l2
        mountPath: /data/lmcache/l2
```

### Verification

**Adapter engaged**

```bash theme={null}
kubectl -n tensormesh-platform get cacheblendengine tensormesh-cacheblend -o jsonpath='{.spec.l2Backend.raw.type}{"\n"}'   # fs_native
E=$(kubectl -n tensormesh-platform get pod -l app.kubernetes.io/instance=tensormesh-cacheblend -o name | head -1)
kubectl -n tensormesh-platform exec $E -- sh -c 'ls /data/lmcache/l2 | head'   # populated after the first requests
```

**Chunks written to and read back from L2**

The engine exposes its counters on the same port as `/status` (`8080` by default):

```bash theme={null}
kubectl -n tensormesh-platform exec $E -- curl -s localhost:8080/metrics \
  | grep -E '^lmcache_mp_l2_(store_completed_objects|prefetch_load_completed)'
```

`lmcache_mp_l2_store_completed_objects_chunks_total` > 0 means chunks were stored to L2;
`lmcache_mp_l2_prefetch_load_completed_chunks_total` > 0 means chunks were read back from L2 into L1
for a lookup. With `storePolicy: default` the second only rises after L1 evictions — send more distinct
long prompts than L1 holds, then repeat an early one.

## Helm values reference

| Value                                        | Type   | Default               | Description                                                                                                                                                                                 |
| -------------------------------------------- | ------ | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `cacheBlend.enabled`                         | bool   | `false`               | Create a `CacheBlendEngine` CR. Requires `webhook.enabled: true` + cert-manager.                                                                                                            |
| `cacheBlend.name`                            | string | `""`                  | Engine CR name. Empty = `<fullname>-cacheblend`.                                                                                                                                            |
| `cacheBlend.namespace`                       | string | `""`                  | Namespace for the engine + its connection ConfigMap. Empty (recommended) = release namespace; opt-in pods bind it cross-namespace.                                                          |
| `cacheBlend.spec.l1.sizeGB`                  | int    | `60`                  | L1 (host-DRAM) chunk cache size, in GiB.                                                                                                                                                    |
| `cacheBlend.spec.image`                      | object | `lmcache/vllm-openai` | Blend-engine image (the lmcache server).                                                                                                                                                    |
| `cacheBlend.spec.injection.payloadImage`     | object | —                     | PRIVATE `cacheblend-plugin` image injected as the init container. Must match the engine's lmcache build.                                                                                    |
| `cacheBlend.spec.injection.imagePullSecrets` | list   | `[]`                  | Pull secret names appended to opt-in pods for the private payload image. The Secret(s) must already exist in the vLLM pod's namespace.                                                      |
| `cacheBlend.spec.l2Backend`                  | object | —                     | L2 tier behind L1: `raw.type`/`raw.config` (adapter, e.g. `fs_native` with `base_path`), `storePolicy` (`default` \| `skip_l1`). See [L2 filesystem offloading](#l2-filesystem-offloading). |
| `cacheBlend.spec.volumes` / `volumeMounts`   | list   | `[]`                  | Extra volumes for the engine pod — where the L2 `base_path` is mounted.                                                                                                                     |
| `webhook.enabled`                            | bool   | `true`                | Enable the CacheBlend mutating webhook. Requires cert-manager.                                                                                                                              |

## Common mistakes

* Running the opt-in vLLM pod in the **operator (release) namespace** — the webhook excludes it, so
  injection never happens
* A **bare** engine ref (`tensormesh-cacheblend`) instead of `<release-ns>/tensormesh-cacheblend` —
  the webhook looks in the pod's own namespace and finds no engine
* Setting a `command:` on the vLLM container — the webhook skips injection (appended args would never
  reach `vllm serve`)
* Workload namespace **not PSS-privileged** — Pod Security rejects the injected `hostIPC`
* The `cacheblend-plugin-pull` secret missing from the **workload** namespace — the injected init
  container `ImagePullBackOff`s (the secret must live where the pod runs, not in the operator namespace)
* **Mismatched** `cacheblend-plugin` and `vllm-openai` images — vLLM crashes at startup with an
  `ImportError`; pin a matched pair
* `webhook.enabled: false` or no cert-manager — the engine deploys but nothing injects the plugin
