> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tensormesh.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Modify an Existing vLLM Deployment

> Patch an existing vLLM deployment to consume the LMCache engine created by the Tensormesh Operator.

Most customers already run vLLM some other way before they install the Tensormesh Platform.
This page covers the **minimum changes** needed to patch an existing vLLM Deployment so it
uses the operator-managed LMCache engine in MP mode.

There are two ways to do it, and they end at the same pod spec:

|                                        | [Webhook injection](#webhook-injection)                                                       | [Modify the Deployment directly](#modify-the-vllm-deployment-directly) |
| -------------------------------------- | --------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- |
| What you change                        | two lines of pod metadata                                                                     | the container's args, env and volumes                                  |
| Who writes the wiring                  | the mutating webhook, at admission                                                            | you, in your own manifests                                             |
| Needs cert-manager + `webhook.enabled` | yes                                                                                           | no                                                                     |
| `lmcache` version matching             | automatic, via [payload injection](#payload-injection-stop-matching-lmcache-versions-by-hand) | yours to manage (or stage the payload yourself)                        |
| Engine in another namespace            | no — the webhook resolves the engine in the pod's own namespace                               | yes — the connector config carries a fully-qualified Service name      |
| Works with `command:` on the container | no — the webhook skips such pods                                                              | yes                                                                    |

## Webhook injection

### Enable injection in the chart

Injection is configured once, on the operator release — nothing here goes into
your vLLM Deployment. Two values matter: the webhook must be on (it is by
default), and the engine must carry a payload image if you want
[payload injection](#payload-injection-stop-matching-lmcache-versions-by-hand)
as well as connection wiring.

```yaml my-values.yaml theme={null}
webhook:
  enabled: true                 # default; requires cert-manager

engine:
  enabled: true
  namespace: my-workload        # NOT the release namespace — see the warning below
  spec:
    l1:
      sizeGB: 60
    image:
      repository: lmcache/vllm-openai
      tag: v0.5.5
    injection:
      payloadImage:
        repository: lmcache/lmcache-payload
        tag: v0.5.5               # must match the engine's lmcache — same wire protocol
        pullPolicy: IfNotPresent
```

```bash theme={null}
helm upgrade --install tensormesh-platform \
  oci://artifacts.tensormesh.ai/tensormesh-production/charts/tensormesh-platform \
  --version 1.0.0 \
  --namespace tensormesh-platform --create-namespace \
  -f my-values.yaml
```

### Opt your Deployment in

Add two lines of metadata on the vLLM pod template:

```yaml theme={null}
spec:
  template:
    metadata:
      labels:
        lmcache.ai/lmcache-inject: "true"      
      annotations:
        lmcache.ai/lmcache-engine: "tensormesh-platform-default-engine"  
```

At admission the webhook adds, to the target container:

| Injected                              | Purpose                                                                                                                                                                             |
| ------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--kv-transfer-config <json>`         | the engine's connection contract, read from `<engine>-connection`                                                                                                                   |
| `/dev/shm`                            | MP-mode shared memory — pod-private in isolated-IPC mode (the default since operator `v0.5.5`); the host's `/dev/shm`, or `hostIPC`, only when the engine sets `isolatedIPC: false` |
| `PYTHONHASHSEED=0`                    | deterministic token hashing, so vLLM and the engine derive the same cache keys                                                                                                      |
| `lmcache.ai/lmcache-injected: "true"` | idempotency stamp — a second admission is a no-op                                                                                                                                   |

<Note>
  **The engine must be in the same namespace as the pod.** The injector resolves
  `<engine>-connection` in the pod's own namespace (`req.Namespace`), so the
  annotation takes a bare engine name, not `<namespace>/<name>`. The chart's own
  release namespace is excluded from the webhook, so run the vLLM Deployment in a
  namespace alongside an engine — that is why the values above put the engine in
  `my-workload`. If you cannot move either one,
  [wire the Deployment directly](#modify-the-vllm-deployment-directly) instead.
</Note>

### Payload injection: stop matching lmcache versions by hand

Ordinarily your vLLM image must already carry a **compatible external `lmcache`**,
matched to the engine's version. That is the awkward part of running an existing
Deployment against a managed engine, and payload injection removes it.

When the engine carries an `injection.payloadImage`
([set it in the chart](#enable-injection-in-the-chart)), the webhook injects
these as well as the connection wiring:

| Injected                                     | Purpose                                                            |
| -------------------------------------------- | ------------------------------------------------------------------ |
| `lmcache-payload-stage` init container       | copies the engine's `lmcache` into a shared volume                 |
| `lmcache-payload` emptyDir + read-only mount | where the staged runtime lands                                     |
| `PYTHONPATH=/lmcache-payload`                | makes vLLM import the **staged** lmcache ahead of its baked-in one |

So an older vLLM image runs the engine's lmcache without being rebuilt. The
payload image is published alongside every engine release and nightly, with
matching tags (`v0.5.5`, `latest-nightly`, …) — pin it to the engine's tag,
because client and server share a versioned protocol.

### Verify the injection

Confirm the pod spec was rewritten:

```bash theme={null}
kubectl get pod <pod> -o jsonpath='{.metadata.annotations}' | grep lmcache
kubectl get pod <pod> -o jsonpath='{.spec.initContainers[*].name}'
kubectl get pod <pod> -o jsonpath='{.spec.containers[0].env[?(@.name=="PYTHONPATH")].value}'
```

Then prove which `lmcache` is actually imported at runtime — the staged one, and
the baked-in one for contrast:

```bash theme={null}
# staged (what vLLM uses)
kubectl exec <pod> -c vllm -- python3 -c 'import lmcache; print(lmcache.__version__, lmcache.__file__)'
#   0.5.5 /lmcache-payload/lmcache/__init__.py

# baked into the image
kubectl exec <pod> -c vllm -- env -u PYTHONPATH python3 -c 'import lmcache; print(lmcache.__version__, lmcache.__file__)'
#   0.5.2 /opt/venv/lib/python3.12/site-packages/lmcache/__init__.py
```

Different paths and versions means payload injection is doing its job.

### When the webhook skips your pod

Injection **fails open**: the pod is always admitted, and the reason is stamped on
`lmcache.ai/lmcache-skip-reason`. Read it before assuming the webhook is broken:

```bash theme={null}
kubectl get pod <pod> -o jsonpath='{.metadata.annotations.lmcache\.ai/lmcache-skip-reason}'
```

| Skip reason                  | Meaning                                         | What to do                                                             |
| ---------------------------- | ----------------------------------------------- | ---------------------------------------------------------------------- |
| `command-override`           | the container sets `command:`                   | move the flags into `args:` so the image's entrypoint runs — see below |
| `engine-not-found`           | no `<engine>-connection` in the pod's namespace | fix the annotation, or move the pod to the engine's namespace          |
| `kv-transfer-config-present` | you already pass `--kv-transfer-config`         | intentional; yours is never clobbered                                  |
| `target-container-not-found` | multi-container pod, wrong target               | name it with `lmcache.ai/lmcache-container`                            |
| `payload-image-unset`        | engine has no `injection.payloadImage`          | set it, or supply `lmcache` in your image                              |

## Modify the vLLM Deployment directly

Use this when webhook injection is not an option — no cert-manager, `webhook.enabled: false`, a
container that sets `command:`, an engine in a different namespace, or simply because you want every
change visible in your own manifests. You write by hand exactly what the webhook would have written.

### What to add

| What the webhook injects      | What you add by hand                                                                                                |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| `--kv-transfer-config <json>` | the `kv-transfer-config.json` value from the `<engine>-connection` ConfigMap, passed as an arg                      |
| `PYTHONHASHSEED=0`            | the same env var on the vLLM container                                                                              |
| `/dev/shm` wiring             | a volume at `/dev/shm` matching [the engine's IPC mode](#match-the-engines-ipc-mode)                                |
| staged `lmcache`              | an `lmcache` in your image matching the engine's version — or [stage the payload yourself](#stage-lmcache-yourself) |

Start by reading the contract the engine published:

```bash theme={null}
kubectl get configmap <engine>-connection -n <engine-namespace> \
  -o jsonpath='{.data.kv-transfer-config\.json}'
```

```json theme={null}
{
  "kv_connector": "LMCacheMPConnector",
  "kv_connector_module_path": "lmcache.integration.vllm.lmcache_mp_connector",
  "kv_role": "kv_both",
  "kv_connector_extra_config": {
    "lmcache.mp.host": "tcp://<engine>.<engine-namespace>.svc.cluster.local",
    "lmcache.mp.port": "5555"
  }
}
```

The host is fully qualified, so — unlike webhook injection — your vLLM pod does not have to run in
the engine's namespace. Mounting the ConfigMap still does; when the namespaces differ, paste the
JSON into your manifest instead.

### Patch your Deployment

Referencing the ConfigMap through an env var keeps the container **args-only**, so nothing about how
your image starts has to change — Kubernetes expands `$(LMCACHE_KV_CONFIG)` in `args` from the env
defined on the same container:

```yaml theme={null}
spec:
  template:
    spec:
      containers:
        - name: vllm
          env:
            # Deterministic token hashing. Without it vLLM and the engine hash the
            # same prompt to different keys, and every lookup misses.
            - name: PYTHONHASHSEED
              value: "0"
            # The engine's connection contract. Read at pod start, so a rolling
            # restart picks up a changed engine port or name.
            - name: LMCACHE_KV_CONFIG
              valueFrom:
                configMapKeyRef:
                  name: tensormesh-platform-default-engine-connection
                  key: kv-transfer-config.json
          args:
            # … your existing vLLM args, unchanged …
            - --kv-transfer-config
            - $(LMCACHE_KV_CONFIG)
          volumeMounts:
            - name: lmcache-dev-shm
              mountPath: /dev/shm
      volumes:
        # isolatedIPC: true (the default) — see the next section.
        - name: lmcache-dev-shm
          emptyDir:
            medium: Memory
```

<Note>
  If your container already runs a shell through `command:`, mount the ConfigMap as a file and read
  it inline instead — this is the form the e2e suite uses:

  ```yaml theme={null}
  args:
    - |
      exec python3 -m vllm.entrypoints.openai.api_server \
        --model <your-model> \
        --kv-transfer-config "$(cat /etc/lmcache/kv-transfer-config.json)"
  volumeMounts:            # on the vLLM container
    - { name: kv-transfer-config, mountPath: /etc/lmcache, readOnly: true }
  volumes:                 # on the pod spec
    - name: kv-transfer-config
      configMap: { name: tensormesh-platform-default-engine-connection }
  ```
</Note>

### Match the engine's IPC mode

The `/dev/shm` wiring is not cosmetic: it is how the engine and vLLM exchange CUDA IPC handles, and
it has to agree with the engine's own setting. Check it with:

```bash theme={null}
kubectl get lmcacheengine <engine> -n <engine-namespace> \
  -o jsonpath='{.spec.isolatedIPC}{" hostIPC="}{.spec.hostIPC}{"\n"}'
```

| Engine                                                               | What the pod needs                                                                                                                                                                                         |
| -------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `isolatedIPC: true` — unset defaults to this for `gpuVendor: nvidia` | a pod-private `/dev/shm`: `emptyDir` with `medium: Memory`, as above. Nothing is shared with the engine; the volume exists because Kubernetes' default 64Mi `/dev/shm` is too small for vLLM's own workers |
| `isolatedIPC: false`                                                 | the **host's** tmpfs: `hostPath` with `path: /dev/shm`, `type: Directory`, mounted at `/dev/shm`                                                                                                           |
| `isolatedIPC: false` **and** `hostIPC: true`                         | `hostIPC: true` on the pod spec, and no `/dev/shm` volume at all                                                                                                                                           |

### Stage lmcache yourself

Without the webhook, your vLLM image must already carry an `lmcache` matching the engine's version —
client and server share a versioned wire protocol. If it does not, replicate what payload injection
does, using the same payload image:

```yaml theme={null}
spec:
  template:
    spec:
      initContainers:
        # No command/args: the payload image's own entrypoint copies its tree
        # into $SHARED_DIR.
        - name: lmcache-payload-stage
          image: lmcache/lmcache-payload:v0.5.5   # pin to the engine's lmcache
          env:
            - name: SHARED_DIR
              value: /lmcache-payload
          volumeMounts:
            - name: lmcache-payload
              mountPath: /lmcache-payload
      containers:
        - name: vllm
          env:
            # Import the staged tree ahead of the image's baked-in lmcache.
            - name: PYTHONPATH
              value: /lmcache-payload
          volumeMounts:
            - name: lmcache-payload
              mountPath: /lmcache-payload
              readOnly: true
      volumes:
        - name: lmcache-payload
          emptyDir: {}
```

Verify it the same way as [the webhook path](#verify-the-injection) — compare
`import lmcache` with and without `PYTHONPATH`.

## OpenShift-specific patch

This applies to **both** approaches. Having the shared-memory wiring in the pod
spec — however it got there — does not grant the pod permission to use it. On
OpenShift the vLLM pod also needs an SCC that permits it: `hostIPC`, or the
`/dev/shm` hostPath used when the engine runs with `isolatedIPC: false`. The
chart handles only the engine pod's ServiceAccount; it does not patch your
pre-existing vLLM Deployment. (An engine left on the `isolatedIPC: true` default
needs none of this — its `/dev/shm` is a plain pod-private `emptyDir`.)

Under webhook injection, a pod whose SCC forbids the injected wiring is
**rejected after** a successful injection, so the skip-reason annotation will be
absent — look at the pod's events instead.

Your vLLM pod therefore needs:

* `hostIPC: true`
* a `serviceAccountName` bound to a suitable SCC

If your vLLM Deployment is in the same namespace as the chart-managed engine, you
can usually reuse the chart-created privileged ServiceAccount:

```yaml theme={null}
spec:
  template:
    spec:
      hostIPC: true
      serviceAccountName: tensormesh-platform-engine-privileged
```

If the Deployment is in another namespace, that namespace needs its own
ServiceAccount plus SCC binding.

## Verify the patched deployment

This applies to both approaches — it checks that KV actually reaches the engine,
not how the pod was wired. After rolling out the Deployment:

```bash theme={null}
kubectl rollout status deploy/<your-vllm-deployment> --timeout=15m
kubectl get pod -l app=<your-label> -o wide
kubectl exec -it <vllm-pod> -- \
  python3 -c 'import lmcache.integration.vllm.lmcache_mp_connector as m; print(m.__file__)'
```

Then send two identical long prompts through the vLLM Service and check the engine logs.

Port-forward the Service in one terminal:

```bash theme={null}
kubectl port-forward -n <your-namespace> svc/<your-vllm-service> 8000:8000
```

In a second terminal, build the request once and send it twice:

```bash theme={null}
PROMPT=$(python3 -c "print('Tell me a long story about a brave knight. ' * 30)")
PAYLOAD=$(jq -n --arg p "$PROMPT" \
  '{model:"Qwen/Qwen3-0.6B", prompt:$p, max_tokens:20, temperature:0}')

# Call 1 — cold, should STORE KV into LMCache
time curl -sS http://localhost:8000/v1/completions \
  -H 'Content-Type: application/json' -d "$PAYLOAD" | jq -r '.choices[0].text'

# Call 2 — identical, should RETRIEVE prefix KV from LMCache
time curl -sS http://localhost:8000/v1/completions \
  -H 'Content-Type: application/json' -d "$PAYLOAD" | jq -r '.choices[0].text'
```

Then inspect the engine logs:

```bash theme={null}
# select by the engine CR's name — the component label (cache-engine) also
# matches CacheBlend engine pods when those are enabled
kubectl logs -n <your-namespace> \
  -l app.kubernetes.io/instance=tensormesh-platform-default-engine --tail=200 \
  | grep -E "Stored [0-9]+ tokens|Prefetch request completed.*prefix hits"
```

Healthy behavior is:

* call 1 prints a normal completion and the engine logs `Stored N tokens`
* call 2 sends the exact same payload and the engine logs
  `Prefetch request completed ... prefix hits=N`
* call 2 is usually noticeably faster than call 1 for a long enough prompt
