Skip to main content
Most customers already run vLLM some other way before they install the Tensormesh Platform. This page covers the minimum changes needed to patch an existing vLLM Deployment so it uses the operator-managed LMCache engine in MP mode. There are two ways to do it, and they end at the same pod spec:

Webhook injection

Enable injection in the chart

Injection is configured once, on the operator release — nothing here goes into your vLLM Deployment. Two values matter: the webhook must be on (it is by default), and the engine must carry a payload image if you want payload injection as well as connection wiring.
my-values.yaml

Opt your Deployment in

Add two lines of metadata on the vLLM pod template:
At admission the webhook adds, to the target container:
The engine must be in the same namespace as the pod. The injector resolves <engine>-connection in the pod’s own namespace (req.Namespace), so the annotation takes a bare engine name, not <namespace>/<name>. The chart’s own release namespace is excluded from the webhook, so run the vLLM Deployment in a namespace alongside an engine — that is why the values above put the engine in my-workload. If you cannot move either one, wire the Deployment directly instead.

Payload injection: stop matching lmcache versions by hand

Ordinarily your vLLM image must already carry a compatible external lmcache, matched to the engine’s version. That is the awkward part of running an existing Deployment against a managed engine, and payload injection removes it. When the engine carries an injection.payloadImage (set it in the chart), the webhook injects these as well as the connection wiring: So an older vLLM image runs the engine’s lmcache without being rebuilt. The payload image is published alongside every engine release and nightly, with matching tags (v0.5.5, latest-nightly, …) — pin it to the engine’s tag, because client and server share a versioned protocol.

Verify the injection

Confirm the pod spec was rewritten:
Then prove which lmcache is actually imported at runtime — the staged one, and the baked-in one for contrast:
Different paths and versions means payload injection is doing its job.

When the webhook skips your pod

Injection fails open: the pod is always admitted, and the reason is stamped on lmcache.ai/lmcache-skip-reason. Read it before assuming the webhook is broken:

Modify the vLLM Deployment directly

Use this when webhook injection is not an option — no cert-manager, webhook.enabled: false, a container that sets command:, an engine in a different namespace, or simply because you want every change visible in your own manifests. You write by hand exactly what the webhook would have written.

What to add

Start by reading the contract the engine published:
The host is fully qualified, so — unlike webhook injection — your vLLM pod does not have to run in the engine’s namespace. Mounting the ConfigMap still does; when the namespaces differ, paste the JSON into your manifest instead.

Patch your Deployment

Referencing the ConfigMap through an env var keeps the container args-only, so nothing about how your image starts has to change — Kubernetes expands $(LMCACHE_KV_CONFIG) in args from the env defined on the same container:
If your container already runs a shell through command:, mount the ConfigMap as a file and read it inline instead — this is the form the e2e suite uses:

Match the engine’s IPC mode

The /dev/shm wiring is not cosmetic: it is how the engine and vLLM exchange CUDA IPC handles, and it has to agree with the engine’s own setting. Check it with:

Stage lmcache yourself

Without the webhook, your vLLM image must already carry an lmcache matching the engine’s version — client and server share a versioned wire protocol. If it does not, replicate what payload injection does, using the same payload image:
Verify it the same way as the webhook path — compare import lmcache with and without PYTHONPATH.

OpenShift-specific patch

This applies to both approaches. Having the shared-memory wiring in the pod spec — however it got there — does not grant the pod permission to use it. On OpenShift the vLLM pod also needs an SCC that permits it: hostIPC, or the /dev/shm hostPath used when the engine runs with isolatedIPC: false. The chart handles only the engine pod’s ServiceAccount; it does not patch your pre-existing vLLM Deployment. (An engine left on the isolatedIPC: true default needs none of this — its /dev/shm is a plain pod-private emptyDir.) Under webhook injection, a pod whose SCC forbids the injected wiring is rejected after a successful injection, so the skip-reason annotation will be absent — look at the pod’s events instead. Your vLLM pod therefore needs:
  • hostIPC: true
  • a serviceAccountName bound to a suitable SCC
If your vLLM Deployment is in the same namespace as the chart-managed engine, you can usually reuse the chart-created privileged ServiceAccount:
If the Deployment is in another namespace, that namespace needs its own ServiceAccount plus SCC binding.

Verify the patched deployment

This applies to both approaches — it checks that KV actually reaches the engine, not how the pod was wired. After rolling out the Deployment:
Then send two identical long prompts through the vLLM Service and check the engine logs. Port-forward the Service in one terminal:
In a second terminal, build the request once and send it twice:
Then inspect the engine logs:
Healthy behavior is:
  • call 1 prints a normal completion and the engine logs Stored N tokens
  • call 2 sends the exact same payload and the engine logs Prefetch request completed ... prefix hits=N
  • call 2 is usually noticeably faster than call 1 for a long enough prompt