Webhook injection
Enable injection in the chart
Injection is configured once, on the operator release — nothing here goes into your vLLM Deployment. Two values matter: the webhook must be on (it is by default), and the engine must carry a payload image if you want payload injection as well as connection wiring.my-values.yaml
Opt your Deployment in
Add two lines of metadata on the vLLM pod template:The engine must be in the same namespace as the pod. The injector resolves
<engine>-connection in the pod’s own namespace (req.Namespace), so the
annotation takes a bare engine name, not <namespace>/<name>. The chart’s own
release namespace is excluded from the webhook, so run the vLLM Deployment in a
namespace alongside an engine — that is why the values above put the engine in
my-workload. If you cannot move either one,
wire the Deployment directly instead.Payload injection: stop matching lmcache versions by hand
Ordinarily your vLLM image must already carry a compatible externallmcache,
matched to the engine’s version. That is the awkward part of running an existing
Deployment against a managed engine, and payload injection removes it.
When the engine carries an injection.payloadImage
(set it in the chart), the webhook injects
these as well as the connection wiring:
So an older vLLM image runs the engine’s lmcache without being rebuilt. The
payload image is published alongside every engine release and nightly, with
matching tags (
v0.5.5, latest-nightly, …) — pin it to the engine’s tag,
because client and server share a versioned protocol.
Verify the injection
Confirm the pod spec was rewritten:lmcache is actually imported at runtime — the staged one, and
the baked-in one for contrast:
When the webhook skips your pod
Injection fails open: the pod is always admitted, and the reason is stamped onlmcache.ai/lmcache-skip-reason. Read it before assuming the webhook is broken:
Modify the vLLM Deployment directly
Use this when webhook injection is not an option — no cert-manager,webhook.enabled: false, a
container that sets command:, an engine in a different namespace, or simply because you want every
change visible in your own manifests. You write by hand exactly what the webhook would have written.
What to add
Start by reading the contract the engine published:
Patch your Deployment
Referencing the ConfigMap through an env var keeps the container args-only, so nothing about how your image starts has to change — Kubernetes expands$(LMCACHE_KV_CONFIG) in args from the env
defined on the same container:
If your container already runs a shell through
command:, mount the ConfigMap as a file and read
it inline instead — this is the form the e2e suite uses:Match the engine’s IPC mode
The/dev/shm wiring is not cosmetic: it is how the engine and vLLM exchange CUDA IPC handles, and
it has to agree with the engine’s own setting. Check it with:
Stage lmcache yourself
Without the webhook, your vLLM image must already carry anlmcache matching the engine’s version —
client and server share a versioned wire protocol. If it does not, replicate what payload injection
does, using the same payload image:
import lmcache with and without PYTHONPATH.
OpenShift-specific patch
This applies to both approaches. Having the shared-memory wiring in the pod spec — however it got there — does not grant the pod permission to use it. On OpenShift the vLLM pod also needs an SCC that permits it:hostIPC, or the
/dev/shm hostPath used when the engine runs with isolatedIPC: false. The
chart handles only the engine pod’s ServiceAccount; it does not patch your
pre-existing vLLM Deployment. (An engine left on the isolatedIPC: true default
needs none of this — its /dev/shm is a plain pod-private emptyDir.)
Under webhook injection, a pod whose SCC forbids the injected wiring is
rejected after a successful injection, so the skip-reason annotation will be
absent — look at the pod’s events instead.
Your vLLM pod therefore needs:
hostIPC: true- a
serviceAccountNamebound to a suitable SCC
Verify the patched deployment
This applies to both approaches — it checks that KV actually reaches the engine, not how the pod was wired. After rolling out the Deployment:- call 1 prints a normal completion and the engine logs
Stored N tokens - call 2 sends the exact same payload and the engine logs
Prefetch request completed ... prefix hits=N - call 2 is usually noticeably faster than call 1 for a long enough prompt

