Everything is set through Helm values:
Cluster access itself needs no configuration — the UI picks it up automatically when running
in-cluster.
Trend Charts: Connect Your Prometheus
Live numbers come straight from the engines; history comes from your Prometheus. Two things must both be true: the UI can reach a Prometheus, and the engine’s metrics are arriving in it. 1. Point the UI at your Prometheus. Find its service withkubectl get svc -A | grep -i prometheus — the name varies by install:
LMCacheEngine), not the UI’s:
ServiceMonitor to scrape it — the operator already creates the
<engine>-metrics Service it selects:
observability.enabled=true on the operator chart is a different pipeline — it pushes
metrics out through an OTel Collector (for example to a managed metrics service). That only
feeds the UI if your collector happens to deliver into the same Prometheus the UI reads. The
endpoint + ServiceMonitor wiring above is self-contained and works on any cluster.No Prometheus in the cluster at all?
No Prometheus in the cluster at all?
Install the community kube-prometheus-stack — it includes the Prometheus Operator that makes
Then
ServiceMonitors work:prometheus.url is http://kube-prometheus-stack-prometheus.monitoring.svc:9090, and
the ServiceMonitor label above is release: kube-prometheus-stack.Per-GPU Detail: Connect GPU Telemetry (GPU clusters)
Per-GPU utilization, memory fill, and hardware-fault detail come from NVIDIA’s DCGM exporter — standard on GPU clusters (it ships with the NVIDIA GPU Operator), but not part of Tensormesh. Find your exporter’s service and pointdcgmService at it (service.namespace:port):
GPU trend charts need Prometheus to scrape DCGM (the live per-GPU cards fall back to
dcgmService directly, but the over-time charts read history from Prometheus). Confirm it’s
being scraped — query DCGM_FI_DEV_GPU_UTIL in Prometheus and expect series back. On some
managed clusters the exporter ships a ServiceMonitor labelled for a different Prometheus
(e.g. CoreWeave labels theirs environment: grafana-monitoring, while kube-prometheus-stack
selects release: <your-release>), so it returns zero series and the GPU trend charts stay
empty. Fix it by adding a ServiceMonitor for the DCGM exporter’s service with the label your
Prometheus selects.Multiple Clusters
One dashboard can watch several clusters — the fleet-wide pages merge every cluster’s engines (labeled· prod / · staging), and per-engine pages scope to the right one.
The cluster the UI runs in is automatic. Each additional cluster needs just two things: its
apiUrl and a read-only token. It’s two steps:
1. Put your clusters (with tokens) in a Secret:
namespace,
prometheusUrl, or dcgmService — and "insecure": true if that cluster’s API uses a
self-signed certificate.
How do I get a read-only token for a cluster?
How do I get a read-only token for a cluster?
Run this in that cluster — it creates a read-only service account (the same access the UI
uses) and prints a token:

