> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tensormesh.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Configuration

> Connect metric history, GPU telemetry, and additional clusters — everything else works out of the box.

Most of the dashboard works the moment it's installed. Two things need connecting: your
**Prometheus** (for the trend charts), and — on GPU clusters — **GPU telemetry** (for per-GPU detail):

| What you see                                                      | Where it comes from             | What you configure                                                                              |
| ----------------------------------------------------------------- | ------------------------------- | ----------------------------------------------------------------------------------------------- |
| Fleet, nodes, health, cache internals, quotas                     | The operator and engines        | No setup needed                                                                                 |
| Trend charts (hit rate, capacity, eviction, throughput over time) | Your Prometheus                 | `prometheus.url` + engine metrics scraping — [two steps](#trend-charts-connect-your-prometheus) |
| Per-GPU cards (utilization, memory, hardware faults)              | NVIDIA's GPU telemetry exporter | `dcgmService` — optional                                                                        |

Everything is set through Helm values:

| Value                | What it does                                                                                        | Default                    |
| -------------------- | --------------------------------------------------------------------------------------------------- | -------------------------- |
| `engineNamespace`    | The namespace where your engines run                                                                | *(the release namespace)*  |
| `prometheus.url`     | Where the UI reads metric history from                                                              | *(unset)*                  |
| `prometheus.service` | Alternative: reach Prometheus through Kubernetes instead of a direct URL (`service.namespace:port`) | *(unset)*                  |
| `dcgmService`        | GPU telemetry source (`service.namespace:port`)                                                     | *(unset)*                  |
| `clusters`           | Watch several clusters in one dashboard                                                             | *(unset — single cluster)* |
| `ingress.enabled`    | Expose the UI beyond port-forward                                                                   | `false`                    |

Cluster access itself needs no configuration — the UI picks it up automatically when running
in-cluster.

***

## Trend Charts: Connect Your Prometheus

Live numbers come straight from the engines; **history** comes from your Prometheus. Two
things must both be true: the UI can reach a Prometheus, **and** the engine's metrics are
arriving in it.

**1. Point the UI at your Prometheus.** Find its service with
`kubectl get svc -A | grep -i prometheus` — the name varies by install:

```yaml theme={null}
prometheus:
  url: http://<prometheus-service>.<namespace>.svc:9090
# e.g. kube-prometheus-stack-prometheus.monitoring.svc      (default kube-prometheus-stack)
# e.g. kube-prom-stack-kube-prome-prometheus.monitoring.svc (name follows your Helm release)
```

**2. Get the engine metrics into that Prometheus.** The engine doesn't expose metrics for
scraping by default. Turn its metrics endpoint on — this goes in the **operator chart's**
values (it configures the `LMCacheEngine`), *not* the UI's:

```yaml theme={null}
engine:
  spec:
    prometheus:
      enabled: true        # the engine serves /metrics on :8080 (the `http` port)
```

then give Prometheus a `ServiceMonitor` to scrape it — the operator already creates the
`<engine>-metrics` Service it selects:

```yaml theme={null}
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: lmcache-engine
  namespace: tensormesh-operator
  labels:
    release: <your-prometheus-release>   # must match your Prometheus' serviceMonitorSelector
spec:
  selector:
    matchLabels:
      app.kubernetes.io/component: cache-engine
  endpoints:
    - port: http
      path: /metrics
      interval: 30s
```

<Note>
  `observability.enabled=true` on the operator chart is a **different pipeline** — it *pushes*
  metrics out through an OTel Collector (for example to a managed metrics service). That only
  feeds the UI if your collector happens to deliver into the same Prometheus the UI reads. The
  endpoint + ServiceMonitor wiring above is self-contained and works on any cluster.
</Note>

<Accordion title="No Prometheus in the cluster at all?">
  Install the community kube-prometheus-stack — it includes the Prometheus Operator that makes
  `ServiceMonitor`s work:

  ```bash theme={null}
  helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
  helm install kube-prometheus-stack prometheus-community/kube-prometheus-stack \
    -n monitoring --create-namespace
  ```

  Then `prometheus.url` is `http://kube-prometheus-stack-prometheus.monitoring.svc:9090`, and
  the ServiceMonitor label above is `release: kube-prometheus-stack`.
</Accordion>

Charts fill once both halves are in place; if they stay empty while everything else works,
see [Troubleshooting](/ui/troubleshooting).

***

## Per-GPU Detail: Connect GPU Telemetry (GPU clusters)

Per-GPU utilization, memory fill, and hardware-fault detail come from NVIDIA's DCGM exporter —
standard on GPU clusters (it ships with the NVIDIA GPU Operator), but not part of Tensormesh.
Find your exporter's service and point `dcgmService` at it (`service.namespace:port`):

```yaml theme={null}
# find yours: kubectl get svc -A | grep -i dcgm
dcgmService: dcgm-exporter.cw-exporters:9400     # CoreWeave clusters
# dcgmService: dcgm-exporter.gpu-operator:9400   # e.g. an NVIDIA GPU Operator install
```

If it's unset or unreachable, the Fleet Map simply shows one row per node instead of per-GPU
cards; everything else is unaffected. If your GPU cluster doesn't have the exporter at all, it
ships with the [NVIDIA GPU Operator](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html) —
install that per NVIDIA's docs and the exporter service appears with it.

<Note>
  **GPU *trend* charts need Prometheus to scrape DCGM** (the live per-GPU cards fall back to
  `dcgmService` directly, but the over-time charts read history from Prometheus). Confirm it's
  being scraped — query `DCGM_FI_DEV_GPU_UTIL` in Prometheus and expect series back. On some
  managed clusters the exporter ships a `ServiceMonitor` labelled for a *different* Prometheus
  (e.g. CoreWeave labels theirs `environment: grafana-monitoring`, while kube-prometheus-stack
  selects `release: <your-release>`), so it returns zero series and the GPU trend charts stay
  empty. Fix it by adding a `ServiceMonitor` for the DCGM exporter's service with the label your
  Prometheus selects.
</Note>

***

## Multiple Clusters

One dashboard can watch several clusters — the fleet-wide pages merge every cluster's engines
(labeled `· prod` / `· staging`), and per-engine pages scope to the right one.

The cluster the UI runs **in** is automatic. Each *additional* cluster needs just two things: its
`apiUrl` and a **read-only token**. It's two steps:

**1. Put your clusters (with tokens) in a Secret:**

```bash theme={null}
kubectl create secret generic tmo-ui-clusters -n tensormesh-operator \
  --from-literal=TMO_CLUSTERS='[
    {"name":"prod","apiUrl":"https://prod-api.example.com","token":"<prod-token>"},
    {"name":"staging","apiUrl":"https://staging-api.example.com","token":"<staging-token>"}
  ]'
```

**2. Point your values file at it:**

```yaml theme={null}
extraEnvFrom:
  - secretRef:
      name: tmo-ui-clusters
```

Reinstall, and every cluster shows up. Each entry can also set its own `namespace`,
`prometheusUrl`, or `dcgmService` — and **`"insecure": true`** if that cluster's API uses a
self-signed certificate.

<Accordion title="How do I get a read-only token for a cluster?">
  Run this **in that cluster** — it creates a read-only service account (the same access the UI
  uses) and prints a token:

  ```bash theme={null}
  kubectl apply -f - <<'EOF'
  apiVersion: v1
  kind: ServiceAccount
  metadata: { name: tmo-ui-reader, namespace: tensormesh-operator }
  ---
  apiVersion: rbac.authorization.k8s.io/v1
  kind: ClusterRole
  metadata: { name: tmo-ui-reader }
  rules:
    - { apiGroups: ["lmcache.lmcache.ai"], resources: ["lmcacheengines", "lmcachecoordinators"], verbs: ["get", "list", "watch"] }
    - { apiGroups: [""], resources: ["pods", "namespaces", "nodes", "events"], verbs: ["get", "list"] }
    - { apiGroups: [""], resources: ["pods/proxy", "services/proxy"], verbs: ["get"] }
  ---
  apiVersion: rbac.authorization.k8s.io/v1
  kind: ClusterRoleBinding
  metadata: { name: tmo-ui-reader }
  roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: tmo-ui-reader }
  subjects:
    - { kind: ServiceAccount, name: tmo-ui-reader, namespace: tensormesh-operator }
  EOF

  kubectl -n tensormesh-operator create token tmo-ui-reader --duration=8760h
  ```
</Accordion>

<Tip>
  **One token for every cluster?** Skip the per-cluster tokens — list the clusters in your **values
  file** and put a single `TMO_KUBE_TOKEN` in the Secret; every cluster inherits it.

  ```yaml theme={null}
  # values.yaml
  clusters:
    - { name: prod, apiUrl: https://prod-api.example.com }
    - { name: dev, apiUrl: https://dev-api.example.com }
    # add `insecure: true` to any entry whose API uses a self-signed certificate
  extraEnvFrom:
    - secretRef: { name: tmo-ui-clusters }
  ```

  ```bash theme={null}
  kubectl create secret generic tmo-ui-clusters -n tensormesh-operator \
    --from-literal=TMO_KUBE_TOKEN='<shared-token>'
  ```
</Tip>

***

## Related

* [Troubleshooting & FAQ](/ui/troubleshooting)
* [Monitoring Your Fleet](/ui/monitoring-your-fleet)
* [Operator Observability](/observability)
