> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tensormesh.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Configuration

> Connect metric history, GPU telemetry, and additional clusters; the rest needs no setup.

Most of the dashboard works the moment it's installed. Two things need connecting: your
**Prometheus** (for the trend charts), and, on GPU clusters, **GPU telemetry** (for per-GPU detail):

| What you see                                                      | Where it comes from             | What you configure                                                                             |
| ----------------------------------------------------------------- | ------------------------------- | ---------------------------------------------------------------------------------------------- |
| Fleet, nodes, health, cache internals, quotas                     | The operator and engines        | No setup needed                                                                                |
| Trend charts (hit rate, capacity, eviction, throughput over time) | Your Prometheus                 | `prometheus.url` + engine metrics scraping: [two steps](#trend-charts-connect-your-prometheus) |
| Per-GPU cards (utilization, memory, hardware faults)              | NVIDIA's GPU telemetry exporter | `dcgmService` (optional)                                                                       |

Everything is set through Helm values:

| Value                | What it does                                                                                        | Default                   |
| -------------------- | --------------------------------------------------------------------------------------------------- | ------------------------- |
| `engineNamespace`    | The namespace where your engines run                                                                | *(the release namespace)* |
| `prometheus.url`     | Where the UI reads metric history from                                                              | *(unset)*                 |
| `prometheus.service` | Alternative: reach Prometheus through Kubernetes instead of a direct URL (`service.namespace:port`) | *(unset)*                 |
| `dcgmService`        | GPU telemetry source (`service.namespace:port`)                                                     | *(unset)*                 |
| `clusters`           | Watch several clusters in one dashboard                                                             | *(unset, single cluster)* |
| `ingress.enabled`    | Expose the UI beyond port-forward                                                                   | `false`                   |

Cluster access itself needs no configuration: the UI picks it up automatically when running
in-cluster.

***

## Trend Charts: Connect Your Prometheus

Live numbers come straight from the engines; **history** comes from your Prometheus. Two
things must both be true: the UI can reach a Prometheus, **and** the engine's metrics are
arriving in it.

**1. Point the UI at your Prometheus.** Find its service with
`kubectl get svc -A | grep -i prometheus`. The name varies by install:

```yaml theme={null}
prometheus:
  url: http://<prometheus-service>.<namespace>.svc:9090
# e.g. kube-prometheus-stack-prometheus.monitoring.svc      (default kube-prometheus-stack)
# e.g. kube-prom-stack-kube-prome-prometheus.monitoring.svc (name follows your Helm release)
```

**2. Make sure the engine's metrics reach that Prometheus.** How they get there depends on your
setup: Prometheus scraping the engine directly, or the engine pushing through an OTel Collector.
Both paths (and how to verify which one you're on) are covered in the operator's
[Metrics](/v1.0.0/observability/metrics) page. The UI just reads what Prometheus already has.

<Accordion title="No Prometheus in the cluster at all?">
  Install the community kube-prometheus-stack. It includes the Prometheus Operator that makes
  `ServiceMonitor`s work:

  ```bash theme={null}
  helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
  helm install kube-prometheus-stack prometheus-community/kube-prometheus-stack \
    -n monitoring --create-namespace
  ```

  Then `prometheus.url` is `http://kube-prometheus-stack-prometheus.monitoring.svc:9090`, and
  the ServiceMonitor label above is `release: kube-prometheus-stack`.
</Accordion>

Charts fill once both halves are in place; if they stay empty while everything else works,
see [Troubleshooting](/v1.0.0/ui/troubleshooting).

***

## Per-GPU Detail: Connect GPU Telemetry (GPU clusters)

Per-GPU utilization, memory fill, and hardware-fault detail come from NVIDIA's DCGM exporter,
standard on GPU clusters (it ships with the NVIDIA GPU Operator), but not part of Tensormesh.
Find your exporter's service and point `dcgmService` at it (`service.namespace:port`):

```yaml theme={null}
# find yours: kubectl get svc -A | grep -i dcgm
dcgmService: dcgm-exporter.cw-exporters:9400     # CoreWeave clusters
# dcgmService: dcgm-exporter.gpu-operator:9400   # e.g. an NVIDIA GPU Operator install
```

If it's unset or unreachable, the Fleet Map shows one row per node instead of per-GPU
cards; everything else is unaffected. If your GPU cluster doesn't have the exporter at all, it
ships with the [NVIDIA GPU Operator](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html):
install that per NVIDIA's docs and the exporter service appears with it.

### GPU trend charts need Prometheus to scrape DCGM

The live per-GPU cards fall back to `dcgmService` directly, but the over-time charts read history
from Prometheus. Confirm DCGM is being scraped: query `DCGM_FI_DEV_GPU_UTIL` in Prometheus and
expect series back. On some managed clusters the exporter ships a `ServiceMonitor` labelled for a
different Prometheus (e.g. CoreWeave labels theirs `environment: grafana-monitoring`, while
kube-prometheus-stack selects `release: <your-release>`), so it returns zero series and the GPU
trend charts stay empty. Fix it by adding a `ServiceMonitor` for the DCGM exporter's service with
the label your Prometheus selects.

***

## Multiple Clusters

One dashboard can watch several clusters: the fleet-wide pages merge every cluster's engines
(labeled `· prod` / `· staging`), and per-engine pages scope to the right one.

The cluster the UI runs **in** is automatic. Each *additional* cluster needs just two things: its
`apiUrl` and a **read-only token**. There are two steps:

**1. Put your clusters (with tokens) in a Secret:**

```bash theme={null}
kubectl create secret generic tmo-ui-clusters -n tensormesh-platform \
  --from-literal=TMO_CLUSTERS='[
    {"name":"prod","apiUrl":"https://prod-api.example.com","token":"<prod-token>"},
    {"name":"staging","apiUrl":"https://staging-api.example.com","token":"<staging-token>"}
  ]'
```

**2. Point your values file at it:**

```yaml theme={null}
extraEnvFrom:
  - secretRef:
      name: tmo-ui-clusters
```

Reinstall, and every cluster shows up. Each entry can also set its own `namespace`,
`prometheusUrl`, or `dcgmService`, and **`"insecure": true`** if that cluster's API uses a
self-signed certificate.

<Accordion title="How do I get a read-only token for a cluster?">
  Run this **in that cluster**: it creates a read-only service account (the same access the UI
  uses) and prints a token:

  ```bash theme={null}
  kubectl apply -f - <<'EOF'
  apiVersion: v1
  kind: ServiceAccount
  metadata: { name: tmo-ui-reader, namespace: tensormesh-platform }
  ---
  apiVersion: rbac.authorization.k8s.io/v1
  kind: ClusterRole
  metadata: { name: tmo-ui-reader }
  rules:
    - { apiGroups: ["lmcache.lmcache.ai"], resources: ["lmcacheengines", "lmcachecoordinators"], verbs: ["get", "list", "watch"] }
    - { apiGroups: [""], resources: ["pods", "namespaces", "nodes", "events"], verbs: ["get", "list"] }
    - { apiGroups: [""], resources: ["pods/proxy", "services/proxy"], verbs: ["get"] }
  ---
  apiVersion: rbac.authorization.k8s.io/v1
  kind: ClusterRoleBinding
  metadata: { name: tmo-ui-reader }
  roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: tmo-ui-reader }
  subjects:
    - { kind: ServiceAccount, name: tmo-ui-reader, namespace: tensormesh-platform }
  EOF

  kubectl -n tensormesh-platform create token tmo-ui-reader --duration=8760h
  ```
</Accordion>

### Sharing one token across clusters

To skip per-cluster tokens, list the clusters in your values file and put a single
`TMO_KUBE_TOKEN` in the Secret; every cluster inherits it.

```yaml theme={null}
# values.yaml
clusters:
  - { name: prod, apiUrl: https://prod-api.example.com }
  - { name: dev, apiUrl: https://dev-api.example.com }
  # add `insecure: true` to any entry whose API uses a self-signed certificate
extraEnvFrom:
  - secretRef: { name: tmo-ui-clusters }
```

```bash theme={null}
kubectl create secret generic tmo-ui-clusters -n tensormesh-platform \
  --from-literal=TMO_KUBE_TOKEN='<shared-token>'
```

***

## Related

* [Troubleshooting & FAQ](/v1.0.0/ui/troubleshooting)
* [Monitoring Your Fleet](/v1.0.0/ui/monitoring-your-fleet)
* [Operator Observability](/v1.0.0/observability)
