> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tensormesh.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Monitoring Your Fleet

> Overview, Fleet Map, Metrics, Health, and Thresholds — what each screen shows and when to use it.

Here's what each screen is for:

| Screen         | What it's for                                                       |
| -------------- | ------------------------------------------------------------------- |
| **Overview**   | The landing page — what needs attention right now                   |
| **Fleet Map**  | Drill down from engines to nodes to individual GPUs                 |
| **Metrics**    | Trends over time — utilization, cache fill, activity, quotas        |
| **Health**     | Inside the engine — the cache pipeline, events, and suggested fixes |
| **Thresholds** | Decide what counts as a problem                                     |

Overview and Fleet Map cover the whole fleet (and every cluster, in a
[multi-cluster](/v1.0.0/ui/configuration#multiple-clusters) setup); the other screens focus
on one engine, chosen with the filter in the page header.

***

## Overview

The landing page: overall fleet health, per-engine cards, and two lists that name what needs
attention. Each entry links straight to the affected node.

* **Alerts** — facts your cluster is reporting: an engine not fully ready, a failing health
  check (with the failing nodes named), a storage backend down.
* **Issues** — *your* [thresholds](/v1.0.0/ui/thresholds-and-notifications) applied to live
  values, each naming the worst offender.

***

## Fleet Map

Every engine, node, and GPU in one view. Pick a lens (Health, Utilization, or Capacity) and the
whole map colors by it, so the affected node stands out. Selecting a node shows its key numbers,
its GPUs (when GPU telemetry is connected), and its recent trends. Nodes can be nicknamed, for
example by the model they serve.

***

## Metrics

Per-engine trends over time across four tabs:
Serving, Cache Capacity, Cache Activity, and Quotas. See [Metrics](/v1.0.0/ui/metrics) for every chart and
the metric behind it.

***

## Health

What's happening inside the engine: health tiles, the cache pipeline, and the event log. Issues &
Fixes turns anything failing into a problem card with a suggested fix that matches the cause (a
storage-backend failure points at the backend, not "restart the pod").

Health here is the engine's own self-assessment, not just pod status, so it catches problems
standard Kubernetes tooling can't see, like a storage connection silently disabled after repeated
failures while the pod still looks fine.

***

## Thresholds

Where you decide what counts as degraded vs. critical: driving the colors, the Issues list,
and notifications. Covered in
[Thresholds & Notifications](/v1.0.0/ui/thresholds-and-notifications).

***

## Related

* [Thresholds & Notifications](/v1.0.0/ui/thresholds-and-notifications)
* [Multi-Tenancy](/v1.0.0/management/multi-tenancy)
* [Troubleshooting & FAQ](/v1.0.0/ui/troubleshooting)
* [CLI Reference](/v1.0.0/reference/cli)
