The UI observes — it never changes anything. It shows you what’s wrong and which
setting to adjust; you make the change with the tools you already use
(
tmo-cli, Helm, or the engine configuration). See
Access & Security for the full picture.What You Can See
Fleet at a Glance
Every engine and node, its readiness and health, and what needs attention right now.
Cache Health
How full each cache tier is, whether storage backends are healthy, and early warning of
problems.
Performance Over Time
Hit rate, eviction, capacity, and throughput trends — confirm a fix worked or catch a
problem building.
Workloads and GPUs
Per-workload cache budgets and who’s over them, plus per-GPU utilization and hardware
health.
Next Steps
Installation
Install the UI next to your operator.
Configuration
Connect metric history, GPU telemetry, and additional clusters.
Monitoring Your Fleet
Triage what needs attention, drill into nodes and GPUs, and track cache and
performance trends.
Thresholds & Notifications
Decide what counts as a problem, and get notified when it happens.

