> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tensormesh.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Metrics

> The cache-engine metrics the operator exposes, and how to read the ones that matter.

Every cache-engine pod exports Prometheus metrics. This page covers how the operator exposes them
and the signals that show whether the cache is reducing recompute.

## How metrics reach Prometheus

Metrics reach Prometheus one of two ways.

The **OpenTelemetry Collector** path is the recommended setup. With `observability.enabled`, each
engine pushes its metrics (and traces) to an OTel Collector, and Prometheus scrapes the Collector.
Enabling it also stops the engine serving `/metrics` directly. See [Observability](/v1.0.0/observability)
for the full setup.

The **engine-endpoint** path is an alternative when you don't run the collector. Each
engine pod serves `/metrics` directly, and the operator creates a metrics `Service`, plus a
`ServiceMonitor` (when `serviceMonitor.enabled` is set) so Prometheus auto-discovers it:

```yaml theme={null}
engine:
  enabled: true
  spec:
    prometheus:
      enabled: true
      port: 9090
      serviceMonitor:
        enabled: true
        labels:
          release: kube-prometheus-stack   # must match your Prometheus' serviceMonitorSelector
```

## Verify metrics are flowing

First, find which path this engine uses. The check is whether the engine is pushing OTLP:

```bash theme={null}
kubectl get lmcacheengine <engine-name> -n <operator-namespace> -o yaml | grep -q otlp-endpoint \
  && echo "OTLP / collector path" || echo "engine-endpoint path"
```

<Tabs>
  <Tab title="OTLP / collector path">
    The engine pushes metrics to an OpenTelemetry Collector, so it exposes no scrapeable
    `lmcache_mp_*` endpoint itself. Read them from the Collector's Prometheus exporter and confirm
    the pipeline in
    [Observability → Verify telemetry is flowing](/v1.0.0/observability#verify-telemetry-is-flowing), or
    just query Prometheus below.
  </Tab>

  <Tab title="Engine-endpoint path">
    The engine serves `/metrics` on its `--prometheus-port` (default `9090`), behind the
    `<engine-name>-metrics` Service:

    ```bash theme={null}
    kubectl get endpoints <engine-name>-metrics -n <operator-namespace>   # one IP per engine pod
    kubectl port-forward -n <operator-namespace> svc/<engine-name>-metrics 9090:9090
    curl -s http://localhost:9090/metrics | grep lmcache_mp_ | head
    ```

    `connection refused` means nothing is serving `/metrics` there; confirm `prometheus.enabled`
    is set and the engine isn't on the OTLP path.
  </Tab>
</Tabs>

The fleet-wide source of truth is your Prometheus, the same data the
[Operator UI](/v1.0.0/ui/introduction) reads:

```bash theme={null}
kubectl port-forward -n <prometheus-namespace> svc/<prometheus> 9091:9090
curl -s 'http://localhost:9091/api/v1/query?query=sum(rate(lmcache_mp_lookup_hit_tokens_total[5m]))/sum(rate(lmcache_mp_lookup_requested_tokens_total[5m]))' \
  | jq -r '.data.result[0].value[1]'
```

A batch of `lmcache_mp_*` lines (or a number from the query) confirms metrics are flowing. Empty
means no traffic has hit the engine yet; send some inference through vLLM first.

## Pull every metric

See the **complete** set straight from Prometheus (works on either path):

```bash theme={null}
# every cache-engine metric name Prometheus holds
curl -s 'http://<prometheus>/api/v1/label/__name__/values' \
  | jq -r '.data[] | select(startswith("lmcache_mp_"))'

# the labels you can slice by (model, tenant, backend, node)
curl -s 'http://<prometheus>/api/v1/series?match[]=lmcache_mp_lookup_hit_tokens_total' \
  | jq '.data[0] | keys'
```

The **[complete metric reference](#complete-metric-reference)** below catalogs every one.

## What to watch first

The signals to check first, in priority order.

### 1. Lookup hit rate

How much of each requested prefix was served from cache instead of recomputed on the GPU.

```promql theme={null}
sum(rate(lmcache_mp_lookup_hit_tokens_total[5m]))
/ sum(rate(lmcache_mp_lookup_requested_tokens_total[5m]))
```

| Metric                                     | Meaning                                |
| ------------------------------------------ | -------------------------------------- |
| `lmcache_mp_lookup_requested_tokens_total` | Tokens the engine was asked to look up |
| `lmcache_mp_lookup_hit_tokens_total`       | Tokens found in cache                  |

### 2. L1 activity

L1 is the [host-DRAM tier](/v1.0.0/configuration/cpu-offloading). These counters show whether it's
serving reads (good) or just churning writes and evictions (undersized).

| Metric                               | Meaning                       |
| ------------------------------------ | ----------------------------- |
| `lmcache_mp_l1_read_chunks_total`    | Chunks served from DRAM       |
| `lmcache_mp_l1_write_chunks_total`   | Chunks written into DRAM      |
| `lmcache_mp_l1_evicted_chunks_total` | Chunks evicted under pressure |

* Reads rising on warm traffic → L1 is doing its job.
* Writes high, reads low → poor reuse, or the hot set doesn't fit — see [CPU offloading sizing](/v1.0.0/configuration/cpu-offloading#sizing).
* Evictions climbing fast → L1 is undersized for the working set.

### 3. L2 activity

Only relevant if you've configured [external storage](/v1.0.0/configuration/filesystem-offloading).
These confirm data is landing in L2 and warm traffic is loading it back.

| Metric                                         | Meaning                                      |
| ---------------------------------------------- | -------------------------------------------- |
| `lmcache_mp_l2_store_completed_requests_total` | Data written to L2                           |
| `lmcache_mp_l2_load_completed_requests_total`  | Data loaded back from L2                     |
| `lmcache_mp_l2_prefetch_failure_chunks_total`  | Prefetches that failed (miss or L1 pressure) |

### 4. Retrieved chunk volume

```promql theme={null}
sum(rate(lmcache_mp_num_chunks_loaded_total[5m]))
```

How much data the engine loads back into service. If hit rate is rising but this stays flat, the
warm path isn't actually loading anything.

## Slice by model and tenant

Many engine metrics carry labels you can group by:

* `model_name` — per-model hit rate and volume
* `cache_salt` — per-tenant isolation
* `l2_name` — which L2 backend

```promql theme={null}
sum by (model_name) (rate(lmcache_mp_lookup_hit_tokens_total[5m]))
/ sum by (model_name) (rate(lmcache_mp_lookup_requested_tokens_total[5m]))
```

Keep `cache_salt` coarse: one salt per tenant is fine, but one per request explodes your
time-series count and hurts cache reuse.

## Complete metric reference

Most metrics carry `model_name`, `cache_salt`, and `l2_name` labels for
[slicing](#slice-by-model-and-tenant). Counters (`_total`) need `rate(...[5m])`; gauges are
instantaneous; histograms expose `_bucket` / `_sum` / `_count`.

A metric only appears once it's been recorded, so an idle engine shows only the always-on gauges;
the activity counters appear after traffic hits the cache.

### Lookup & hit rate

| Metric                                     | Type    | Measures                                                      |
| ------------------------------------------ | ------- | ------------------------------------------------------------- |
| `lmcache_mp_lookup_requested_tokens_total` | counter | Tokens the engine was asked to look up (hit-rate denominator) |
| `lmcache_mp_lookup_hit_tokens_total`       | counter | Tokens found in cache (hit-rate numerator)                    |

### L1 — CPU offloading (host DRAM)

| Metric                                        | Type    | Measures                                               |
| --------------------------------------------- | ------- | ------------------------------------------------------ |
| `lmcache_mp_l1_usage_ratio`                   | gauge   | L1 fill level, `0..1`                                  |
| `lmcache_mp_l1_memory_usage_bytes`            | gauge   | Bytes currently held in L1                             |
| `lmcache_mp_l1_read_chunks_total`             | counter | Chunks served from L1 (reuse activity)                 |
| `lmcache_mp_l1_write_chunks_total`            | counter | Chunks written into L1 (ingest)                        |
| `lmcache_mp_l1_evicted_chunks_total`          | counter | Chunks evicted from L1 under pressure                  |
| `lmcache_mp_l1_read_failure_chunks_total`     | counter | L1 reads that failed                                   |
| `lmcache_mp_l1_eviction_loop_ticks_total`     | counter | Eviction-loop iterations (every cycle)                 |
| `lmcache_mp_l1_eviction_loop_triggered_total` | counter | Eviction-loop iterations where the policy actually ran |

### L2 — external storage

| Metric                                                       | Type    | Measures                                                                |
| ------------------------------------------------------------ | ------- | ----------------------------------------------------------------------- |
| `lmcache_mp_l2_usage_bytes`                                  | gauge   | Bytes currently held in L2, per `l2_name`                               |
| `lmcache_mp_l2_adapters`                                     | gauge   | L2 adapters on the store controller, by `state` (`active` / `draining`) |
| `lmcache_mp_num_inflight_l2_stores`                          | gauge   | Store operations in flight to L2                                        |
| `lmcache_mp_num_inflight_l2_loads`                           | gauge   | Load operations in flight from L2                                       |
| `lmcache_mp_inflight_load_memory_usage_bytes`                | gauge   | Memory held by in-flight L2 loads                                       |
| `lmcache_mp_l2_store_submitted_requests_total`               | counter | L2 store requests submitted                                             |
| `lmcache_mp_l2_store_submitted_objects_chunks_total`         | counter | Chunks submitted for L2 store                                           |
| `lmcache_mp_l2_store_completed_requests_total`               | counter | L2 store requests completed                                             |
| `lmcache_mp_l2_store_completed_objects_chunks_total`         | counter | Chunks written to L2                                                    |
| `lmcache_mp_l2_load_completed_requests_total`                | counter | L2 load requests completed                                              |
| `lmcache_mp_l2_prefetch_lookup_requests_total`               | counter | L2 prefetch lookup requests submitted                                   |
| `lmcache_mp_l2_prefetch_lookup_objects_chunks_total`         | counter | Chunks submitted for L2 prefetch lookup                                 |
| `lmcache_mp_l2_prefetch_hit_chunks_total`                    | counter | Prefix chunks found in an L2 lookup                                     |
| `lmcache_mp_l2_prefetch_load_submitted_requests_total`       | counter | L2 prefetch load requests submitted                                     |
| `lmcache_mp_l2_prefetch_load_submitted_objects_chunks_total` | counter | Chunks submitted for L2 load                                            |
| `lmcache_mp_l2_prefetch_load_completed_chunks_total`         | counter | Chunks successfully loaded from L2                                      |
| `lmcache_mp_l2_prefetch_failure_chunks_total`                | counter | Prefetches that failed (`reason` = `l1_oom` / `not_found`)              |
| `lmcache_mp_num_chunks_loaded_total`                         | counter | Chunks loaded back into service (the warm path)                         |

### Throughput

Histograms in **GB/s**: compute the live rate as `rate(<name>_sum[5m]) / rate(<name>_count[5m])`.

| Metric                                            | Direction                    |
| ------------------------------------------------- | ---------------------------- |
| `lmcache_mp_l0_l1_store_throughput_GB_per_second` | GPU → CPU (store into L1)    |
| `lmcache_mp_l0_l1_load_throughput_GB_per_second`  | CPU → GPU (load from L1)     |
| `lmcache_mp_l2_store_throughput_GB_per_second`    | L1 → L2 (store to external)  |
| `lmcache_mp_l2_load_throughput_GB_per_second`     | L2 → L1 (load from external) |

### Prefetch & event bus (health)

| Metric                                      | Type    | Measures                                                                           |
| ------------------------------------------- | ------- | ---------------------------------------------------------------------------------- |
| `lmcache_mp_active_prefetch_jobs`           | gauge   | Prefetch jobs currently running                                                    |
| `lmcache_mp_active_p2p_lookup_jobs`         | gauge   | Active peer-to-peer lookup jobs (with [P2P](/v1.0.0/configuration/p2p) enabled)    |
| `lmcache_mp_event_bus_queue_depth`          | gauge   | Depth of the internal event queue                                                  |
| `lmcache_mp_event_bus_drain_lag_seconds`    | gauge   | Seconds since the oldest queued event; rising = the drain thread is falling behind |
| `lmcache_mp_event_bus_dropped_events_total` | counter | Events dropped (queue overflow, a saturation signal)                               |

### Reuse gap

How long cached data lives before it's reused, the input to sizing decisions.

| Metric                              | Type      | Measures                                     |
| ----------------------------------- | --------- | -------------------------------------------- |
| `lmcache_mp_real_reuse_gap_seconds` | histogram | Time between storing a chunk and reusing it  |
| `lmcache_mp_real_reuse_gap_objects` | histogram | Reuse-gap distribution across cached objects |

### Lifecycle histograms

Per-tier block/chunk timing: how long cached data lives, sits idle, and waits between reuses.
Newer engine builds expose more of these. All are histograms (`_bucket` / `_sum` / `_count`).

| Metric                                          | Tier      | Measures                                   |
| ----------------------------------------------- | --------- | ------------------------------------------ |
| `lmcache_mp_l0_block_lifetime_seconds`          | GPU (L0)  | Block lifetime, allocation → eviction      |
| `lmcache_mp_l0_block_idle_before_evict_seconds` | GPU (L0)  | Idle time before a block is evicted        |
| `lmcache_mp_l0_block_reuse_gap_seconds`         | GPU (L0)  | Gap between consecutive block accesses     |
| `lmcache_mp_l1_chunk_lifetime_seconds`          | L1 (DRAM) | Chunk lifetime, allocation → eviction      |
| `lmcache_mp_l1_chunk_idle_before_evict_seconds` | L1 (DRAM) | Idle time before a chunk is evicted        |
| `lmcache_mp_l1_chunk_reuse_gap_seconds`         | L1 (DRAM) | Gap between consecutive touches of a chunk |
| `lmcache_mp_l1_chunk_evict_reuse_gap_seconds`   | L1 (DRAM) | Time from chunk eviction to next reuse     |

vLLM exports its own `vllm:*` metrics on a separate port (TTFT, running/waiting requests, HBM KV
usage). Pair those with the cache-engine metrics above for the full request-to-cache picture.

## When to go deeper

The catalog above is complete for operating the cache. For low-level internals (exact
histogram-bucket boundaries, precise semantics, and cardinality notes), see the upstream
multiprocess-engine reference: [MP observability](https://docs.lmcache.ai/mp/observability.html).

## Next steps

<CardGroup cols={3}>
  <Card title="Operator UI" icon="browser" href="/v1.0.0/ui/introduction">
    An interface that visualizes these metrics: fleet, hit rate, capacity, health.
  </Card>

  <Card title="Observability stack" icon="chart-line" href="/v1.0.0/observability">
    Ship metrics and traces through an OpenTelemetry Collector to Prometheus and Tempo.
  </Card>

  <Card title="CPU offloading" icon="memory" href="/v1.0.0/configuration/cpu-offloading">
    Tune the L1 tier the L1 metrics above describe.
  </Card>
</CardGroup>
