Skip to main content
This page deploys an OpenTelemetry Collector for Datadog. It runs one pod per node, scrapes vLLM, the DCGM exporter, node-exporter, and kube-state-metrics, receives LMCache metrics from the engine over OTLP, reads node CPU, memory, disk, and network, and sends all of it to Datadog through the datadog exporter. Each node appears as a host in Datadog Infrastructure, tagged with your cluster name and environment. The collector also serves LMCache metrics on port 8889, the same as the chart’s collector in Observability, so Prometheus and the Operator UI keep working.

Prerequisites

  • Tensormesh Platform installed with Helm.
  • The OpenTelemetry Operator and the Prometheus Operator, as listed in Observability → Prerequisites:
    Two lines means both are installed.
  • ServiceMonitors for the DCGM exporter, node-exporter, and kube-state-metrics. kube-prometheus-stack creates the node-exporter and kube-state-metrics ones. For DCGM, see Operator UI → Configuration.
  • A Datadog API key saved in a file, and your Datadog site.

Step 1: Create the API key Secret

Create a namespace for the collector if you don’t have one, then store the key in it:
The collector reads the key from this Secret when it starts, so it never appears in the manifest. To rotate the key, run the second command again with the new file, then restart the collector:

Step 2: Deploy the collector

Save as datadog-collector.yaml. It creates the collector, the permissions its target allocator needs to read ServiceMonitors, and a ServiceMonitor for vLLM.
datadog-collector.yaml
Replace the placeholders: If your Datadog site is not US1, change site: to your site’s value.
The filter/datadog list decides which metric names reach Datadog. Anything not on it is dropped.

Step 3: Choose what gets scraped

The target allocator picks up every ServiceMonitor labelled monitoring: datadog. The vLLM one in the file already has the label. Add it to your DCGM exporter, node-exporter, and kube-state-metrics ServiceMonitors. The first command lists their names and namespaces:
The vLLM ServiceMonitor scrapes every Service labelled monitoring: datadog on the port named http, which is how the E2E Quickstart names it. Label each Service that serves vLLM metrics, including any router in front of vLLM:
If your vLLM Service names its port differently, change port: http in the file.

Step 4: Point the engine at the collector

Add these values to your values file:
my-values.yaml
If you already set engine.spec.extraArgs, add these three entries to your list, replacing any --enable-tracing or --otlp-endpoint entries already there. Then run the Helm upgrade. The engine pods restart to pick up the new endpoint. The engine sends OTLP to one collector. If your Prometheus reads LMCache metrics from the chart’s collector, point that ServiceMonitor at the datadog-metrics-collector Service instead, as described in Exporting metrics.

Verify metrics are flowing

Check the target allocator found your ServiceMonitors:
In a second terminal:
Expect one job per labelled ServiceMonitor, plus otel-collector. In Datadog, open Infrastructure → Hosts and filter on kube_cluster_name:<cluster-name>. Each node appears as <node-name>-<cluster-name>. Then query one metric per source in Metrics → Explorer, scoped to kube_cluster_name:<cluster-name>: vLLM and LMCache metrics need traffic. Send some through vLLM first, as in the E2E Quickstart, and allow a few minutes. If the collector pods restart and their logs show API Key validation failed, the key or the site is wrong. Fix the Secret and the pods pick it up on their next restart, or fix site: and apply the file again.

What you see in Datadog

LMCache names differ from Prometheus because Prometheus adds units and _total to OTLP names and Datadog does not. What each LMCache metric means is covered in Metrics.

Cost

Each node running the collector reports as a host in Datadog. The metrics it scrapes and receives count as Datadog custom metrics, once per unique combination of metric name and tag values. Histograms count more than once. See Datadog’s custom metrics billing. resource_attributes_as_tags: true turns resource attributes into tags, which adds combinations. To send less, remove lines from filter/datadog and apply the file again. Check the count on Datadog’s usage page after a day.

Rollback

Remove the three extraArgs entries from your values file, put back any entries they replaced, and run the Helm upgrade. If you pointed a Prometheus ServiceMonitor at datadog-metrics-collector in step 4, point it back to the chart’s collector. Then delete the collector and the Secret:
The monitoring: datadog labels on your ServiceMonitors and vLLM Service do nothing once the collector is gone. Remove them with the same kubectl label commands, ending the label in -.