datadog exporter.
Each node appears as a host in Datadog Infrastructure, tagged with your cluster name and
environment. The collector also serves LMCache metrics on port 8889, the same as the chart’s
collector in Observability, so Prometheus and the
Operator UI keep working.
Prerequisites
- Tensormesh Platform installed with Helm.
-
The OpenTelemetry Operator and the Prometheus Operator, as listed in
Observability → Prerequisites:
Two lines means both are installed.
- ServiceMonitors for the DCGM exporter, node-exporter, and kube-state-metrics. kube-prometheus-stack creates the node-exporter and kube-state-metrics ones. For DCGM, see Operator UI → Configuration.
- A Datadog API key saved in a file, and your Datadog site.
Step 1: Create the API key Secret
Create a namespace for the collector if you don’t have one, then store the key in it:Step 2: Deploy the collector
Save asdatadog-collector.yaml. It creates the collector, the permissions its target allocator
needs to read ServiceMonitors, and a ServiceMonitor for vLLM.
datadog-collector.yaml
If your Datadog site is not US1, change
site: to your site’s value.
filter/datadog list decides which metric names reach Datadog. Anything not on it is dropped.
Step 3: Choose what gets scraped
The target allocator picks up every ServiceMonitor labelledmonitoring: datadog. The vLLM one
in the file already has the label. Add it to your DCGM exporter, node-exporter, and
kube-state-metrics ServiceMonitors. The first command lists their names and namespaces:
monitoring: datadog on the port named
http, which is how the E2E Quickstart names it. Label each Service that
serves vLLM metrics, including any router in front of vLLM:
port: http in the file.
Step 4: Point the engine at the collector
Add these values to your values file:my-values.yaml
engine.spec.extraArgs, add these three entries to your list, replacing any
--enable-tracing or --otlp-endpoint entries already there. Then run the
Helm upgrade. The engine pods restart to pick up the new endpoint.
The engine sends OTLP to one collector. If your Prometheus reads LMCache metrics from the
chart’s collector, point that ServiceMonitor at the datadog-metrics-collector Service instead,
as described in Exporting metrics.
Verify metrics are flowing
Check the target allocator found your ServiceMonitors:otel-collector.
In Datadog, open Infrastructure → Hosts and filter on kube_cluster_name:<cluster-name>.
Each node appears as <node-name>-<cluster-name>. Then query one metric per source in
Metrics → Explorer, scoped to kube_cluster_name:<cluster-name>:
vLLM and LMCache metrics need traffic. Send some through vLLM first, as in the
E2E Quickstart, and allow a few minutes.
If the collector pods restart and their logs show
API Key validation failed, the key or the
site is wrong. Fix the Secret and the pods pick it up on their next restart, or fix site: and
apply the file again.
What you see in Datadog
LMCache names differ from Prometheus because Prometheus adds units and
_total to OTLP names and
Datadog does not. What each LMCache metric means is covered in Metrics.
Cost
Each node running the collector reports as a host in Datadog. The metrics it scrapes and receives count as Datadog custom metrics, once per unique combination of metric name and tag values. Histograms count more than once. See Datadog’s custom metrics billing.resource_attributes_as_tags: true turns resource attributes into tags, which adds
combinations. To send less, remove lines from filter/datadog and apply the file again. Check
the count on Datadog’s usage page after a day.
Rollback
Remove the threeextraArgs entries from your values file, put back any entries they replaced,
and run the Helm upgrade. If you pointed a Prometheus
ServiceMonitor at datadog-metrics-collector in step 4, point it back to the chart’s collector.
Then delete the collector and the Secret:
monitoring: datadog labels on your ServiceMonitors and vLLM Service do nothing once the
collector is gone. Remove them with the same kubectl label commands, ending the label in -.
