> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tensormesh.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Export to Datadog

> Send vLLM, GPU, node, Kubernetes, and LMCache metrics to Datadog Infrastructure through an OpenTelemetry Collector.

This page deploys an **OpenTelemetry Collector for Datadog**. It runs one pod per node, scrapes
vLLM, the DCGM exporter, node-exporter, and kube-state-metrics, receives LMCache metrics from the
engine over OTLP, reads node CPU, memory, disk, and network, and sends all of it to Datadog
through the `datadog` exporter.

Each node appears as a host in Datadog Infrastructure, tagged with your cluster name and
environment. The collector also serves LMCache metrics on port `8889`, the same as the chart's
collector in [Observability](/observability), so Prometheus and the
[Operator UI](/ui/introduction) keep working.

## Prerequisites

* Tensormesh Platform installed with [Helm](/installation/helm).
* The OpenTelemetry Operator and the Prometheus Operator, as listed in
  [Observability → Prerequisites](/observability#prerequisites):

  ```bash theme={null} theme={null}
  kubectl get crd | grep -E 'opentelemetrycollectors|servicemonitors.monitoring'
  ```

  Two lines means both are installed.
* ServiceMonitors for the DCGM exporter, node-exporter, and kube-state-metrics.
  kube-prometheus-stack creates the node-exporter and kube-state-metrics ones. For DCGM, see
  [Operator UI → Configuration](/ui/configuration).
* A Datadog API key saved in a file, and your
  [Datadog site](https://docs.datadoghq.com/getting_started/site/).

## Step 1: Create the API key Secret

Create a namespace for the collector if you don't have one, then store the key in it:

```bash theme={null} theme={null}
kubectl create namespace <collector-namespace>

kubectl -n <collector-namespace> create secret generic datadog-api-key \
  --from-file=api-key=<path-to-key-file> \
  --dry-run=client -o yaml | kubectl apply --server-side -f -
```

The collector reads the key from this Secret when it starts, so it never appears in the manifest.
To rotate the key, run the second command again with the new file, then restart the collector:

```bash theme={null} theme={null}
kubectl -n <collector-namespace> rollout restart ds/datadog-metrics-collector
```

## Step 2: Deploy the collector

Save as `datadog-collector.yaml`. It creates the collector, the permissions its target allocator
needs to read ServiceMonitors, and a ServiceMonitor for vLLM.

```yaml datadog-collector.yaml theme={null} theme={null}
apiVersion: v1
kind: ServiceAccount
metadata:
  name: datadog-metrics-targetallocator
  namespace: <collector-namespace>
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: datadog-metrics-targetallocator
rules:
  - apiGroups: [""]
    resources: [nodes, nodes/metrics, services, endpoints, pods, namespaces]
    verbs: [get, list, watch]
  - apiGroups: [""]
    resources: [configmaps, secrets]
    verbs: [get]
  - apiGroups: [discovery.k8s.io]
    resources: [endpointslices]
    verbs: [get, list, watch]
  - apiGroups: [networking.k8s.io]
    resources: [ingresses]
    verbs: [get, list, watch]
  - apiGroups: [monitoring.coreos.com]
    resources: [servicemonitors, podmonitors, scrapeconfigs, probes]
    verbs: ["*"]
  - nonResourceURLs: [/metrics, /api, /api/*, /apis, /apis/*]
    verbs: [get]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: datadog-metrics-targetallocator
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: datadog-metrics-targetallocator
subjects:
  - kind: ServiceAccount
    name: datadog-metrics-targetallocator
    namespace: <collector-namespace>
---
apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
  name: datadog-metrics
  namespace: <collector-namespace>
spec:
  mode: daemonset
  image: otel/opentelemetry-collector-contrib:0.159.0

  env:
    - name: DD_API_KEY
      valueFrom:
        secretKeyRef:
          name: datadog-api-key
          key: api-key
    - name: NODE_NAME
      valueFrom:
        fieldRef:
          fieldPath: spec.nodeName

  # Waits until the target allocator answers, so the first job-list request does not fail.
  initContainers:
    - name: wait-for-target-allocator
      image: busybox:1.36
      command:
        - sh
        - -c
        - for i in $(seq 120); do wget -q -O /dev/null http://datadog-metrics-targetallocator/scrape_configs && exit 0; sleep 1; done; exit 0

  volumes:
    - name: hostfs
      hostPath:
        path: /
        type: Directory
  volumeMounts:
    - name: hostfs
      mountPath: /hostfs
      readOnly: true
      mountPropagation: HostToContainer

  targetAllocator:
    enabled: true
    allocationStrategy: per-node
    serviceAccount: datadog-metrics-targetallocator
    prometheusCR:
      enabled: true
      scrapeInterval: 30s
      serviceMonitorNamespaceSelector: {}
      serviceMonitorSelector:
        matchLabels:
          monitoring: datadog
      podMonitorNamespaceSelector: {}
      podMonitorSelector:
        matchLabels:
          monitoring: datadog

  config:
    receivers:
      otlp:
        protocols:
          grpc:
            endpoint: 0.0.0.0:4317
          http:
            endpoint: 0.0.0.0:4318
      prometheus:
        config:
          scrape_configs:
            - job_name: otel-collector
              scrape_interval: 10s
              static_configs:
                - targets: ["0.0.0.0:8888"]
      host_metrics:
        collection_interval: 10s
        root_path: /hostfs
        scrapers:
          cpu:
            metrics:
              system.cpu.utilization:
                enabled: true
              system.cpu.time:
                enabled: false
          memory: {}
          load: {}
          paging:
            metrics:
              system.paging.utilization:
                enabled: true
          disk: {}
          network: {}
          processes: {}
          filesystem:
            metrics:
              system.filesystem.utilization:
                enabled: true
            exclude_fs_types:
              fs_types: [autofs, binfmt_misc, bpf, cgroup, cgroup2, configfs, debugfs, devpts, devtmpfs, fusectl, hugetlbfs, iso9660, mqueue, nsfs, overlay, proc, procfs, pstore, securityfs, selinuxfs, squashfs, sysfs, tracefs, tmpfs]
              match_type: strict
            exclude_mount_points:
              mount_points: [/var/lib/kubelet/.*, /var/lib/containers/.*, /run/.*, /proc/.*, /sys/.*, /dev/.*]
              match_type: regexp

    processors:
      memory_limiter:
        check_interval: 1s
        limit_percentage: 75
        spike_limit_percentage: 15
      filter/datadog:
        error_mode: ignore
        metrics:
          include:
            match_type: regexp
            metric_names:
              - ^up$
              - ^lmcache.*
              - ^system\..*
              - ^node_memory_MemAvailable_bytes$
              - ^node_memory_MemTotal_bytes$
              - ^node_load1$
              - ^node_load5$
              - ^node_load15$
              - ^node_network_receive_bytes_total$
              - ^node_network_transmit_bytes_total$
              - ^node_uname_info$
              - ^kube_pod_container_status_restarts_total$
              - ^kube_pod_container_status_waiting_reason$
              - ^kube_pod_container_status_terminated_reason$
              - ^kube_pod_status_phase$
              - ^kube_pod_status_ready$
              - ^kube_deployment_status_replicas$
              - ^kube_deployment_status_replicas_available$
              - ^kube_deployment_status_replicas_updated$
              - ^kube_deployment_spec_replicas$
              - ^kube_statefulset_status_replicas$
              - ^kube_statefulset_status_replicas_ready$
              - ^kube_daemonset_status_number_ready$
              - ^kube_daemonset_status_desired_number_scheduled$
              - ^otelcol_.*
              - ^vllm:avg_ttft$
              - ^vllm:generation_tokens_total$
              - ^vllm:input_throughput$
              - ^vllm:inter_token_latency_seconds$
              - ^vllm:kv_cache_usage_perc$
              - ^vllm:num_incoming_requests_total$
              - ^vllm:num_requests_running$
              - ^vllm:output_throughput$
              - ^vllm:prompt_tokens_total$
              - ^vllm:time_to_first_token_seconds$
              - ^vllm:e2e_request_latency_seconds$
              - ^vllm:external_prefix_cache_hits_total$
              - ^vllm:external_prefix_cache_queries_total$
              - ^vllm:prefix_cache_hits_total$
              - ^vllm:prefix_cache_queries_total$
              - ^vllm:request_success_total$
              - ^vllm:num_requests_waiting$
              - ^vllm:request_queue_time_seconds$
              - ^vllm:num_requests_waiting_by_reason$
              - ^vllm:request_inference_time_seconds$
              - ^vllm:request_prefill_time_seconds$
              - ^vllm:request_decode_time_seconds$
              - ^vllm:iteration_tokens_total$
              - ^vllm:current_qps$
              - ^vllm:avg_latency$
              - ^vllm:avg_itl$
              - ^vllm:num_prefill_requests$
              - ^vllm:num_decoding_requests$
              - ^vllm:num_preemptions_total$
              - ^vllm:request_prompt_tokens$
              - ^vllm:request_generation_tokens$
              - ^vllm:healthy_pods_total$
              - ^DCGM_FI_DEV_FB_FREE$
              - ^DCGM_FI_DEV_FB_USED$
              - ^DCGM_FI_DEV_FB_TOTAL$
              - ^DCGM_FI_DEV_GPU_UTIL$
              - ^DCGM_FI_DEV_MEM_COPY_UTIL$
              - ^DCGM_FI_DEV_GPU_TEMP$
              - ^DCGM_FI_DEV_MEMORY_TEMP$
              - ^DCGM_FI_DEV_POWER_USAGE$
              - ^DCGM_FI_DEV_XID_ERRORS$
              - ^DCGM_FI_DEV_ECC_SBE_VOL_TOTAL_total$
              - ^DCGM_FI_DEV_ECC_DBE_VOL_TOTAL_total$
              - ^DCGM_FI_DEV_ECC_SBE_AGG_TOTAL_total$
              - ^DCGM_FI_DEV_ECC_DBE_AGG_TOTAL_total$
              - ^DCGM_FI_DEV_THERMAL_VIOLATION_total$
              - ^DCGM_FI_DEV_POWER_VIOLATION_total$
              - ^DCGM_FI_DEV_CLOCK_THROTTLE_REASONS_total$
              - ^DCGM_FI_DEV_CORRECTABLE_REMAPPED_ROWS_total$
              - ^DCGM_FI_DEV_UNCORRECTABLE_REMAPPED_ROWS_total$
              - ^DCGM_FI_DEV_ROW_REMAP_FAILURE$
              - ^DCGM_FI_PROF_SM_ACTIVE$
              - ^DCGM_FI_PROF_GR_ENGINE_ACTIVE$
              - ^DCGM_FI_PROF_PIPE_TENSOR_ACTIVE$
              - ^DCGM_FI_PROF_DRAM_ACTIVE$
      metrics_transform/system_cpu:
        transforms:
          - include: system.cpu.utilization
            match_type: strict
            action: update
            operations:
              - action: aggregate_labels
                label_set: [state]
                aggregation_type: mean
      resource/datadog:
        attributes:
          - key: k8s.cluster.name
            value: <cluster-name>
            action: insert
          - key: deployment.environment
            value: <env>
            action: upsert
          - key: datadog.host.name
            value: ${env:NODE_NAME}-<cluster-name>
            action: upsert
      batch: {}

    exporters:
      prometheus:
        endpoint: 0.0.0.0:8889
      datadog:
        hostname: ${env:NODE_NAME}-<cluster-name>
        api:
          key: ${env:DD_API_KEY}
          site: datadoghq.com        # your Datadog site
          fail_on_invalid_key: true
        host_metadata:
          enabled: true
          hostname_source: config_or_system
          tags:
            - env:<env>
            - kube_cluster_name:<cluster-name>
        metrics:
          resource_attributes_as_tags: true
          histograms:
            send_aggregation_metrics: true

    service:
      pipelines:
        metrics:
          receivers: [prometheus, otlp, host_metrics]
          processors: [memory_limiter, filter/datadog, metrics_transform/system_cpu, resource/datadog, batch]
          exporters: [datadog]
        metrics/prometheus:
          receivers: [otlp]
          processors: [memory_limiter, batch]
          exporters: [prometheus]
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: vllm
  namespace: <collector-namespace>
  labels:
    monitoring: datadog
spec:
  namespaceSelector:
    any: true
  selector:
    matchLabels:
      monitoring: datadog
  endpoints:
    - port: http
```

Replace the placeholders:

| Placeholder             | Value                                                                                    |
| ----------------------- | ---------------------------------------------------------------------------------------- |
| `<collector-namespace>` | The namespace from step 1                                                                |
| `<cluster-name>`        | A name for this cluster. It becomes the `kube_cluster_name` tag and the host name suffix |
| `<env>`                 | The Datadog `env` tag, for example `prod`                                                |

If your Datadog site is not US1, change `site:` to your site's value.

```bash theme={null} theme={null}
kubectl apply -f datadog-collector.yaml
kubectl -n <collector-namespace> rollout status ds/datadog-metrics-collector --timeout=180s
kubectl -n <collector-namespace> rollout status deploy/datadog-metrics-targetallocator --timeout=180s
```

The `filter/datadog` list decides which metric names reach Datadog. Anything not on it is dropped.

## Step 3: Choose what gets scraped

The target allocator picks up every ServiceMonitor labelled `monitoring: datadog`. The vLLM one
in the file already has the label. Add it to your DCGM exporter, node-exporter, and
kube-state-metrics ServiceMonitors. The first command lists their names and namespaces:

```bash theme={null} theme={null}
kubectl get servicemonitor -A
kubectl -n <namespace> label servicemonitor <servicemonitor-name> monitoring=datadog
```

The vLLM ServiceMonitor scrapes every Service labelled `monitoring: datadog` on the port named
`http`, which is how the [E2E Quickstart](/installation/example) names it. Label each Service that
serves vLLM metrics, including any router in front of vLLM:

```bash theme={null} theme={null}
kubectl -n <vllm-namespace> label service <vllm-service> monitoring=datadog
```

If your vLLM Service names its port differently, change `port: http` in the file.

## Step 4: Point the engine at the collector

Add these values to your values file:

```yaml my-values.yaml theme={null} theme={null}
engine:
  spec:
    extraArgs:
      - --enable-tracing
      - --otlp-endpoint
      - http://datadog-metrics-collector.<collector-namespace>.svc:4317
```

If you already set `engine.spec.extraArgs`, add these three entries to your list, replacing any
`--enable-tracing` or `--otlp-endpoint` entries already there. Then run the
[Helm upgrade](/installation/helm#upgrade). The engine pods restart to pick up the new endpoint.

The engine sends OTLP to one collector. If your Prometheus reads LMCache metrics from the
chart's collector, point that ServiceMonitor at the `datadog-metrics-collector` Service instead,
as described in [Exporting metrics](/observability#exporting-metrics).

## Verify metrics are flowing

Check the target allocator found your ServiceMonitors:

```bash theme={null} theme={null}
kubectl -n <collector-namespace> port-forward svc/datadog-metrics-targetallocator 8080:80
```

In a second terminal:

```bash theme={null} theme={null}
curl -s localhost:8080/jobs | jq -r 'keys[]'
```

Expect one job per labelled ServiceMonitor, plus `otel-collector`.

In Datadog, open **Infrastructure → Hosts** and filter on `kube_cluster_name:<cluster-name>`.
Each node appears as `<node-name>-<cluster-name>`. Then query one metric per source in
**Metrics → Explorer**, scoped to `kube_cluster_name:<cluster-name>`:

| Source     | Metric                      |
| ---------- | --------------------------- |
| Node       | `system.cpu.utilization`    |
| GPU        | `DCGM_FI_DEV_GPU_UTIL`      |
| Kubernetes | `kube_pod_status_ready`     |
| vLLM       | `vllm_num_requests_running` |
| LMCache    | `lmcache_mp.l1_usage_ratio` |

vLLM and LMCache metrics need traffic. Send some through vLLM first, as in the
[E2E Quickstart](/installation/example), and allow a few minutes.

If the collector pods restart and their logs show `API Key validation failed`, the key or the
site is wrong. Fix the Secret and the pods pick it up on their next restart, or fix `site:` and
apply the file again.

## What you see in Datadog

| Source                                         | Name in Datadog                                                                                                 |
| ---------------------------------------------- | --------------------------------------------------------------------------------------------------------------- |
| Node (`host_metrics`)                          | `system.*`                                                                                                      |
| GPU (DCGM exporter)                            | Unchanged, for example `DCGM_FI_DEV_GPU_UTIL`                                                                   |
| Kubernetes (kube-state-metrics, node-exporter) | Unchanged, for example `kube_pod_status_ready`, `node_load1`                                                    |
| vLLM                                           | `vllm:` becomes `vllm_`. Histograms add `.sum` and `.count`, for example `vllm_time_to_first_token_seconds.sum` |
| LMCache (engine OTLP)                          | `lmcache_mp.*`, for example `lmcache_mp.l1_read` for `lmcache_mp_l1_read_chunks_total`                          |

LMCache names differ from Prometheus because Prometheus adds units and `_total` to OTLP names and
Datadog does not. What each LMCache metric means is covered in [Metrics](/observability/metrics).

## Cost

Each node running the collector reports as a host in Datadog. The metrics it scrapes and
receives count as Datadog custom metrics, once per unique combination of metric name and tag
values. Histograms count more than once. See Datadog's
[custom metrics billing](https://docs.datadoghq.com/account_management/billing/custom_metrics/).

`resource_attributes_as_tags: true` turns resource attributes into tags, which adds
combinations. To send less, remove lines from `filter/datadog` and apply the file again. Check
the count on Datadog's usage page after a day.

## Rollback

Remove the three `extraArgs` entries from your values file, put back any entries they replaced,
and run the [Helm upgrade](/installation/helm#upgrade). If you pointed a Prometheus
ServiceMonitor at `datadog-metrics-collector` in step 4, point it back to the chart's collector.
Then delete the collector and the Secret:

```bash theme={null} theme={null}
kubectl -n <collector-namespace> delete opentelemetrycollector datadog-metrics
kubectl delete -f datadog-collector.yaml --ignore-not-found
kubectl -n <collector-namespace> delete secret datadog-api-key
```

The `monitoring: datadog` labels on your ServiceMonitors and vLLM Service do nothing once the
collector is gone. Remove them with the same `kubectl label` commands, ending the label in `-`.
