Skip to content

NVIDIA GPUs

The collector monitors NVIDIA GPUs via NVML using the go-nvml library, which loads libnvidia-ml.so at runtime. No CUDA toolkit or DCGM daemon is needed.

  • Linux with NVIDIA GPU drivers installed
  • libnvidia-ml.so present on the host (installed with the NVIDIA driver)
  • For Docker: NVIDIA Container Toolkit
MetricDescription
hw.gpu.utilizationCompute, encoder, and decoder utilization (0.0–1.0) via hw.gpu.task
hw.gpu.memory.utilizationFraction of GPU memory used (usage / limit)
hw.gpu.memory.controller.utilizationMemory controller busy fraction (extension)
hw.gpu.memory.limitTotal VRAM (bytes)
hw.gpu.memory.usageUsed VRAM (bytes)
hw.gpu.memory.freeFree VRAM (bytes)
hw.temperatureDie and memory temperature (°C) via hw.sensor_location
hw.fan.speed_ratioFan speed as fraction of max (NVML %; rpm not available)
hw.powerCurrent power draw (W)
hw.power.limitPower cap (W)
hw.energyCumulative energy (J)
hw.gpu.speedClock frequency in Hz (hw.gpu.clock_domain=graphics/sm/memory)
hw.statusHardware status (ok / degraded / failed)
hw.errorsECC correctable/uncorrectable errors and PCIe replay errors
Terminal window
docker run -d \
--name otel-gpu-collector \
--gpus all \
--pid=host \
-e OTEL_SERVICE_NAME=my-app \
-e OTEL_RESOURCE_ATTRIBUTES='deployment.environment=production' \
-e OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318 \
ghcr.io/thinkfleetai/otel-gpu-collector:latest

--pid=host is required for per-process GPU attribution (cmdline, PID, zombie state).

services:
otel-gpu-collector:
image: ghcr.io/thinkfleetai/otel-gpu-collector:latest
pid: host
environment:
OTEL_SERVICE_NAME: my-app
OTEL_RESOURCE_ATTRIBUTES: deployment.environment=production
OTEL_EXPORTER_OTLP_ENDPOINT: http://otel-collector:4318
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
restart: always

To monitor GPUs on every node in a cluster, deploy the collector as a DaemonSet:

apiVersion: apps/v1
kind: DaemonSet
metadata:
name: otel-gpu-collector
namespace: monitoring
spec:
selector:
matchLabels:
app: otel-gpu-collector
template:
metadata:
labels:
app: otel-gpu-collector
spec:
hostPID: true
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
containers:
- name: collector
image: ghcr.io/thinkfleetai/otel-gpu-collector:latest
env:
- name: OTEL_SERVICE_NAME
value: gpu-collector
- name: OTEL_RESOURCE_ATTRIBUTES
value: deployment.environment=production
- name: OTEL_EXPORTER_OTLP_ENDPOINT
value: http://otel-collector.monitoring.svc.cluster.local:4318
resources:
limits:
nvidia.com/gpu: 1
securityContext:
privileged: false