AMD GPUs
The collector monitors AMD GPUs directly from the Linux kernel’s sysfs and hwmon interfaces. No ROCm, no user-space libraries, and no additional drivers are needed beyond the standard AMDGPU kernel module.
Requirements
Section titled “Requirements”- Linux with the
amdgpukernel driver - Kernel 5.x+ (sysfs/hwmon paths are stable from 5.x onwards)
Collected metrics
Section titled “Collected metrics”| Metric | Description |
|---|---|
hw.gpu.utilization | Compute utilization (0.0–1.0) |
hw.gpu.memory.utilization | Fraction of GPU memory used (usage / limit) |
hw.gpu.memory.controller.utilization | Memory controller busy fraction (extension) |
hw.gpu.memory.limit | Total VRAM (bytes) |
hw.gpu.memory.usage | Used VRAM (bytes) |
hw.gpu.memory.free | Free VRAM (bytes) |
hw.temperature | Die temperature (°C) |
hw.fan.speed | Fan speed (RPM) |
hw.power | Current power draw (W) |
hw.power.limit | Power cap (W) |
hw.energy | Cumulative energy (J) |
Docker
Section titled “Docker”docker run -d \ --name otel-gpu-collector \ --device /dev/kfd:/dev/kfd \ --device /dev/dri:/dev/dri \ --pid=host \ -e OTEL_SERVICE_NAME=my-app \ -e OTEL_RESOURCE_ATTRIBUTES='deployment.environment=production' \ -e OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318 \ ghcr.io/thinkfleetai/otel-gpu-collector:latest--pid=host is required for per-process GPU attribution (cmdline, PID, zombie state).
Docker Compose
Section titled “Docker Compose”services: otel-gpu-collector: image: ghcr.io/thinkfleetai/otel-gpu-collector:latest pid: host environment: OTEL_SERVICE_NAME: my-app OTEL_RESOURCE_ATTRIBUTES: deployment.environment=production OTEL_EXPORTER_OTLP_ENDPOINT: http://otel-collector:4318 devices: - /dev/kfd:/dev/kfd - /dev/dri:/dev/dri restart: alwaysKubernetes (DaemonSet)
Section titled “Kubernetes (DaemonSet)”apiVersion: apps/v1kind: DaemonSetmetadata: name: otel-gpu-collector namespace: monitoringspec: selector: matchLabels: app: otel-gpu-collector template: metadata: labels: app: otel-gpu-collector spec: hostPID: true containers: - name: collector image: ghcr.io/thinkfleetai/otel-gpu-collector:latest env: - name: OTEL_SERVICE_NAME value: gpu-collector - name: OTEL_RESOURCE_ATTRIBUTES value: deployment.environment=production - name: OTEL_EXPORTER_OTLP_ENDPOINT value: http://otel-collector.monitoring.svc.cluster.local:4318 securityContext: privileged: false volumeMounts: - name: sys mountPath: /sys readOnly: true - name: dri mountPath: /dev/dri volumes: - name: sys hostPath: path: /sys - name: dri hostPath: path: /dev/dri