Installation
Docker (recommended)
Section titled “Docker (recommended)”The easiest way to run the collector. The image is published to GitHub Container Registry and supports linux/amd64 and linux/arm64.
docker pull ghcr.io/thinkfleetai/otel-gpu-collector:latestFor per-process GPU attribution (cmdline, PID, zombie/process.state, owner), run with host PID namespace access: Docker --pid=host, Compose pid: host, or Kubernetes hostPID: true. Device-level hw.gpu.* metrics work without it.
| Tag | Description |
|---|---|
latest | Most recent release |
1.2.3 | Specific version |
1.2 | Latest patch of a minor version |
Pre-built binaries
Section titled “Pre-built binaries”Download a binary for your platform from the GitHub Releases page. Binaries are available for:
| Platform | Architecture |
|---|---|
| Linux | amd64, arm64, armv7 |
| macOS | amd64 (Intel), arm64 (Apple Silicon) |
| Windows | amd64, arm64 |
# Example: Linux amd64curl -L https://github.com/ThinkfleetAI/Shield360/releases/latest/download/opentelemetry-gpu-collector-<version>-linux-amd64 \ -o opentelemetry-gpu-collectorchmod +x opentelemetry-gpu-collector
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317 \./opentelemetry-gpu-collectorVerify the SHA256 checksum from the SHA256SUMS.txt file in the release:
sha256sum -c SHA256SUMS.txt --ignore-missingBuild from source
Section titled “Build from source”Requirements: Go 1.21+, CGO enabled (required for NVML on Linux).
git clone https://github.com/ThinkfleetAI/Shield360.gitcd shield360/opentelemetry-gpu-collectormake build./opentelemetry-gpu-collectorFor eBPF CUDA tracing support, also run:
make setup-bpf # installs bpftool, generates vmlinux.hmake generate # runs bpf2go code generationmake buildKubernetes DaemonSet
Section titled “Kubernetes DaemonSet”Run one collector per GPU node. Use the OpenTelemetry-recommended pattern: downward API → K8S_NODE_NAME → OTEL_RESOURCE_ATTRIBUTES with host.name and k8s.node.name. On GKE, AKS, and EKS the collector also auto-detects k8s.cluster.name, cloud.provider, and host.type (instance type) from the Kubernetes Node object and/or cloud metadata.
apiVersion: v1kind: ServiceAccountmetadata: name: otel-gpu-collector---apiVersion: rbac.authorization.k8s.io/v1kind: ClusterRolemetadata: name: otel-gpu-collector-node-getrules: - apiGroups: [""] resources: ["nodes"] verbs: ["get"]---apiVersion: rbac.authorization.k8s.io/v1kind: ClusterRoleBindingmetadata: name: otel-gpu-collector-node-getroleRef: apiGroup: rbac.authorization.k8s.io kind: ClusterRole name: otel-gpu-collector-node-getsubjects: - kind: ServiceAccount name: otel-gpu-collector namespace: default # change to your namespace---apiVersion: apps/v1kind: DaemonSetmetadata: name: otel-gpu-collectorspec: selector: matchLabels: app: otel-gpu-collector template: metadata: labels: app: otel-gpu-collector spec: serviceAccountName: otel-gpu-collector hostPID: true containers: - name: otel-gpu-collector image: ghcr.io/thinkfleetai/otel-gpu-collector:latest env: - name: K8S_NODE_NAME valueFrom: fieldRef: fieldPath: spec.nodeName - name: OTEL_SERVICE_NAME value: otel-gpu-collector - name: OTEL_RESOURCE_ATTRIBUTES value: "host.name=$(K8S_NODE_NAME),k8s.node.name=$(K8S_NODE_NAME)" - name: OTEL_EXPORTER_OTLP_ENDPOINT value: "http://otel-collector:4317" # Optional for on-prem / self-managed clusters: # - name: K8S_CLUSTER_NAME # value: my-cluster # eBPF CUDA tracing is on by default on Linux; set false to disable: # - name: OTEL_GPU_EBPF_ENABLED # value: "false" volumeMounts: - name: pod-resources mountPath: /var/lib/kubelet/pod-resources readOnly: true # AMD/Intel DRM: # - name: dri # mountPath: /dev/dri securityContext: capabilities: add: ["SYS_ADMIN"] # or privileged / CAP_BPF+CAP_PERFMON for eBPF volumes: - name: pod-resources hostPath: path: /var/lib/kubelet/pod-resources # - name: dri # hostPath: # path: /dev/driSee Configuration for identity env vars and detection order.
Upgrade
Section titled “Upgrade”Docker
Section titled “Docker”docker pull ghcr.io/thinkfleetai/otel-gpu-collector:latestdocker stop otel-gpu-collectordocker rm otel-gpu-collector# re-run with same flagsBinary
Section titled “Binary”Download the new binary from the Releases page, replace the existing file, and restart the process.
Uninstall
Section titled “Uninstall”Docker
Section titled “Docker”docker stop otel-gpu-collectordocker rm otel-gpu-collectordocker rmi ghcr.io/thinkfleetai/otel-gpu-collector:latestBinary
Section titled “Binary”rm /usr/local/bin/opentelemetry-gpu-collectorTroubleshooting
Section titled “Troubleshooting”No GPU metrics - collector starts but reports no hw.gpu.* metrics
- Confirm the host has a supported GPU:
lspci | grep -E 'VGA|3D|Display' - For NVIDIA: verify
libnvidia-ml.sois present:ldconfig -p | grep nvidia-ml - For Docker: ensure
--gpus all(NVIDIA) or--device /dev/dri(AMD/Intel) is passed - Check logs:
docker logs otel-gpu-collectorfor"discovered GPU"entries
eBPF tracing not working
- On Linux it is enabled by default; confirm it is not disabled via
OTEL_GPU_EBPF_ENABLED=false - Check kernel version:
uname -r(requires 5.8+) - The process needs
CAP_BPFandCAP_PERFMON, or run as root - For Docker:
--cap-add CAP_BPF --cap-add CAP_PERFMON,--pid=host, and--ulimit memlock=-1:-1(BPF maps need locked memory; no CUDA mount needed) - For Kubernetes:
hostPID: trueplus BPF capabilities (orSYS_ADMIN); raise memlock if map create returns EPERM - If no CUDA process is running yet, the collector rescans
/procevery 30s - Verify a workload has loaded CUDA:
grep -E 'libcudart|libcuda.so' /proc/*/maps 2>/dev/null | head
Connection refused / no data reaching backend
- Verify
OTEL_EXPORTER_OTLP_ENDPOINTis reachable from the container:curl http://<endpoint>/health - For Docker networking: use the host IP or service name, not
localhost - Check if gRPC vs HTTP/protobuf matches the backend: set
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuffor HTTP backends (port 4318)
Intel GPU not detected
- Verify the i915 or Xe driver is loaded:
lsmod | grep -E 'i915|xe' - Check DRM entries exist:
ls /sys/class/drm/ - Requires Linux kernel 5.10+ for sysfs metric exposure
- Fan speed requires kernel 6.16+