Skip to content

Quickstart

In this guide you’ll pull the collector Docker image, point it at your OTel backend, and start seeing GPU and host metrics within minutes.

  • Linux host with NVIDIA, AMD, or Intel GPU (for GPU metrics)
  • Docker installed
  • An OpenTelemetry-compatible backend (Shield360, Grafana, Datadog, or any OTLP endpoint)
Don't have an OTel backend? Start Shield360 locally
Terminal window
docker run -d \
--name shield360 \
-p 3000:3000 \
-p 4318:4318 \
ghcr.io/thinkfleetai/shield360:latest

Then use http://localhost:4318 as your OTEL_EXPORTER_OTLP_ENDPOINT.

Pull the collector image
Terminal window
docker pull ghcr.io/thinkfleetai/otel-gpu-collector:latest
Run the collector
Terminal window
docker run -d \
--name otel-gpu-collector \
--gpus all \
--pid=host \
-e OTEL_SERVICE_NAME=my-app \
-e OTEL_RESOURCE_ATTRIBUTES='deployment.environment=production' \
-e OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 \
ghcr.io/thinkfleetai/otel-gpu-collector:latest

Requires the NVIDIA Container Toolkit on the host. --pid=host is required for per-process GPU attribution (cmdline, PID, zombie state).

Verify it's running
Terminal window
docker logs otel-gpu-collector

You should see output like:

time=2024-01-01T00:00:00Z level=INFO msg="starting opentelemetry-gpu-collector"
time=2024-01-01T00:00:00Z level=INFO msg="discovered GPU" address=0000:01:00.0 vendor=nvidia
time=2024-01-01T00:00:00Z level=INFO msg="system metrics collector initialized"
time=2024-01-01T00:00:00Z level=INFO msg="process metrics collector initialized"
time=2024-01-01T00:00:00Z level=INFO msg="collector running"
View metrics in your backend

Open your OTel backend and look for metrics in the hw.gpu.*, system.*, and process.* namespaces.

If using Shield360, navigate to http://localhost:3000 and go to the Metrics section.

Add the collector as a service alongside your existing stack:

services:
otel-gpu-collector:
image: ghcr.io/thinkfleetai/otel-gpu-collector:latest
pid: host
environment:
OTEL_SERVICE_NAME: my-app
OTEL_RESOURCE_ATTRIBUTES: deployment.environment=production
OTEL_EXPORTER_OTLP_ENDPOINT: http://otel-collector:4318
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
depends_on:
- otel-collector
restart: always

pid: host (Docker --pid=host) is required so the collector can see host workload PIDs under /proc for per-process GPU metrics, cmdline, and zombie detection.