Skip to content

GPU Performance Monitoring

Get started

Using the SDK
Deploy Shield360
Git clone Shield360 repository
Terminal window
git clone git@github.com:ThinkfleetAI/Shield360.git
Start Docker Compose

From the root directory of the Shield360 distribution, Run the below command:

Terminal window
docker compose up -d
Install Shield360

Open your command line or terminal and run:

Terminal window
pip install shield360
Initialize Shield360 in your Application
Terminal window
# Start GPU monitoring instantly
shield360-instrument --collect-system-metrics python your_app.py
# With custom settings
shield360-instrument \
--otlp-endpoint http://127.0.0.1:4318 \
--service-name my-gpu-app \
--environment production \
--collect-system-metrics \
python your_app.py

To send metrics to other Observability tools, refer to the supported destinations.

For more advanced configurations and application use cases, visit the SDK configuration reference.

Using the Collector
Deploy Shield360
Git clone Shield360 repository
Terminal window
git clone git@github.com:ThinkfleetAI/Shield360.git
Start Docker Compose

From the root directory of the Shield360 distribution, Run the below command:

Terminal window
docker compose up -d
Pull `otel-gpu-collector` Docker Image

You can quickly start using the OTel GPU Collector by pulling the Docker image:

Terminal window
docker pull ghcr.io/thinkfleetai/otel-gpu-collector:latest
Run `otel-gpu-collector` Docker container

You can quickly start using the OTel GPU Collector by pulling the Docker image: Here’s a quick example showing how to run the container with the required environment variables:

Terminal window
docker run --gpus all --pid=host \
-e OTEL_SERVICE_NAME='chatbot' \
-e OTEL_RESOURCE_ATTRIBUTES='deployment.environment=staging' \
-e OTEL_EXPORTER_OTLP_ENDPOINT="http://127.0.0.1:4318" \
ghcr.io/thinkfleetai/otel-gpu-collector:latest

--pid=host is required for per-process GPU attribution (cmdline, PID, zombie state).

For more advanced configurations of the collector, visit the GPU Collector documentation.

Note: If you’ve deployed Shield360 using Docker Compose (docker-compose.yml in the Shield360 distribution), make sure to use the host’s IP address or add OTel GPU Collector to the Docker Compose (docker-compose.yml in the Shield360 distribution):

Docker Compose: Add the following config under `services`
otel-gpu-collector:
image: ghcr.io/thinkfleetai/otel-gpu-collector:latest
pid: host
environment:
OTEL_SERVICE_NAME: 'chatbot'
OTEL_RESOURCE_ATTRIBUTES: 'deployment.environment=staging'
OTEL_EXPORTER_OTLP_ENDPOINT: "http://otel-collector:4318"
device_requests:
- driver: nvidia
count: all
capabilities: [gpu]
depends_on:
- otel-collector
restart: always
Host IP: Use the Host IP to connect to OTel Collector
Terminal window
OTEL_EXPORTER_OTLP_ENDPOINT="http://192.168.10.15:4318"