DCGM exporter tells you a GPU is hot. It won't tell you whose job is frying it. l9gpu closes the loop. One agent per node emits vendor-neutral OTLP with workload attribution baked in — Kubernetes pod, namespace, deployment; Slurm job, user, partition.
https://github.com/last9/gpu-telemetry