TGViewer
DevOps&SRE Library DevOps&SRE Library @devopslibrary · 19.9K subscribers
Post #7590 2.66K
What Does 4.4% GPU Utilization Actually Mean?

A few weeks after publishing the 1M token/s post, I spent a weekend helping my good friend Milko Ilari set up vLLM on his shiny new DGX Spark with Gemma 4. My first in-person reaction was “It’s Champagne” (from the old days of PC Perspective) The Spark is a wild little machine — 128 GB of unified memory in a box you can hold with one hand, running the same Blackwell architecture as the datacenter B200s. But its memory bandwidth is 273 GB/s. The B200s in our cluster do 8,000 GB/s. Almost 30x less.

Watching the numbers on that tiny machine got me thinking. The benchmark I ran on GKE Autopilot with 96 B200 GPUs had reported 4.4% FLOPS utilization. 10.9% memory bandwidth. Tensor cores active 1.5% of the time. The GPUs looked almost idle while pushing a million tokens per second. Was something wrong?

No. And honestly, figuring out why turned out to be more interesting than the benchmark itself.

That first post covers the journey — every optimization, and many failure 🫠. This one covers the physics.


https://medium.com/google-cloud/what-does-4-4-gpu-utilization-actually-mean-ee61fabebbf0
More from @devopslibrary
  1. Sep 23, 2026From Ingress to Gateway API: How We Modernized Networking on Our GKE Cluster We recently m…
  2. Sep 23, 2026Building a Real k6 Test Suite Against a Live Kubernetes App In part 1 I covered k6's philo…
  3. Sep 22, 2026My Experiments with MCP: Moving Beyond the "Agent Wrapper" I'm currently working with a cl…
  4. Sep 22, 2026Your AI just deleted the wrong deployment. Now what? Picture this. A developer asks an AI…
  5. Sep 21, 2026Kafka on Kubernetes: Performance Lessons for Any Disk-Heavy Data Service We recently start…
  6. Sep 21, 2026What the Popularity of Emerging Tools Tells Us About Kubernetes' Future Kubernetes has mat…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →