TGViewer
DevOps&SRE Library DevOps&SRE Library @devopslibrary · 19.9K subscribers
Post #7592 2.69K
Client’s GKE Cluster Ate Their Entire VPC

GKE pod IP exhaustion is one of the few failure modes that gives you no warning before it goes terminal. I recently stepped into a war room where a client’s primary scaling group had flatlined — workloads cordoned, deployments stuck in Pending, and the estimated cost of the stall nearing $15k per hour in lost transaction volume. The culprit wasn’t traffic. It was a /20 subnet that had quietly run out of address space, and a set of GKE allocation defaults nobody had questioned at design time.


The IP Math I Uncovered During Triage: https://www.rack2cloud.com/gke-pod-ip-exhaustion-triage-part-1

The Class E Rescue: https://www.rack2cloud.com/gke-ip-exhaustion-fix-part-2
More from @devopslibrary
  1. Sep 23, 2026From Ingress to Gateway API: How We Modernized Networking on Our GKE Cluster We recently m…
  2. Sep 23, 2026Building a Real k6 Test Suite Against a Live Kubernetes App In part 1 I covered k6's philo…
  3. Sep 22, 2026My Experiments with MCP: Moving Beyond the "Agent Wrapper" I'm currently working with a cl…
  4. Sep 22, 2026Your AI just deleted the wrong deployment. Now what? Picture this. A developer asks an AI…
  5. Sep 21, 2026Kafka on Kubernetes: Performance Lessons for Any Disk-Heavy Data Service We recently start…
  6. Sep 21, 2026What the Popularity of Emerging Tools Tells Us About Kubernetes' Future Kubernetes has mat…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →