Is Your Cluster Underutilized?
"In clusters with 50 CPUs or more, only 13% of the CPUs that were provisioned were utilized, on average. Memory utilization was slightly higher at 20%, on average," - 2024 Kubernetes Cost Benchmark Report.
This aligns with what I've seen: often there are no enough resources to deploy an app, even the cluster resource usage is less than 50%.
The most common reasons for that:
✏️ Incorrect requests and limits. Kubernetes reserves cluster resources based on requests parameters, while limits prevent a service from consuming more than expected. Typical issues are:
- Requests = Limits. Resources are exclusively reserved for peak loads and cannot be shared with other services. This often happens with memory configuration (standard explanation is to prevent OOM), but mostly all modern languages can manage memory dynamically (e.g., Java since v17 with G1 and Shenandoah GC, Go, and C++ handle this natively).
- Requests > Average Usage. It may seem counterintuitive, but requests should be set below average usage. By the law of large numbers, peak loads is balanced across multiple deployments. Probability that all services hit peak usage at the same time is relatively small.
✏️ Mismatched between resource settings and runtime parameters. Requests and limits should align with language-specific configurations (e.g., GOMEMLIMIT for Golang, Xms and Xmx for java).
✏️ High startup resource requirements. Some services need a lot of CPU and memory just to start, even though they consume far less after that. This requires high resource requests just to make deployment possible. A common example is Spring applications, which consume significant resources to load all beans at startup. Using native compilation or more efficient technologies can help.
As you can see the real problem is somewhere between developers and deployment engineers. Technical decisions and implementation details directly impact resource utilization and infrastructure costs.
#engineering
Post #140
322