Currently, there's considerable interest in automatic autoscaling, largely driven by the cost of resource usage. Let's verify how it works with dynamic load inside Kubernetes. There exists a naive belief that the HPA will deliver equivalent performance to pre-configured replicas when demand arises.
Consider a scenario where we set CPU utilization target at 80%. During peak events such as Black Friday or intense invoicing periods, service demands surge, causing CPU utilization growth to 100%.
The HPA operates on a logic where desired replicas are calculated using the formula:
desiredReplicas = ceil[currentReplicas * (currentMetricValue / desiredMetricValue)]
For instance, if we have 5 replicas and the CPU utilization grows to 100% out of the desired 80%, the calculation will result in 6 replicas as a target. Consequently, the deployment scales up to 6 replicas. However, this scaling process involves a delay as the HPA controller waits for all pods to be ready, which can be time-consuming depending on the size of the service. Additionally, there's a stabilization window with a default of 5 minutes.
After the initial scaling, it's highly likely that the adjustment won't be sufficient (let’s say you need around 50 replicas to handle the load). Therefore, subsequent iterations of scaling will occur at intervals of at least 5 minutes until the desired 80% utilization is achieved or until the maximum allowed number of replicas is reached. This delay might not meet the business requirements.
In summary, autoscaling requires time to adapt to the load. While it functions effectively in scenarios where the load changes gradually, it's less suitable for handling scheduled, heavy loads. In such cases, it's advisable to explore alternative options and proactively prepare the environment for the anticipated load.
#architecture #performance #scalability