Post #2686
2.91K
I found a good example of why autoscaling based only on CPU utilization can cause an outage.
About a week ago, Twingate had an incident that affected us as a client. They've published a postmortem, and it's a good example of why CPU isn't a good metric to rely on when autoscaling your services.
So, from the CPU utilization perspective, everything was OK, but the number of processed requests decreased.
https://status.twingate.com/incidents/49qvqk7swjpq
Twingate Twingate Service Incident Twingate's Status Page - Twingate Service Incident. About a week ago, Twingate had an incident that affected us as a client. They've published a postmortem, and it's a good example of why CPU isn't a good metric to rely on when autoscaling your services.
The incident was triggered by elevated network latency affecting communication paths used by the Authorization service. As requests took longer to complete, individual service instances were able to process fewer requests than normal.
This reduction in throughput exposed a limitation in our auto-scaling configuration, which primarily relied on CPU utilization to determine service capacity requirements.
So, from the CPU utilization perspective, everything was OK, but the number of processed requests decreased.
https://status.twingate.com/incidents/49qvqk7swjpq
- 👍 6
- 🔥 2