TGViewer
TechLead Bits TechLead Bits @techleadbits · 517 subscribers
Post #63 312
Addressing Cascading Failures

A cascading failure is a failure that grows over time: failure of one or a few parts of a system triggers a domino effect, leading to the progressive failure of other parts. Cascading failures are one of the biggest challenges to the reliability of distributed systems.

Google SRE Book has a separate chapter about potential root causes of this type of failure, design recommendations and immediate steps to fix.

Let's start with typical causes of the problem:
- Server Overload: there are more requests then server can handle
- Resource Exhaustion: running out of CPU, RAM, threads, file descriptors
- Service Unavailability: container crash, failed readiness probes, errors
- Slow startup and cold caching

Common triggers of cascading failures:
- Maintenance procedures: updates, new rollouts, planned changes in infrastructure
- Organic growth of the load
- System usage changes: more users, increased non-typical usage scenarios
- Resource limits: clusters are usually overprovisioned, so some heavy operation can occupy resources and impact other services

Design recommendations that can help to avoid cascading failures:
✏️ In case of overload, perform load shedding and graceful degradation: reject requests that cannot be served, serve degraded results (less data, data only from cache, etc.)
✏️ Instrument higher-level systems to reject requests, rather than overload servers
✏️ Perform accurate capacity planning. Capacity planning reduces the probability of triggering a cascading failure, but it is not sufficient for protection
✏️ Smart retries implementation:
- Use randomized exponential backoff when scheduling retries
- Limit retries per request. Don’t retry a given request indefinitely.
- Consider having a server-wide retry budget. For example, only allow 60 retries per minute in a process, and if the retry budget is exceeded, don’t retry; just fail the request.
- Use clear response codes and consider how different failure modes should be handled. Don’t retry permanent errors or malformed requests in a client, because neither will ever succeed.
✏️ Set requests timeouts (they called them deadlines). Implement deadline propagation, where each service in the call chain checks whether the deadline has already been exceeded. Based on this, the service can decide whether to proceed or terminate the process.

Immediate steps to fix cascading failures on problem environment:
- Increase resources
- Restart servers
- Limit or drop traffic
- Enter degraded mode (should be supported on service level)
- Decrease batch load: Some services have load that is important, but not critical. Consider turning off those sources of load.

When systems are overloaded, something has to give to remedy the situation. If a service reaches its limit, it's better to let some errors or lower-quality results than try to serve all requests. Understanding system limits and how the system behaves under the load is critical to implement protection from cascading failures.

#architecture #systemdesign #reliability
  • ❤‍🔥 2
  • 👍 1
More from @techleadbits
  1. Oct 7, 2026AI & Repository Strategy For many years, there has been an ongoing debate between monorepo…
  2. Oct 1, 2026Tracer Bullets Continuing the topic from the previous post, let's talk in more detail abou…
  3. Sep 28, 2026Why Software Factories Fail "Read the Code!" is one of the key ideas from Dex Horthy's tal…
  4. Sep 21, 2026Illustrations from The Culture Map showing how different cultures compare on the scales. #…
  5. Sep 21, 2026The Culture Map Have you ever worked in international distributed teams? Or collaborated w…
  6. Sep 10, 2026Loop Engineering from First Principles Continuing the topic of Loop Engineering, I'd like…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →