The Evolution of SRE at Google
Google not only pioneered SRE practices, they constantly improve SRE approach to keep Google's large-scale production up and running. SLOs\SLIs, error budgets, isolation strategies, postmortems are well-known tools for reliability.
But Google's team went beyond that and adopted systems theory and control theory. The main idea is to focus on the complex system as a whole rather than on individual elements and their failures.
A new approach is based on System-Theoretic Accident Model and Processes (STAMP) framework. In complex systems, most accidents are the result of interactions between components that are all functioning well, but collectively produce an unsafe state. STAMP shifts analysis from a traditional linear chain of failures (root cause) to a "control problem" (I think STAMP itself deserves a separate overview).
From very basic perspective it analyzes the system in terms of control-feedback loops to find issues both in the control path and the feedback path. The feedback path is usually less well understood, but just as important as control path from a system safety perspective.
For example, imagine the service "resizer" that changes resource quotas according to the resource usage. The logic responsible for calculating resource usage and delivering these values to the resizer is probably even more critical than the logic that changes the quota itself.
Google SRE team performs such analysis on a regular basis for their global services. Even the first results showed a set of scenarios that can potentially produce hazards in the future. But as they are "predicted" before the real incidents, improvements can be planned without rush as any other feature.
Overall, I think the new approach is a really great example of system theory in practice. Moreover, I’ve discovered there’s a whole area of reliability theory behind this that I plan to explore in more detail in the future.
#engineering #reliability
Post #235
242