More follow-ups for the AWS outage (Azure outage didn't generate that much press).
Lorin Hochstein analyzes the postmortem from the complexity point of view and comes to quite interesting conclusions that you can absolutely apply to your incidents and postmortems as well.
tl;dr is that incidents (especially bigger ones) are often unique. So, when reasoning about the preventive measures, you need not only to prevent similar incidents, but also get prepared to handle incidents in general, because the next incident may be not the same as the present one.
#reliability #sre #aws
Post #2791
1.69K