Spotify has released a postmortem for their outage that happened on 16th of April, and was almost global.
In nutshell, it was a combination of a bug, and a cascading issue caused by user retries. Here's an interesting bit:
> This change was deemed low risk and as such we applied it to all regions at the same time.
This is something what burned a lot of engineers. So, the take-away is probably never consider any change low-risk, especially if you already have the architecture for gradual rollouts. However, it's much easier to be said than done.
#postmortem #sre
Post #2685
2.02K