tl;dr: config changes. Config changes can be dangerous too. Despite there were successful deploys between the update of CrowdStrike Scanner and the outage, it seems like a new type of config was deployed which caused the entire clusterfuck.
This line is also interesting:
June 4th, Red Hat released a KB relating to kernel panics that were caused by the Crowdstrike sensor
process. This was a bug in the Linux kernel itself, that the sensor was
triggering and wasn’t Crowdstrike’s fault. However it does prove that config that has passed the Content Validator can cause kernel panics.
UPD: I think the most important take-away here is not what caused the outage or how the deployment process at CrowdStrike looks like. It's the fact that problems can be obscure enough. When something goes wrong big times, it's easy to "blame" a "big thing": the whole deployment process, or code quality, or people behind the software. This is much more comforting than the idea that any small change can cause a butterfly-effect and take your whole system down. This was true for CrowdStrike and this is true for you as well.
#postmortem #crowdstrike #windows