What We Can Learn from UK E-Gate Failure
In the beginning of May, the UK E-Gate system experienced a 4-hour outage. This system handles automatic passport control, with over 270 automated gates spread across 15 airports and rail ports in the UK. These gates use biometric data and advanced facial recognition technology to allow entry into the country. The failure caused major airports to cease operations, leading to delays for passengers attempting to cross the UK border.
The outage was not the result of a cyber attack but rather a network issue. The Home Office was conducting a software update that exceeded the data limits outlined in the contract, prompting the network provider to shut down the service. This network was also used for a connection between the gates and the database used for passport verifications.
Dave Farley addressed potential architectural flaws in the E-Gate system in a recent video and offered recommendations for designing distributed systems with failure in mind.
So how E-Gate architecture can be improved to prevent such failures:
- Add redundancy (reserved network channel)
- Cache copies of the data
- Limit the ways things can fail (e.g., use local network instead)
If something can go wrong, it will go wrong. And we as engineers must be prepared for potential failures, designing systems that can handle problems instead of trying to make them perfect.
Here are some general recommendations to consider when building distributed systems:
✔️Assume things can go wrong
✔️ Limit the blast radius of failure
✔️ Use Chaos Engineering techniques
Just a few days ago on May 30th, the E-Gate system encountered another failure🤦♂️, resulting in disruptions at railway stations and leaving passengers queued for approximately 4 hours. It appears that the fundamental issues with the system remain unresolved.
#architecture #reliability #resilience #usecase
Post #17
210