Reliability is the top priority feature for modern systems. Your customers always expect service reliability, even if they don't realize it. Nobody interests in the super cool feature that cannot be used (because of service unavailability or bad performance for example).
Reliability can be defined as the ability of a system to carry out its intended function without interruption. Good definition, but not really actionable. I prefer Google definition
Your service is reliable when your customers are happy.
The logic is simple: if system is not reliable, users will not use it. If users don't use it, it's worth nothing. So reliability matters.
Let's check what reliable architecture usually includes:
📍Measurable Reliability Targets: SLO, SLI and error budget
📍High-Availability:
- Redundancy: multiple replicas for the same service
- Self-Healing: the ability to remediate issues without manual interventions
- Graceful Degradation: degrade service levels gracefully when overloaded
- Fail Safe: be ready for unexpected failure, no data or system corruption
- Retriable APIs: make your operations idempotent, allow retries
- Critical Dependencies Minimization: the reliability level of a service is defined by the reliability of its least reliable component or dependency.
- Multiple Availability Zones (AZ): spread instances across multiple AZ, ability to survive in case of AZ outage
📍Disaster Recovery:
- Multiple Regions: spread instances across multiple regions (each region has multiple AZ), ability to survive in case of region failure
- Data Replication Across Regions
📍Scalability: ability to scale for increased workload
📍Observability: code instrumentation, tools for data collection and analysis, fast failure detection
📍Recovery Procedures: rollback strategies, recovery from outages
📍Chaos Engineering: practices to test failures internally
📍Operational Excellence: a fully automated operational experience with minimal manual steps and low cognitive complexity
References:
- Google Cloud Architecture Framework: Reliability
- AWS Well-Architected Framework
- AWS Well-Architected Framework: Reliability Pillar
- Azure Well-Architected Framework: Reliability
#architecture #systemdesign #reliability