TGViewer
TechLead Bits TechLead Bits @techleadbits · 517 subscribers
Post #8 261
Reliable Stateful Systems at Netflix

There is quite an interesting long-read article explaining how Netflix implements the reliability of its stateful services.

To understand what reliability actually means the author proposes answering three fundamental questions:
- How often does the system fail?
- When it fails, how large is the blast radius?
- How long does it take to recover from an outage?
Ideally, systems don’t fail, have minimal impact on failure and recover very quickly😀. That is the simple part.

So how is that achieved? That is actually the hard part:
📍Single tenancy. There are no multi-tenant data stores. It minimizes blast impact.
📍Capacity management. Netflix programs special workload capacity models that generate specifications for a cluster according to the system requirements and SLOs.
📍Data replication. Data is replicated in 12 availability zones across 4 regions.
📍Overprovisioning. In case of one region degradation, the traffic is spread across other regions so each region must keep an extra 33% capacity reserved for failover.
📍Snapshot restoration. To replace one instance with another - load snapshot from s3 and then apply delta.
📍Performance monitoring. Continuous monitoring is essential to detect failures quickly, remediate them and recover.
📍Cache in front of services. The idea is to use caching for complex business logic rather than underlying data. Service with business logic is quite expensive to operate, so approach shifts load to the cache.
📍Reliable clients. Quite complex approach to inform clients what timeouts and level of service that they can rely on. Better to read the original.
📍Load Balancing. Netflix uses improved choice-of-2 algorithms with weight requests. It takes into account availability zones, replicas, health state.
📍Stateful APIs. All requests are idempotent with built-in work pagination. That approach requires additional idempotent tokens or global transaction ids.

#architecture #reliability #usecase
InfoQ How Netflix Ensures Highly-Reliable Online Stateful Systems Building reliable stateful services at scale isn’t a matter of building reliability into the servers, the clients, or the APIs in isolation. By combining smart and meaningful choices for each of these three components, we can build massively scalable, SLO…
  • ❤ 2
More from @techleadbits
  1. Oct 7, 2026AI & Repository Strategy For many years, there has been an ongoing debate between monorepo…
  2. Oct 1, 2026Tracer Bullets Continuing the topic from the previous post, let's talk in more detail abou…
  3. Sep 28, 2026Why Software Factories Fail "Read the Code!" is one of the key ideas from Dex Horthy's tal…
  4. Sep 21, 2026Illustrations from The Culture Map showing how different cultures compare on the scales. #…
  5. Sep 21, 2026The Culture Map Have you ever worked in international distributed teams? Or collaborated w…
  6. Sep 10, 2026Loop Engineering from First Principles Continuing the topic of Loop Engineering, I'd like…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →