📍Round Robin. Simple, stable, well-suited for servers with identical replicas.
📍Weighted Round Robin. Adds a `weight
for each replica, weight` usually correlates with the server capacity. Improves resource utilization in heterogeneous infrastructure.📍Least connections\Load. Directs network traffic to the server with the fewest active connections\load. It can be really effective with long-lived sessions or tasks.
📍Weighted Least Connection. It’s the previous one enriched with capacity weights.
📍Hash Ring. Each host is mapped into a circle using its hashed address, each request is routed to a host by hashing some property of the request. So the balancer finds the nearest host clockwise to match requests with the server. It can work well if there is a good request attribute to hash.
📍Random Subsetting. Each client randomly shuffles the list of hosts and fills its subset by selecting available backends from the list. In real cases the load may be distributed unevenly.
📍Deterministic Subsetting. Google improvement for Random Subsetting: adds client assignment rounds, server random shuffles and allows servers to adjust weights back to clients. It has greater stability and evenness in distribution.
📍Random Choice of 2. The algorithms suggest picking up 2 servers randomly and selecting one with the least load. Simple, quite effective in many cases.
Let’s go back to that last post and figure out why Netflix wasn't satisfied and had to cook up another load balancing algorithm.
The main issue is that stateful services are asymmetric (postgres leader and followers, zookeeper leader and followers, cassandra topology with quorums, etc.), and it does matter which host to connect to. Plus, the traffic between availability zones was substantial, leading to additional latency.
So what actually was done (on Cassandra client balancing sample):
- Random Choice of 2 is extended to Random Choice of 8
- Weight for nodes is added and based on rack topology, replicas info and health state
- Selected nodes are sorted according to the weight
The improvement reduces latency by up to 40% for Netflix scenarios, which I believe is a significant achievement. Unfortunately I couldn't find the original video detailing the implementation, but there's a presentation with measurements available.
#architecture #reliability #network