One server is now overwhelmed. Two options: make it bigger (vertical scaling) or add more of them (horizontal scaling). Horizontal scaling is what most large systems actually do - there's a ceiling on how big one machine can get, but no real ceiling on how many machines you can add.
βββββββββββ
βββββββΆβ Server 1β
β βββββββββββ
[Client] β [Load Balancer]
β βββββββββββ
βββββββΆβ Server 2β
β βββββββββββ
β βββββββββββ
βββββββΆβ Server 3β
βββββββββββ
The load balancer sits in front of your servers and distributes incoming requests, usually via:
πΉ Round robin - requests go to servers in rotating order
πΉ Least connections - send to whichever server currently has the fewest active requests
πΉ IP hash - same client always routed to the same server (useful for session data)
Here's the important follow-up question interviewers ask: "if a user's session data is stored in Server 2's memory, what happens if the load balancer routes their next request to Server 3?"
That's the concept of statelessness - well-designed servers shouldn't hold session data locally at all. Instead, store session state in a shared cache (like Redis) or a database, so ANY server can handle ANY request. This is a foundational principle behind horizontally scalable systems.
Why do you think "statelessness" matters so much in distributed systems? π