Last week: client → server → database. Now let's fix the bottleneck.
Say your database gets hit with the same query a thousand times a second - like fetching a popular product page. Hitting disk every time is wasteful.
Enter the cache: a fast, in-memory layer that sits between your server and database.
[Client] → [Server] → [Cache] → [Database]
↑
(checked first)
Flow:
1️⃣ Server checks cache first
2️⃣ Cache hit → return immediately (fast!)
3️⃣ Cache miss → query database, store result in cache, return it
Popular tools: Redis, Memcached.
Two concepts you MUST be able to explain in an interview:
🔹 Cache eviction (LRU) - cache has limited memory, so when it's full, Least Recently Used items get kicked out to make room for new ones.
🔹 Cache invalidation - the hardest part. If the underlying data changes, how does the cache know to update? (Famous quote: "There are only two hard things in computer science: cache invalidation and naming things.")
Common strategies: TTL (time-to-live expiry), write-through (update cache and DB together), or explicit invalidation on writes.
Next System Design Monday: what happens when ONE server can't handle the traffic anymore - load balancing.
What would you cache first in a system like Instagram? 👇