Time for our first full end-to-end design. Requirements:
โ Given a long URL, generate a short one
โ Given a short URL, redirect to the original
โ High read traffic (redirects happen constantly), lower write traffic (new URLs created less often)
Step 1 - the core challenge: how do we generate a short, unique code for each URL?
Option A: Hash the URL (like MD5) and take the first 7 characters. Risk: collisions, requires checking uniqueness.
Option B (better): use an auto-incrementing ID from the database, then convert it to base62 (a-z, A-Z, 0-9). ID
125 becomes something like "cb" - dramatically shorter, always unique by construction.Step 2 - the architecture:
[Client] โ [Load Balancer] โ [App Servers] โ [Cache] โ [Database]
(hot URLs cached
for fast redirects)
Step 3 - the schema:
urls
+--------+--------------------------+---------------------+
| id | short_code | long_url | created_at |
+--------+------------+-------------+---------------------+
| 125 | cb | example.com | 2024-01-01 10:00:00 |
Step 4 - the follow-ups interviewers actually ask:
๐น "What if two servers generate the same ID at the same time?" โ Use a centralized ID generator service, or database auto-increment, or a distributed ID scheme like Twitter's Snowflake.
๐น "How do you handle a URL that goes viral overnight?" โ This is exactly why we cache hot URLs - the cache absorbs the read spike so the database doesn't get hammered.
๐น "Should short codes expire?" โ Product decision, not purely technical - worth explicitly asking your interviewer this instead of assuming.
The key skill being tested isn't "do you know TinyURL" - it's whether you can reason from requirements to architecture out loud, and handle the follow-up curveballs.
If you were designing this, would you go with hash-based or counter-based short codes? ๐