Not everything needs an instant response. Sending a confirmation email, processing an uploaded video, generating a report - these can happen "in the background" instead of forcing the user to wait.
[Client] → [App Server] → [Message Queue] → [Worker Service] → [Database/Storage]
(responds "got it!" (processes asynchronously,
immediately) at its own pace)
Flow: instead of doing slow work directly in the request, the app server drops a "job" onto a queue (like RabbitMQ, Kafka, or SQS) and immediately responds to the user. Separate worker processes pull jobs off the queue and do the actual heavy lifting, independent of the original request's timeline.
Why this matters at scale:
✅ The user gets a fast response instead of waiting on slow work
✅ If a worker crashes mid-job, the message can be safely retried instead of being lost
✅ Workers can be scaled independently from your main app servers - heavy video processing doesn't need to compete for resources with fast web requests
⚠️ The follow-up interviewers ask: "What if the same job gets processed twice?" This is the concept of idempotency - designing your job handlers so processing the same message multiple times produces the same end result as processing it once (e.g., "set status to complete" rather than "increment counter by one," which would double-count on a retry).
Real example: when you upload a video to a platform like YouTube, encoding it into multiple resolutions happens exactly this way - asynchronously, off a queue, while you're immediately told "upload successful, processing now."
Where in a system you've worked on could background processing via a queue have improved things? 👇