Overview
This is the run-of-show for the system design round: what to say, what to decide, and the numbers to have ready. Skim it the night before; use the focused sheets in this category for depth on any single topic.
The 45-minute playbook
The interview is a structured conversation with a clock. Interviewers grade how you drive it, not just what you draw. Follow this timeline and you will never run out of time with an empty whiteboard.
| Minutes | Phase | What the interviewer is grading | The mistake that sinks people |
|---|---|---|---|
| 0-5 | Clarify requirements | Do you build the right thing, or the thing in your head? | Jumping straight to boxes and arrows |
| 5-10 | Estimate scale | Can you turn vague scope into numbers that drive decisions? | Skipping math, then guessing at architecture |
| 10-20 | High-level design | Can you draw a simple end-to-end system that works? | Adding Kafka and shards before anything works |
| 20-35 | Deep dive | Real depth in 1-2 components (interviewer usually picks) | Staying shallow everywhere; touring buzzwords |
| 35-43 | Scale and failure | Do you know what breaks first and what to do about it? | "It scales because it's microservices" |
| 43-45 | Wrap up | Can you summarize trade-offs and name what you'd do next? | Trailing off instead of closing |
Two habits that raise every phase: think out loud (a silent minute feels like five), and check in ("does this match what you had in mind, or should I go deeper here?").
Clarifying questions that score points
The first five minutes are free points. Interviewers deliberately give vague prompts to see if you ask. Ask these, then write the answers in a corner of the board.
- Who uses this and what is the core action? (One sentence. "Users shorten a URL and share it.")
- How many users? Daily actives, not registered accounts.
- Read-heavy or write-heavy? Ask for the ratio. This one answer drives caching, replicas, and DB choice.
- How fast must it feel? Interactive (under ~200 ms) vs background (seconds is fine).
- How correct must it be? Is stale data acceptable for a few seconds, or never (money, inventory)?
- What scale of data? Size per item and how long we keep it.
- Any hard constraints? Regions, compliance, mobile-only, budget.
- What is explicitly out of scope? Say it out loud so you are not graded on it.
When the interviewer says "you decide": decide, state it as an assumption, and move on. "I'll assume 10M daily actives and 100:1 reads to writes - tell me if you want different numbers." Deciding quickly with stated assumptions reads as senior; fishing for hints reads as junior.
Back-of-envelope math that never fails
You need about six numbers and one method. Round aggressively - the goal is the right power of ten, not the right digit.
The shortcuts
| Fact | Use it for |
|---|---|
| 1 day ≈ 100,000 seconds (really 86,400) | Requests/day → requests/second |
| 1M requests/day ≈ 12/sec | Instant QPS conversion |
| Peak ≈ 2-5x average | Size for peak, not average |
| 1 char ≈ 1 byte; typical DB row or tweet ≈ 1 KB | Storage per item |
| Photo ≈ 500 KB, minute of video ≈ 10-50 MB | Media-heavy estimates |
| Million × KB = GB; billion × KB = TB | Skip the zeros: 10^6 × 10^3 bytes = 10^9 |
The latency ladder
The only hardware numbers worth memorizing:
| Operation | Time | Takeaway |
|---|---|---|
| Read from memory | ~100 ns | Memory is effectively free |
| Read from SSD | ~100 µs | 1,000x slower than memory |
| Round trip, same datacenter | ~0.5 ms | Chatty services add up fast |
| Disk seek (HDD) | ~10 ms | Avoid on hot paths |
| Round trip, across continents | ~150 ms | Why CDNs and regions exist |
Worked example, 60 seconds
Photo app: 10M daily actives, each views 50 photos and uploads 2/day at 500 KB.
- Write QPS: 20M uploads / 100k sec = 200/sec, peak ~1,000/sec. Modest.
- Read QPS: 500M views / 100k sec = 5,000/sec, peak ~25,000/sec. Read-heavy: cache + CDN.
- Storage: 20M × 500 KB = 10 TB/day ≈ 3.7 PB/year. Object storage, not a database.
Close with a sanity check: "so this is a read-heavy system whose hard problem is storage and delivery, not write throughput." That sentence - math driving the design - is what the estimate phase is actually for.
Building blocks at a glance
Every design is the same dozen blocks arranged differently. For each one, know its one-line job, when to reach for it, and the trade-off to say out loud when you add it.
| Block | One-line job | Reach for it when | Trade-off to say out loud |
|---|---|---|---|
| Load balancer | Spread requests across identical servers | More than one app server | Servers must be stateless (sessions move to Redis) |
| CDN | Serve static files from near the user | Global users, images/video/JS | Cache invalidation delay on updates |
| Cache (Redis) | Answer repeat reads from memory | Read-heavy, tolerates slight staleness | Invalidation strategy and cold-start misses |
| Message queue | Absorb work now, process it later | Slow work in the request path; traffic spikes | Eventual completion; consumers must be idempotent |
| Relational DB | Source of truth with transactions | Almost always, at the start | Vertical limits first; scaling reads is easy, writes hard |
| Key-value store | Tiny values by exact key, huge scale | Sessions, counters, rate limits | No queries, no joins - key lookups only |
| Wide-column store | Massive write volume, time-ordered data | Feeds, logs, messages at scale | Design tables around queries; no ad-hoc queries later |
| Object store (S3) | Cheap, durable files of any size | Anything over ~1 MB: images, video, backups | Store the URL in the DB, never the file |
| Search index | Text search and filtering | "Search by keyword" appears in requirements | A copy of the data, seconds behind the source of truth |
| WebSocket gateway | Hold open connections, push to clients | Chat, live updates, presence | Connections are state: need sticky routing or a registry |
| Scheduler / cron | Run work on a timer | Cleanup, digests, retries, reports | Needs a lock or leader so two workers don't double-run |
Rule of thumb: introduce a block only when a number or requirement demands it, and say which one. "At 25k reads/sec we need a cache" beats "let's add Redis" every time.
Choosing a database in 60 seconds
Say the senior line first: "I'd start with one Postgres instance - it handles more than people think - and here is what would force a change." Then choose by access pattern, not by fashion.
| Your data looks like | Pick | Because |
|---|---|---|
| Entities with relations, transactions, flexible queries | Postgres / MySQL | Joins, ACID, 50 years of tooling |
| Self-contained documents, shape varies per record | Document DB (MongoDB) | Schema flexibility, reads whole objects |
| Billions of tiny lookups by exact key | Key-value (Redis / DynamoDB) | Predictable ~ms reads at any scale |
| Relentless writes, time-ordered, known queries | Wide-column (Cassandra) | Write-optimized, linear scale-out |
| Keyword search, ranking, filters | Search index (Elasticsearch) | Inverted index; runs beside the real DB |
| Files and media | Object store + metadata row in SQL | DBs are terrible at blobs |
Signals that force a move off single-Postgres, in the order they usually arrive: read load (add replicas - still Postgres), dataset too big for one machine or write QPS in the tens of thousands (shard, or move that table to a wide-column store), one access pattern dominating (peel it off to a specialized store).
The trap to avoid: picking NoSQL "for scale" on day one, then discovering you need transactions and joins. Consistency needs decide too - money and inventory want ACID; likes and view counts don't.
The scaling ladder
Scale in this order. Each rung is cheaper and less complex than the next; climbing early is how systems get over-engineered. Name the rung, say what breaks next.
- One bigger machine (vertical). Boring and correct. Breaks: hardware ceiling, single point of failure.
- Stateless app servers behind a load balancer. Move sessions to Redis first. Breaks: the database is now the bottleneck.
- Cache + CDN. Kill repeat reads before they hit the DB; 80-90%+ hit rates are normal. Breaks: writes and cache misses still land on one DB.
- Read replicas. Writes go to the leader, reads to followers. Breaks: replication lag (read-your-own-writes bugs), and write volume still hits one machine.
- Queues for async work. Take everything slow out of the request path; absorb spikes. Breaks: eventually nothing - until raw write volume does.
- Shard the database. Split data across machines by a key (user ID: even spread, but cross-user queries scatter; geography: locality, but hotspots). Last resort: you lose joins and cross-shard transactions, and resharding is painful.
When asked "how does this scale?", walking this ladder with numbers is a complete answer. Saying "shard it" at rung one is a red flag - most systems retire before rung 6.
Consistency and CAP, said correctly
CAP in two sentences: network partitions will happen, so partition tolerance is not optional. During a partition you choose - keep answering with possibly stale data (availability) or refuse/wait until nodes agree (consistency).
Saying "pick 2 of 3" is the classic mistake; the real choice is C vs A, only during a partition. The follow-up worth knowing: PACELC - even with no partition, you still trade latency vs consistency (waiting for replicas to confirm is slower).
| Level | Guarantee | Example that wants it |
|---|---|---|
| Strong consistency | Every read sees the latest write | Bank balance, inventory count, auction bids |
| Read-your-own-writes | You see your update; others may lag | Profile edits, posting a comment |
| Eventual consistency | All replicas converge, eventually | Like counts, follower counts, view tallies |
In the room, classify each piece of data: "the payment ledger needs strong consistency; the like counter can be eventual." Per-data-type reasoning is exactly what's being tested.
Idempotency, the favorite follow-up
Networks retry, so every operation may arrive twice. Design writes so a repeat is harmless: client sends a unique idempotency key, server stores processed keys and returns the saved result for duplicates. Any time you say "retry", say "idempotent" in the same breath.
Failure handling in one breath
Whatever you design, the next question is "what happens when X dies?" Have one sentence ready per pattern - the pattern name plus what it protects against.
- Timeouts: "Every remote call has a timeout - without one, a slow dependency freezes my whole thread pool."
- Retries with backoff and jitter: "Retry on failure, wait longer each time, add randomness so a thousand clients don't retry in sync and stampede the recovering service. Retries require idempotency."
- Circuit breaker: "After repeated failures, stop calling the broken service and fail fast; probe occasionally and close the circuit when it recovers. This stops cascading failure."
- Rate limiting: "Protect from abuse and overload. Token bucket allows short bursts; leaky bucket smooths to a constant rate; fixed window is simple but leaks 2x at window edges; sliding window fixes that for more memory."
- Graceful degradation: "Drop nice-to-have features under load instead of dying - serve the feed without recommendations, show cached prices."
- Redundancy: "No single point of failure: every tier has 2+ instances, the database has a failover replica, and I'd ask whether we need to survive a whole-region outage."
The complete answer to "what if the database goes down?" is a chain: "reads fail over to a replica; writes queue or fail fast behind a circuit breaker; the app degrades to cached data; alerts fire. We lose freshness, not the site."
Picking a communication style
Choose by the sentence that matches, and always answer sync vs async first: does the caller need the result to respond (sync), or can it acknowledge and process later (queue)? Async-by-default is a scalability superpower.
| The need | Pick | One-line why |
|---|---|---|
| Public API, standard CRUD | REST | Universal, cacheable, everyone's default |
| Many client types each wanting different fields | GraphQL | Clients query exactly what they need; costs server complexity |
| Internal service-to-service, high volume | gRPC | Binary + HTTP/2: fast, typed contracts |
| Two-way real-time (chat, games, live docs) | WebSocket | One open connection, both sides push |
| Server pushes one-way (live scores, notifications) | SSE | Push without WebSocket's complexity; auto-reconnects |
| Freshness within a minute is fine | Polling | Simplest thing that works; say it, then upgrade if pushed |
Common interview moment: "how do clients get updates?" Weak answer names a technology. Strong answer: "at this update frequency, polling every 30s is fine and simplest; if we need sub-second, I'd move to SSE - WebSockets only if clients also send in real time."
Classic problems: the one insight each
Most design questions are one of ~10 classics in costume. Each has a single insight that unlocks it and a trap that sinks candidates. Recognize the costume, lead with the insight.
| Problem | The unlocking insight | The trap |
|---|---|---|
| URL shortener | It's a key-value read problem: Base62-encode a counter (or pre-generate keys). Cache hot URLs. | Over-engineering hash collisions; forgetting 302-vs-301 (redirect type decides if you get analytics) |
| News feed | Fan-out on write (push to followers' feeds) for normal users; fan-out on read for celebrities. Hybrid. | One strategy for everyone - 100M-follower accounts make pure push explode |
| Chat | WebSockets + a per-conversation ordering key; store the message, then push it. | Forgetting offline users - undelivered messages need a queue/inbox, not just a socket |
| Video streaming | Transcode once into a quality ladder, chunk it, serve via CDN; player adapts bitrate. | Streaming video from your own servers or - worse - the database |
| Ride sharing | Geospatial index (geohash/quadtree) over constantly-updating driver locations; match nearby. | Storing lat/lng in a plain table and radius-scanning every request |
| Rate limiter | Token bucket per key in Redis, checked atomically (Lua). | Per-server counters - clients bypass the limit by hitting different servers |
| Notification system | A queue feeding per-channel workers (push/email/SMS), plus user preferences and dedupe. | Sending synchronously in the request path |
| Autocomplete | Precompute top-k suggestions per prefix; serve from memory; cache aggressively. | Running a LIKE query per keystroke |
| Payment system | Idempotency keys everywhere + an append-only double-entry ledger + reconciliation jobs. | Floats for money; assuming exactly-once delivery exists |
| Web crawler | It's a queue problem: URL frontier with per-domain politeness + a Bloom filter for "seen?" | Ignoring robots.txt/politeness and hammering one domain |
Sound senior: trade-off phrases
Interviewers grade judgment, and judgment sounds like conditions and crossover points, not technology names. Same knowledge, different sentence, different level.
| Sounds junior | Sounds senior |
|---|---|
| "I'd use Kafka." | "Uploads can complete async, so a queue decouples them; at our volume Redis Streams would do - Kafka earns its complexity if we need replay or many consumer groups." |
| "NoSQL scales better." | "Postgres covers this to tens of thousands of QPS. The signal to move is write volume on this one table - I'd peel it off to Cassandra then." |
| "I'd cache everything." | "The read ratio is 100:1, so caching the feed cuts 99% of DB reads. TTL of 30s is fine here because staleness is acceptable; the cart is not cacheable." |
| "Microservices, obviously." | "I'd start modular-monolith. I'd split a service out when a team or a scaling profile demands it - the search indexer is the first candidate." |
| "It's strongly consistent." | "The ledger is strongly consistent; counters are eventual. Paying the latency cost everywhere would be a mistake." |
| "That won't fail." | "When this fails - and it will - the breaker trips, we serve cached results, and we page someone. Degraded beats down." |
The pattern in every senior phrase: a number or condition, a default choice, and the trigger that would change it.
Red flags and rescue moves
Interviewers see the same failure modes every week. Each has a recovery line - the round is graded on trajectory, so a caught mistake costs almost nothing.
| Red flag | Rescue move |
|---|---|
| Drawing architecture in the first two minutes | Stop: "Let me back up and nail requirements first, or I'll build the wrong thing." |
| Over-engineering a small system | "Actually, at 12 requests/sec this is one server and a backup. Let me solve today's problem and note the scale path." |
| Buzzword with no depth ("we'd use Kafka") | Self-probe: "concretely, what that buys us here is X; the cost is Y." If you can't finish that sentence, don't name the tech. |
| Long silence while thinking | Narrate: "I'm weighing push vs pull for the feed - let me talk through both." |
| Steamrolling past hints | Interviewer hints are steering, not small talk. "Good point - let me reconsider that part." |
| No numbers anywhere | Anchor even roughly: "ballpark, 10M actives means ~100 QPS - so modest scale, and simplicity wins." |
| Ignoring failure until asked | Volunteer it at the end of high-level design: "before deep-diving - the single points of failure here are X and Y; here's the plan for each." |
Numbers to keep in your head
The one-screen recall table. If you remember nothing else, remember these.
| Number | Value |
|---|---|
| Seconds per day | ~100,000 (86,400) |
| 1M requests/day | ~12/sec (peak: ×2-5) |
| Memory read / SSD read / cross-continent round trip | 100 ns / 100 µs / 150 ms |
| Same-datacenter round trip | ~0.5 ms |
| 99.9% vs 99.99% uptime | ~9 hours vs ~1 hour of downtime per year |
| Good cache hit rate | 80-90%+ |
| Redis, single node | ~100k ops/sec |
| Postgres, single beefy node | ~10k+ simple queries/sec |
| Typical row / photo / minute of video | 1 KB / 500 KB / 10-50 MB |
| "Interactive" latency budget | ~200 ms end to end |
And the five sentences that carry the whole round: clarify before drawing; let the math pick the architecture; start simple and name the upgrade trigger; classify each data type's consistency; volunteer the failure story before being asked.