LearnThatStack Ace your next interview

Software Architecture & System Design.
Interview cheat sheet.

Quick reference for Software Architecture & System Design - sectioned for fast scanning. Skim the part you're shaky on, walk in confident.

System Design Concepts 14-section reference ~13 min read

Overview

This is the run-of-show for the system design round: what to say, what to decide, and the numbers to have ready. Skim it the night before; use the focused sheets in this category for depth on any single topic.

The 45-minute playbook

The interview is a structured conversation with a clock. Interviewers grade how you drive it, not just what you draw. Follow this timeline and you will never run out of time with an empty whiteboard.

Minutes Phase What the interviewer is grading The mistake that sinks people
0-5 Clarify requirements Do you build the right thing, or the thing in your head? Jumping straight to boxes and arrows
5-10 Estimate scale Can you turn vague scope into numbers that drive decisions? Skipping math, then guessing at architecture
10-20 High-level design Can you draw a simple end-to-end system that works? Adding Kafka and shards before anything works
20-35 Deep dive Real depth in 1-2 components (interviewer usually picks) Staying shallow everywhere; touring buzzwords
35-43 Scale and failure Do you know what breaks first and what to do about it? "It scales because it's microservices"
43-45 Wrap up Can you summarize trade-offs and name what you'd do next? Trailing off instead of closing

Two habits that raise every phase: think out loud (a silent minute feels like five), and check in ("does this match what you had in mind, or should I go deeper here?").

Clarifying questions that score points

The first five minutes are free points. Interviewers deliberately give vague prompts to see if you ask. Ask these, then write the answers in a corner of the board.

  1. Who uses this and what is the core action? (One sentence. "Users shorten a URL and share it.")
  2. How many users? Daily actives, not registered accounts.
  3. Read-heavy or write-heavy? Ask for the ratio. This one answer drives caching, replicas, and DB choice.
  4. How fast must it feel? Interactive (under ~200 ms) vs background (seconds is fine).
  5. How correct must it be? Is stale data acceptable for a few seconds, or never (money, inventory)?
  6. What scale of data? Size per item and how long we keep it.
  7. Any hard constraints? Regions, compliance, mobile-only, budget.
  8. What is explicitly out of scope? Say it out loud so you are not graded on it.

When the interviewer says "you decide": decide, state it as an assumption, and move on. "I'll assume 10M daily actives and 100:1 reads to writes - tell me if you want different numbers." Deciding quickly with stated assumptions reads as senior; fishing for hints reads as junior.

Back-of-envelope math that never fails

You need about six numbers and one method. Round aggressively - the goal is the right power of ten, not the right digit.

The shortcuts

Fact Use it for
1 day ≈ 100,000 seconds (really 86,400) Requests/day → requests/second
1M requests/day ≈ 12/sec Instant QPS conversion
Peak ≈ 2-5x average Size for peak, not average
1 char ≈ 1 byte; typical DB row or tweet ≈ 1 KB Storage per item
Photo ≈ 500 KB, minute of video ≈ 10-50 MB Media-heavy estimates
Million × KB = GB; billion × KB = TB Skip the zeros: 10^6 × 10^3 bytes = 10^9

The latency ladder

The only hardware numbers worth memorizing:

Operation Time Takeaway
Read from memory ~100 ns Memory is effectively free
Read from SSD ~100 µs 1,000x slower than memory
Round trip, same datacenter ~0.5 ms Chatty services add up fast
Disk seek (HDD) ~10 ms Avoid on hot paths
Round trip, across continents ~150 ms Why CDNs and regions exist

Worked example, 60 seconds

Photo app: 10M daily actives, each views 50 photos and uploads 2/day at 500 KB.

  • Write QPS: 20M uploads / 100k sec = 200/sec, peak ~1,000/sec. Modest.
  • Read QPS: 500M views / 100k sec = 5,000/sec, peak ~25,000/sec. Read-heavy: cache + CDN.
  • Storage: 20M × 500 KB = 10 TB/day ≈ 3.7 PB/year. Object storage, not a database.

Close with a sanity check: "so this is a read-heavy system whose hard problem is storage and delivery, not write throughput." That sentence - math driving the design - is what the estimate phase is actually for.

Building blocks at a glance

Every design is the same dozen blocks arranged differently. For each one, know its one-line job, when to reach for it, and the trade-off to say out loud when you add it.

Block One-line job Reach for it when Trade-off to say out loud
Load balancer Spread requests across identical servers More than one app server Servers must be stateless (sessions move to Redis)
CDN Serve static files from near the user Global users, images/video/JS Cache invalidation delay on updates
Cache (Redis) Answer repeat reads from memory Read-heavy, tolerates slight staleness Invalidation strategy and cold-start misses
Message queue Absorb work now, process it later Slow work in the request path; traffic spikes Eventual completion; consumers must be idempotent
Relational DB Source of truth with transactions Almost always, at the start Vertical limits first; scaling reads is easy, writes hard
Key-value store Tiny values by exact key, huge scale Sessions, counters, rate limits No queries, no joins - key lookups only
Wide-column store Massive write volume, time-ordered data Feeds, logs, messages at scale Design tables around queries; no ad-hoc queries later
Object store (S3) Cheap, durable files of any size Anything over ~1 MB: images, video, backups Store the URL in the DB, never the file
Search index Text search and filtering "Search by keyword" appears in requirements A copy of the data, seconds behind the source of truth
WebSocket gateway Hold open connections, push to clients Chat, live updates, presence Connections are state: need sticky routing or a registry
Scheduler / cron Run work on a timer Cleanup, digests, retries, reports Needs a lock or leader so two workers don't double-run

Rule of thumb: introduce a block only when a number or requirement demands it, and say which one. "At 25k reads/sec we need a cache" beats "let's add Redis" every time.

Choosing a database in 60 seconds

Say the senior line first: "I'd start with one Postgres instance - it handles more than people think - and here is what would force a change." Then choose by access pattern, not by fashion.

Your data looks like Pick Because
Entities with relations, transactions, flexible queries Postgres / MySQL Joins, ACID, 50 years of tooling
Self-contained documents, shape varies per record Document DB (MongoDB) Schema flexibility, reads whole objects
Billions of tiny lookups by exact key Key-value (Redis / DynamoDB) Predictable ~ms reads at any scale
Relentless writes, time-ordered, known queries Wide-column (Cassandra) Write-optimized, linear scale-out
Keyword search, ranking, filters Search index (Elasticsearch) Inverted index; runs beside the real DB
Files and media Object store + metadata row in SQL DBs are terrible at blobs

Signals that force a move off single-Postgres, in the order they usually arrive: read load (add replicas - still Postgres), dataset too big for one machine or write QPS in the tens of thousands (shard, or move that table to a wide-column store), one access pattern dominating (peel it off to a specialized store).

The trap to avoid: picking NoSQL "for scale" on day one, then discovering you need transactions and joins. Consistency needs decide too - money and inventory want ACID; likes and view counts don't.

The scaling ladder

Scale in this order. Each rung is cheaper and less complex than the next; climbing early is how systems get over-engineered. Name the rung, say what breaks next.

  1. One bigger machine (vertical). Boring and correct. Breaks: hardware ceiling, single point of failure.
  2. Stateless app servers behind a load balancer. Move sessions to Redis first. Breaks: the database is now the bottleneck.
  3. Cache + CDN. Kill repeat reads before they hit the DB; 80-90%+ hit rates are normal. Breaks: writes and cache misses still land on one DB.
  4. Read replicas. Writes go to the leader, reads to followers. Breaks: replication lag (read-your-own-writes bugs), and write volume still hits one machine.
  5. Queues for async work. Take everything slow out of the request path; absorb spikes. Breaks: eventually nothing - until raw write volume does.
  6. Shard the database. Split data across machines by a key (user ID: even spread, but cross-user queries scatter; geography: locality, but hotspots). Last resort: you lose joins and cross-shard transactions, and resharding is painful.

When asked "how does this scale?", walking this ladder with numbers is a complete answer. Saying "shard it" at rung one is a red flag - most systems retire before rung 6.

Consistency and CAP, said correctly

CAP in two sentences: network partitions will happen, so partition tolerance is not optional. During a partition you choose - keep answering with possibly stale data (availability) or refuse/wait until nodes agree (consistency).

Saying "pick 2 of 3" is the classic mistake; the real choice is C vs A, only during a partition. The follow-up worth knowing: PACELC - even with no partition, you still trade latency vs consistency (waiting for replicas to confirm is slower).

Level Guarantee Example that wants it
Strong consistency Every read sees the latest write Bank balance, inventory count, auction bids
Read-your-own-writes You see your update; others may lag Profile edits, posting a comment
Eventual consistency All replicas converge, eventually Like counts, follower counts, view tallies

In the room, classify each piece of data: "the payment ledger needs strong consistency; the like counter can be eventual." Per-data-type reasoning is exactly what's being tested.

Idempotency, the favorite follow-up

Networks retry, so every operation may arrive twice. Design writes so a repeat is harmless: client sends a unique idempotency key, server stores processed keys and returns the saved result for duplicates. Any time you say "retry", say "idempotent" in the same breath.

Failure handling in one breath

Whatever you design, the next question is "what happens when X dies?" Have one sentence ready per pattern - the pattern name plus what it protects against.

  • Timeouts: "Every remote call has a timeout - without one, a slow dependency freezes my whole thread pool."
  • Retries with backoff and jitter: "Retry on failure, wait longer each time, add randomness so a thousand clients don't retry in sync and stampede the recovering service. Retries require idempotency."
  • Circuit breaker: "After repeated failures, stop calling the broken service and fail fast; probe occasionally and close the circuit when it recovers. This stops cascading failure."
  • Rate limiting: "Protect from abuse and overload. Token bucket allows short bursts; leaky bucket smooths to a constant rate; fixed window is simple but leaks 2x at window edges; sliding window fixes that for more memory."
  • Graceful degradation: "Drop nice-to-have features under load instead of dying - serve the feed without recommendations, show cached prices."
  • Redundancy: "No single point of failure: every tier has 2+ instances, the database has a failover replica, and I'd ask whether we need to survive a whole-region outage."

The complete answer to "what if the database goes down?" is a chain: "reads fail over to a replica; writes queue or fail fast behind a circuit breaker; the app degrades to cached data; alerts fire. We lose freshness, not the site."

Picking a communication style

Choose by the sentence that matches, and always answer sync vs async first: does the caller need the result to respond (sync), or can it acknowledge and process later (queue)? Async-by-default is a scalability superpower.

The need Pick One-line why
Public API, standard CRUD REST Universal, cacheable, everyone's default
Many client types each wanting different fields GraphQL Clients query exactly what they need; costs server complexity
Internal service-to-service, high volume gRPC Binary + HTTP/2: fast, typed contracts
Two-way real-time (chat, games, live docs) WebSocket One open connection, both sides push
Server pushes one-way (live scores, notifications) SSE Push without WebSocket's complexity; auto-reconnects
Freshness within a minute is fine Polling Simplest thing that works; say it, then upgrade if pushed

Common interview moment: "how do clients get updates?" Weak answer names a technology. Strong answer: "at this update frequency, polling every 30s is fine and simplest; if we need sub-second, I'd move to SSE - WebSockets only if clients also send in real time."

Classic problems: the one insight each

Most design questions are one of ~10 classics in costume. Each has a single insight that unlocks it and a trap that sinks candidates. Recognize the costume, lead with the insight.

Problem The unlocking insight The trap
URL shortener It's a key-value read problem: Base62-encode a counter (or pre-generate keys). Cache hot URLs. Over-engineering hash collisions; forgetting 302-vs-301 (redirect type decides if you get analytics)
News feed Fan-out on write (push to followers' feeds) for normal users; fan-out on read for celebrities. Hybrid. One strategy for everyone - 100M-follower accounts make pure push explode
Chat WebSockets + a per-conversation ordering key; store the message, then push it. Forgetting offline users - undelivered messages need a queue/inbox, not just a socket
Video streaming Transcode once into a quality ladder, chunk it, serve via CDN; player adapts bitrate. Streaming video from your own servers or - worse - the database
Ride sharing Geospatial index (geohash/quadtree) over constantly-updating driver locations; match nearby. Storing lat/lng in a plain table and radius-scanning every request
Rate limiter Token bucket per key in Redis, checked atomically (Lua). Per-server counters - clients bypass the limit by hitting different servers
Notification system A queue feeding per-channel workers (push/email/SMS), plus user preferences and dedupe. Sending synchronously in the request path
Autocomplete Precompute top-k suggestions per prefix; serve from memory; cache aggressively. Running a LIKE query per keystroke
Payment system Idempotency keys everywhere + an append-only double-entry ledger + reconciliation jobs. Floats for money; assuming exactly-once delivery exists
Web crawler It's a queue problem: URL frontier with per-domain politeness + a Bloom filter for "seen?" Ignoring robots.txt/politeness and hammering one domain

Sound senior: trade-off phrases

Interviewers grade judgment, and judgment sounds like conditions and crossover points, not technology names. Same knowledge, different sentence, different level.

Sounds junior Sounds senior
"I'd use Kafka." "Uploads can complete async, so a queue decouples them; at our volume Redis Streams would do - Kafka earns its complexity if we need replay or many consumer groups."
"NoSQL scales better." "Postgres covers this to tens of thousands of QPS. The signal to move is write volume on this one table - I'd peel it off to Cassandra then."
"I'd cache everything." "The read ratio is 100:1, so caching the feed cuts 99% of DB reads. TTL of 30s is fine here because staleness is acceptable; the cart is not cacheable."
"Microservices, obviously." "I'd start modular-monolith. I'd split a service out when a team or a scaling profile demands it - the search indexer is the first candidate."
"It's strongly consistent." "The ledger is strongly consistent; counters are eventual. Paying the latency cost everywhere would be a mistake."
"That won't fail." "When this fails - and it will - the breaker trips, we serve cached results, and we page someone. Degraded beats down."

The pattern in every senior phrase: a number or condition, a default choice, and the trigger that would change it.

Red flags and rescue moves

Interviewers see the same failure modes every week. Each has a recovery line - the round is graded on trajectory, so a caught mistake costs almost nothing.

Red flag Rescue move
Drawing architecture in the first two minutes Stop: "Let me back up and nail requirements first, or I'll build the wrong thing."
Over-engineering a small system "Actually, at 12 requests/sec this is one server and a backup. Let me solve today's problem and note the scale path."
Buzzword with no depth ("we'd use Kafka") Self-probe: "concretely, what that buys us here is X; the cost is Y." If you can't finish that sentence, don't name the tech.
Long silence while thinking Narrate: "I'm weighing push vs pull for the feed - let me talk through both."
Steamrolling past hints Interviewer hints are steering, not small talk. "Good point - let me reconsider that part."
No numbers anywhere Anchor even roughly: "ballpark, 10M actives means ~100 QPS - so modest scale, and simplicity wins."
Ignoring failure until asked Volunteer it at the end of high-level design: "before deep-diving - the single points of failure here are X and Y; here's the plan for each."

Numbers to keep in your head

The one-screen recall table. If you remember nothing else, remember these.

Number Value
Seconds per day ~100,000 (86,400)
1M requests/day ~12/sec (peak: ×2-5)
Memory read / SSD read / cross-continent round trip 100 ns / 100 µs / 150 ms
Same-datacenter round trip ~0.5 ms
99.9% vs 99.99% uptime ~9 hours vs ~1 hour of downtime per year
Good cache hit rate 80-90%+
Redis, single node ~100k ops/sec
Postgres, single beefy node ~10k+ simple queries/sec
Typical row / photo / minute of video 1 KB / 500 KB / 10-50 MB
"Interactive" latency budget ~200 ms end to end

And the five sentences that carry the whole round: clarify before drawing; let the math pick the architecture; start simple and name the upgrade trigger; classify each data type's consistency; volunteer the failure story before being asked.

Found this useful? Pass it on.
Pro · $10/mo

The sheet is free. Pro goes deeper.

Pro opens the full question library behind every sheet, every refresher and a monthly AI allowance. One subscription, all formats.

Full question library All refreshers Cancel anytime