The interviewer names the system, and the first decision is what type of system it is. It is read-heavy, write-heavy, real-time fan-out, or batch, and that type sets the numbers and tradeoffs you will defend for the rest of the interview. Review this sheet in one sitting the night before, and check each row against the prompt during the interview.
Which type of system is the problem
Before you draw anything, answer three questions about the prompt:
- What is higher: reads, writes, or both at the same time?
- Do users need the result right now, or is near real time or delayed fine?
- Is the value in serving live traffic, or in processing data later?
The answers show what type of system the prompt is. Name it out loud in the interview. The rest of the design gets easier to defend when the interviewer can follow your reasoning.
| Type of system | Cues in the prompt | First component you name | Tradeoff you must state |
|---|---|---|---|
| Read-heavy (feed, catalog, lookup) | "users view", "look up", reads far outnumber writes | Cache in front of the store | Hit rate and invalidation; what a cache miss costs |
| Write-heavy, event-driven (orders, logs, clickstream) | "every action is recorded", "audit", "event" | Queue or append-only log in front of workers | Delivery semantics; at-least-once means idempotent consumers |
| Real-time fan-out (chat, live scores, notifications) | "push to all followers", "within seconds", "live" | Fan-out decision, then a gateway | Fan out on write versus on read; per-connection cost |
| Batch and analytics (reports, recommendations) | "end of day", "dashboard", "aggregate" | Ingestion store, then a processing job | Freshness versus cost; batch size versus run frequency |
If two types fit, say which one dominates and why. A prompt that mixes them (a feed with live updates) usually means the read path and the update path are separate designs, and naming that split early earns points.
The numbers an interviewer will ask you to defend
You do not need exact numbers. You need an estimate in the right range and arithmetic you can show.
Uptime (nines). Seconds per day is 86,400. Downtime budget per day = 86,400 x (1 - availability). Seconds per year is 31,536,000 (86,400 x 365).
| Availability | Downtime per day | Downtime per year |
|---|---|---|
| 99.9% (3 nines) | 86.4 s, about 1.4 minutes | 31,536 s, about 8.8 hours |
| 99.99% (4 nines) | 8.64 s | 3,153.6 s, about 53 minutes |
| 99.999% (5 nines) | 0.864 s | 315.36 s, about 5.3 minutes |
If the prompt says the service must be up 99.99% of the time, your downtime budget is 8.64 seconds per day. Check every design choice against that number: a synchronous call through a flaky dependency, a manual failover step, a deploy that takes ten minutes.
QPS. Requests per day divided by 86,400. Example: 10,000,000 daily active users x 20 requests per day = 200,000,000 requests per day. 200,000,000 / 86,400 = about 2,300 QPS average. Plan peak at 2 to 3 times average. That factor is a planning rule, not a measurement. In this example the peak is about 5,000 to 7,000 QPS.
Storage. Writes per day x size per write. Example: 1,000,000 photos per day x 1 MB = 1 TB per day, about 365 TB per year. If the prompt says data is kept forever, say so out loud. It changes the storage tier decision.
Bandwidth. QPS x response size. Example: 2,300 QPS x 50 KB per response = 115,000 KB/s, about 115 MB/s. If responses are large media, this number, not QPS, drives your network and CDN choices.
Fan-out. 1,000,000 users with 100 followers each. Fan out on write: one post becomes 100 writes, so 100,000 posts per day is 10,000,000 fan-out writes per day. Fan out on read: the write stays at 1, and every read gathers the posts from the reader's follow list. The cue that decides: huge follow counts or a write rate that dwarfs the read rate push you toward read-side fan-out plus caching.
Latency budget. Pick a budget (say 200 ms) and count sequential hops. Each hop is a round trip. A request that goes client, load balancer, app, database and back crosses three round trips. At 10 ms each, the network alone is 30 ms before the database does any work. Three calls that take 20 ms each cost 60 ms in sequence and about 20 ms in parallel. Every sequential dependency you add is a number you must defend.
Cache. The hit rate tells you how much read traffic reaches the database. At 100,000 read QPS with a 90% hit rate, the database takes 10,000 QPS, not 100,000. Then size the hot set: 1,000,000 hot objects x 1 KB each = 1 GB. If the hot set fits in memory, a cache in front of the database takes most of the read load. If it does not, you need a policy (TTL or eviction) and you say so.
Core terms with the definitions attached
- Availability: the fraction of time the system is up and responding correctly. The nines table is the math you quote when the prompt names a target.
- Consistency: the rule for how in sync multiple copies of data must be.
- Strong consistency: every read sees the latest write. The cue: money, balances, inventory counts, "exactly once".
- Eventual consistency: copies agree given time. The cue: feeds, profiles, reads where a second of staleness is acceptable.
- Idempotency: running an operation many times has the same effect as running it once. The cue: retries, at-least-once delivery, payment operations.
- RPO: how much data you can afford to lose, measured in time. An RPO of 5 minutes means losing at most 5 minutes of work is acceptable.
- RTO: how long you can be down before the damage is unacceptable. An RTO of 1 hour means you must be serving again within an hour.
Reliability: attach a number and a component name
A reliability answer is a pair: the failure you are handling, and the number that bounds it.
- S3 provides 99.99% availability and 99.999999999% (11 nines) data durability by default (AWS S3). Durability is about data loss; availability is about uptime. When the prompt names an availability target, the number you quote is the availability one.
- An RDS Multi-AZ DB instance deployment keeps a synchronous standby replica in a different Availability Zone, which AWS provisions and maintains automatically (AWS RDS documentation). On the cost side, write and commit latency can be higher than a single-AZ deployment because of the synchronous replication (same source). State which side of the tradeoff your prompt needs.
- Replication topology is a decision you name. A single leader accepts writes and replicates to the others. That gives you one write path and simple failover. Leaderless has no leader: any node accepts writes, and a read is valid only when it meets a quorum. The rule is arithmetic: write quorum W plus read quorum R must exceed the node count N, so W + R > N.
- Delivery semantics are a guarantee you choose. Some queue implementations deliver at least once, so a consumer can see the same message twice after a retry. At-least-once delivery means the consumer must be idempotent, and the idempotency key is the event ID.
Vendor numbers change. Recheck the official page before you quote one in an interview.
The tradeoffs to say out loud, and the cue that picks them
The pattern is the same every time: name the tradeoff, state which side you picked, and cite the cue in the prompt that picked it.
| Tradeoff | Pick the left side when the prompt says | Pick the right side when the prompt says |
|---|---|---|
| SQL versus NoSQL | "joins", "transactions", "audit trail" | "flexible schema", "one table must scale horizontally", "eventually consistent reads" |
| Strong versus eventual consistency | "money moves", "balances", "exactly once" | "feed", "profile", "a second of staleness is fine" |
| Synchronous versus asynchronous | "the user needs the result before continuing" | "another service acts later", "audit", "notification" |
| Vertical versus horizontal scaling | "one machine's CPU or memory is the limit" | "stateless compute must grow", "data will not fit one node" |
| Fan-out on write versus on read | "read rate dwarfs write rate", "follow counts are modest" | "write rate dwarfs read rate", "follow counts are huge", "writes are expensive or must stay exactly once" |
| Durable event versus best-effort event | "no event may be lost", "payment", "order" | "telemetry", "analytics", "best effort is acceptable" |
Every row in the table is a default, and a prompt constraint can flip it. A payments feed that must be both strongly consistent and highly available is the case where you say the conflict out loud and pick the side the business cares about, with the number attached (the RPO if you pick availability, the audit cost if you pick consistency).
A worked design step: making event publication durable
Requirement. An orders API commits an order to the database and must publish an OrderCreated event to a broker. The broker is sometimes down. Committed orders must not lose their event, and a redelivered event must not apply twice.
The weak version, and why it fails. Publish the event in the same request and retry until the broker accepts. A retry inside the request cannot outlive a broker outage longer than the request timeout, and even when it does, the database write and the publish are two separate steps. A crash between the two loses the event. When you present the weak version, say which step in the design makes the event survive a crash.
The decision. Write the order and an outbox record in one database transaction. A separate publisher reads the outbox, sends the row to the broker, and marks it published. The order and the intent to publish commit together, so a crash between the two can no longer drop the event. The publisher retries safely, because the outbox is the source of truth and a row that was sent but not marked can be resent. The consumer is idempotent on the event ID, so a redelivery is a no-op.
The tradeoff the interviewer will probe. You added a background publisher and a second write inside the request transaction. What it costs: the publisher's index scan, and a publish delay measured in outbox drain time instead of request time. Say the order rate is 1,000 per minute, about 17 per second. A publisher that drains the outbox every 100 ms sends batches of about 2 events, so the steady-state publish lag is bounded by that 100 ms interval plus one batch's send time. If the event is best-effort telemetry, the outbox is extra machinery and inline publish is the right call. If the consumer needs strict per-order ordering, partition events by order ID so one order's events arrive in order.
The weak answer fails because it treats the network call as the durable step. The broker is a delivery mechanism, and you make it safe with the outbox and with idempotent consumers.
The 16 decision areas to check the night before
CramHQ's system design interview path organizes practice around 16 decision areas. Use the list as a checklist the night before. If you cannot state the tradeoff for an area in one sentence, that area goes on your review list for the week.
- Requirements and scope: turn the ambiguous prompt into explicit requirements and constraints.
- Capacity and constraint: traffic, concurrency, storage, bandwidth, fan-out, compute.
- API, data model, and access pattern: where the source of truth lives, indexes, schema.
- Service communication and traffic: protocols, gateways, load balancing, health checks, failover.
- Read path and caching: cache placement, invalidation, CDNs, hot keys.
- Write path and async processing: queues, workers, retries, dead-letter handling, backpressure.
- Partitioning and hotspot: the partition key, hot partitions, rebalancing.
- Consistency and concurrency: consistency models, idempotency, transactions, sagas.
- Reliability and degradation: failure modes, graceful degradation, replay, disaster recovery.
- Performance and tail latency: the critical path, precomputation, bounded fan-out.
- Observability and operations: SLOs, metrics, logs, traces, alerts.
- Security, abuse, and privacy: where authentication and authorization sit, rate limits, isolation.
- Cost and operational tradeoff: the dominant cost driver, build versus buy.
- Migration and rollout: dual writes, backfills, cutovers, rollback.
- Specialized systems: search, media, time series, ranking, rate control.
- Synthesis: the highest-leverage repair, and defending the tradeoffs across the whole design.
If you want the list graded instead of self-checked, take the free system design assessment. No card required. It returns a Readiness Report with your top decision gaps and a targeted repair preview. The full interview path adds the Decision Drills and the System Design Readiness Exam for $49.99 (list $74.99), with 12 months of access.
How to use this the night before
One pass before bed: pick one prompt you have practiced, and recompute its three numbers (QPS, storage, bandwidth) from scratch in under 3 minutes. If you cannot reproduce the arithmetic without looking, redo it before the interview.
The morning of: classify five prompts into the four types of system out loud, and state one tradeoff for each. Keep it to a minute each.
After the interview or the next day: take a 45-minute timed prompt and grade yourself against the 16-area list. The areas where you could not state a tradeoff are your study list for the next week, in the order they came up.
Then run the loop again with a new prompt.
