Numbers worth memorising
A back-of-the-envelope estimate is only useful because of the sentence that follows it. These are the constants that let you get to that sentence quickly, plus the specific figures from this playbook that are worth having ready.
1. Latency, ordered
| Operation | Time | Relative |
|---|---|---|
| L1 cache reference | 1 ns | 1× |
| Branch misprediction | 3 ns | |
| L2 cache reference | 4 ns | |
| Mutex lock/unlock | 17 ns | |
| Main memory reference | 100 ns | 100× L1 |
| Compress 1 KB (Snappy) | 2 μs | |
| Read 1 MB sequentially from RAM | 3 μs | |
| Round trip within a datacenter | 500 μs | |
| Read 1 MB sequentially from NVMe SSD | 50 μs | |
| SSD random read | 16 μs | |
| Read 1 MB sequentially from disk (HDD) | 825 μs | |
| Disk seek | 2 ms | |
| Round trip, same region | 1–2 ms | |
| Round trip, cross-continent (US↔EU) | 80–100 ms | |
| Round trip, antipodal (US↔Australia) | 150–200 ms |
The two that matter most in interviews:
- A datacenter round trip is ~0.5ms; a cross-continent round trip is ~100ms — 200× more. This single ratio justifies edge caching, regional replicas, local rate-limit enforcement, and asynchronous cross-region reconciliation.
- Memory is ~100× faster than SSD and ~10,000× faster than a disk seek. This is why "keep the index in RAM" appears in almost every design here.
2. Throughput rules of thumb
| Component | Rough capacity per node | Notes |
|---|---|---|
| Redis | 100k+ ops/sec/core, 1M+ with pipelining | Single-threaded per shard; Lua scripts are atomic |
| Postgres | 10–50k simple queries/sec | Falls sharply with contention on hot rows |
| Cassandra/Scylla | 50–200k writes/sec/node | Excellent for append-heavy, point-read workloads |
| Kafka | 100k–1M msg/sec/broker | Sequential disk writes; partition count is the parallelism unit |
| Elasticsearch | 1–10k queries/sec/node | Depends enormously on query complexity |
| Nginx/Envoy | 50–100k req/sec/node | Proxying, not application work |
| WebSocket server (Go/Erlang) | 100k–2M connections/node | Bounded by RAM per connection, ~2–10 KB |
| HTTP app server | 1–10k req/sec/node | Application logic dominates |
| ClickHouse | Billions of rows/sec scanned | Columnar; the scan rate is the useful number |
Use these as order-of-magnitude anchors, and say so. "Redis does roughly 100k ops/sec per shard, so 4M checks/sec needs about 40 shards plus headroom" is a good sentence; claiming precision you do not have is not.
3. Capacity arithmetic
Seconds in a day 86,400 ≈ 10^5
Seconds in a year 31.5M ≈ 3 × 10^7
Requests/day -> QPS divide by 10^5
Peak factor 3-5× average (diurnal); 40× for event-driven spikes
1 million/day ≈ 12 QPS
1 billion/day ≈ 12,000 QPS
1 trillion/year ≈ 32,000 QPS
Storage
1 KB × 1M/day = 1 GB/day ≈ 365 GB/year
1 KB × 1B/day = 1 TB/day ≈ 365 TB/year
UUID 16 bytes Timestamp 8 bytes
Typical row ~100-500 bytes with indexes
Index overhead add 20-50% on top of raw data
Bandwidth
1 Gbps = 125 MB/s = ~10 TB/day
1 MB/s sustained = ~86 GB/day
Video at 2 Mbps: one stream = 0.9 GB/hour
Compression ratios worth knowing, because they change conclusions:
| Data | Typical ratio | Note |
|---|---|---|
| Text/JSON (gzip/zstd) | 5–10× | Almost always worth it |
| Columnar analytics (Parquet + zstd) | 10–20× | Plus column pruning at query time |
| Time-series floats (Gorilla) | ~12× | 16 B/point → ~1.37 B/point |
| Already-compressed media | 1× | Do not bother |
4. Memory footprints of the usual structures
| Structure | Size | Use |
|---|---|---|
| Bloom filter, 1% FP | 9.6 bits/element | 1B URLs = 1.2 GB |
| Bloom filter, 0.1% FP | 14.4 bits/element | Trade RAM for fewer disk checks |
| HyperLogLog, 2% error | 12 KB, any cardinality | Unique counts; mergeable |
| Count-Min Sketch | w × d × 4 B; ε = e/w, δ = e^-d | 2^20 × 5 × 4 B = 20 MB |
| Roaring bitmap (dense ints) | ~0.5–2 bits/element | Exact, supports deletion |
| t-digest / DDSketch | 1–5 KB | Mergeable percentiles |
| Java object overhead | 16 B header + padding | Why 50M objects ≠ 50M × payload |
| Redis key overhead | ~50–100 B/key | Small values are dominated by overhead |
Say the trade, not just the size. "A Bloom filter at 1% costs 9.6 bits per key, so a billion URLs fits in 1.2 GB of RAM — but a false positive means we never crawl a real URL, so I verify against RocksDB rather than trusting the filter alone."
5. Figures from these designs
Numbers that already did work on a page, ready to reuse:
| Figure | Where | What it proves |
|---|---|---|
| 62^7 ≈ 3.5 trillion | Bitly | 7 Base62 chars ≈ 95 years of runway at 100M/day |
| 400M followers ÷ 100k writes/s = 66 min | Why pure fan-out-on-write dies on celebrities | |
| 100k comments/s × 5M viewers = 5×10^11 msg/s | Live comments | Why sampling is the design, not an optimisation |
| Manhattan vs rural Kansas ≈ 600,000× density | Yelp | Why no fixed grid resolution works |
| 1.25M location writes/s vs 230 rides/s | Uber | Two storage strategies in one system |
| ~900 PB/day egress | YouTube | Why owning the CDN is existential |
| ~1.37 bytes/point (Gorilla) | Metrics | 138 TB/day → 12 TB/day |
| 115 TB exact vs 16 GB sketched | YouTube Top K | The case for probabilistic counting |
| 30M segments → ~200 candidates | Strava | 150,000× pruning from an R-tree |
| 12B fetches/day → 240M | Price tracking | 50× from adaptive polling |
| ~320 KB KV per token | ChatGPT | Concurrency is bounded by VRAM, not FLOPs |
| 2M concurrent users, 50k total writes | Ticketmaster | Contention, not throughput |
| 500M connections × 10 KB = 5 TB RAM | Why the runtime choice matters | |
| RS(10,4) = 1.4× vs 3× replication | Dropbox | 480 PB instead of 900 PB |
6. Availability arithmetic
| SLA | Downtime/year | Downtime/month | Downtime/week |
|---|---|---|---|
| 99% | 3.65 days | 7.2 hours | 1.7 hours |
| 99.9% | 8.77 hours | 43 min | 10 min |
| 99.95% | 4.38 hours | 22 min | 5 min |
| 99.99% | 52.6 min | 4.4 min | 1 min |
| 99.999% | 5.26 min | 26 s | 6 s |
Two consequences worth stating:
- Dependencies multiply. Five services at 99.9% each, all required, give 99.5% — 1.8 days/year. This is the argument for partial results and graceful degradation, not for making every dependency more reliable.
- 99.99% leaves 4.4 minutes a month. That is less than most deploy windows and less than a typical failover. Achieving it means the failover itself must be automatic and fast, which is a design requirement, not an ops aspiration.
7. Cost anchors
Order-of-magnitude only, but enough to reason about which decisions are expensive:
CDN egress $0.01-0.08 / GB (heavily negotiated at volume)
Cloud egress $0.05-0.09 / GB (the reason people build their own CDN)
Object storage $0.02 / GB-month (standard); $0.004 archive
Block storage (SSD) $0.10 / GB-month (5× object storage)
Compute $0.03-0.05 / vCPU-hour
GPU (H100 class) $2-10 / GPU-hour (the dominant cost in LLM serving)
Managed database 3-10× raw compute for the same work
SMS ~$0.0075 / message (vs ~$0.0000004 for a push)
The generalisable insight: at scale, bandwidth and GPUs dominate; storage is cheap; compute is in between. That ordering explains why YouTube optimises codecs before servers, why ChatGPT is a memory-allocation problem, and why Dropbox chooses erasure coding over replication.
8. Using numbers well
- Round aggressively. 86,400 is 10^5. 3.15×10^7 is 3×10^7. Nobody wants four significant figures on an assumption.
- Always follow the number with a conclusion. "40 TB/year — comfortable for a sharded relational store" is useful; "40 TB/year" alone is trivia.
- State the peak factor explicitly. Average QPS is a planning fiction; the system has to survive the peak, and for event-driven systems the peak can be 40× the mean.
- Sanity-check against something you know. If your estimate says one server handles all of Twitter, or that you need a million machines to serve a blog, redo it.
- Say which numbers you are unsure about. "I'm assuming 2 MB per post including thumbnails — if it's 10× that, the CDN bill dominates and I'd revisit the media strategy" is a stronger answer than false confidence.