Skip to main content

Choosing the technology

Naming a technology is only worth credit if you can say what it buys you and what it costs. This page is the short version of those arguments, organised by the decision you are actually making.

The meta-rule first: pick the boring option until a specific requirement rules it out, and say which requirement that is. "Postgres until we exceed roughly 50k writes/sec on this table, at which point the sharding operational cost is worth paying" is a far better answer than reaching for Cassandra in minute three.

1. Databases​

You needChooseBecauseWatch out for
Transactions, joins, correctnessPostgreSQLACID, rich indexing, mature. Handles far more load than people assumeSingle-writer; sharding is manual and painful
Global transactions across regionsSpanner / CockroachDBExternal consistency with distributed transactionsLatency cost of consensus on every write
High write throughput, point readsCassandra / ScyllaDBLSM trees, linear scaling, no single writerNo joins; you must design the partition key for the query
Managed key-value at any scaleDynamoDBPredictable single-digit-ms, conditional writes, TTLCost model punishes scans; global tables are last-writer-wins
Analytical queries over billions of rowsClickHouse / Pinot / DruidColumnar, vectorised, sub-second GROUP BYNot for point mutations; eventual consistency
Full-text and multi-attribute searchElasticsearch / OpenSearchInverted index + BKD trees + geo in one queryOperationally demanding; not a system of record
Time seriesTimescaleDB / VictoriaMetrics / M3Compression, retention tiers, downsampling built inCardinality is the failure mode, not volume
Financial ledger at high throughputTigerBeetlePurpose-built double-entry with enforced invariantsNarrow purpose; not a general database
Vectors / similarityFaiss / Milvus / pgvectorANN indexes over embeddingsIndex rebuild cost; recall/latency trade-off
Cache and data structuresRedisSorted sets, HLL, Lua atomicity, Pub/SubSingle-threaded per shard; memory-bound

Choosing a partition key is the real decision, and it is worth more than the product name. Say it out loud: "partitioned by namespace_id, because every query is scoped to a namespace and a namespace never spans shards."

2. Queues and streams​

You needChooseBecause
Replayable, ordered, high-throughput logKafkaPartitions give ordering + parallelism; retention enables replay; the substrate for stream processing
Simple managed work queueSQSZero operations, native delay queues, dead-letter queues
Complex routing, per-message ackRabbitMQExchanges, priorities, DLX for delayed retry
Ultra-low-latency in-process handoffAeron / Disruptor / ring buffersNanoseconds; used inside matching engines
Pub/Sub fan-out, loss-tolerantRedis Pub/Sub / NATSCheap and fast; at-most-once, so never for durable work

The distinction to state: Kafka is a log (replayable, ordered per partition, consumers track their own offset); SQS and RabbitMQ are queues (a message is consumed and gone). If you need replay, audit, or multiple independent consumers, you need a log.

3. Stream processing​

You needChooseBecause
Event-time windows, large keyed state, exactly-onceFlinkWatermarks, RocksDB state, 2PC sinks. The default for correctness-sensitive streaming
Lightweight, Kafka-nativeKafka StreamsA library, not a cluster; good when the topology is simple
Unified batch and streamSpark Structured StreamingReuses batch code; micro-batch latency (seconds, not milliseconds)
Simple stateless transformationPlain consumersDo not deploy Flink to filter a topic

The question that decides it: do you need state, and does that state need to survive? Stateless filtering needs no framework. Windowed aggregation with exactly-once semantics needs Flink.

4. Coordination​

You needChooseNotes
Leader election, leases, config watchetcdRaft, linearizable CAS, ModRevision is a free fencing token
The same, in a JVM ecosystemZooKeeperMature; ephemeral znodes and watches; heavier to operate
Advisory locks with a TTLRedis (SET NX PX)Fast but not safe under process pauses without a fence
Durable multi-step workflowsTemporal / CadenceDurable execution; the workflow's own state survives worker death

Never use a Redis lock alone for correctness. It gives mutual exclusion while your process is healthy, which is not the case you are defending against. Pair it with a fencing token or use etcd.

5. Caching​

LayerChooseNotes
In-process (L1)Caffeine (JVM) / Ristretto (Go)W-TinyLFU admission control; the fix for hot keys
Distributed (L2)Redis Cluster or MemcachedRedis for data structures; Memcached for multi-threaded raw KV throughput
EdgeCloudFront KVS / Workers KVSub-millisecond, tiny, eventually consistent. See the specifics
HTTPCDN with correct Cache-ControlThe cheapest cache is the one you do not operate

6. Compute and isolation​

You needChooseNotes
Run untrusted codeFirecracker microVM or gVisorContainers share the host kernel — insufficient alone
Long-lived servicesKubernetesOnly when you have enough services to justify it
Bursty, short jobsLambda / Cloud RunCold starts matter; not for latency-critical paths
Cheap idempotent batch workSpot / preemptible instances60–90% cheaper; only for retryable, checkpointable work
GPU inferencevLLM / TensorRT-LLM / SGLangPagedAttention and continuous batching are table stakes

7. Anti-patterns​

Things that read as inexperience:

  • Naming a technology with no justification. "We'll use Kafka" without saying what ordering, replay, or decoupling property you need.
  • A graph database for a social graph at scale. Facebook built TAO — a cache over sharded MySQL with four query shapes — precisely because general graph databases do not survive 10^9 point lookups per second.
  • MongoDB as a default. Not wrong, but rarely the reason a design works; if it is your answer, say what document-model property you are using.
  • Microservices for everything. Each service boundary is a network call, a failure mode, and a deployment. Split where teams or scaling profiles genuinely diverge.
  • Reaching for eventual consistency to avoid thinking. "Eventually consistent" is not an answer to "what does the user see, and for how long?"
  • A distributed system where one machine would do. A single Postgres box handles a great deal. Say what forces you off it.

8. The sentence pattern​

Whenever you name a technology, use this shape:

"I'd use X because I need specific property. The cost is concrete downside, which is acceptable here because requirement. If condition changed, I'd switch to Y."

Concretely:

"Cassandra, because the write pattern is append-only at 200k/sec with point reads by a known key, and there are no joins on the read path. The cost is that any query I didn't design a partition key for becomes a full scan, so the access patterns have to be fixed up front. If we later needed ad-hoc analytical queries, I'd stream into ClickHouse rather than trying to make Cassandra do it."

That structure demonstrates the three things being assessed: you know the tool, you know its cost, and you know when the decision would change.