Skip to main content

Ticketmaster (High-Demand Event Booking)

Sharpened prompt. Design a ticketing platform where 2M people attempt to buy 50,000 seats at 10:00:00 sharp, no seat is ever sold twice, abandoned carts return to inventory within seconds, bots do not take the whole allocation, the seat map stays interactive under 500k QPS, and a failed payment does not silently destroy a seat.

This is the purest contention problem in the playbook. The interesting question is not how to store tickets — it is how to stop 2M people from touching the same 50,000 rows in the same second.

1. Problem framing​

Functional requirements​

  • Browse events; view a live seat map with availability.
  • Hold specific seats for a bounded time while the buyer checks out.
  • Purchase held seats; release them automatically on abandonment.
  • Support general admission (quantity-based) as well as reserved seating.

Non-functional requirements​

PropertyTargetConsequence
OversellingExactly zero, everThe hold operation must be atomic and authoritative
Onsale concurrency2M concurrent users, 50k seatsAdmission control before the booking engine, not inside it
Hold expiryReleased within seconds of timeoutA reliable delayed-execution mechanism, not a nightly job
Seat map reads500k QPS, ≤ 2s staleReads and writes are entirely different systems
FairnessDefensible ordering; bot share minimisedQueue position must be assigned before the rush, not won by speed

Back-of-the-envelope​

Onsale: 2M users arrive within ~60s = ~33k arrivals/sec
Each polls/refreshes ~5× = ~165k req/sec of pure impatience
Inventory: 50,000 seats. Even at 100% conversion this is 50k successful writes TOTAL.
The write volume is trivial; the CONTENTION is the entire problem.
Seat map: 50k seats × ~40 B of state = 2 MB. Small enough to broadcast.
Holds: peak ~50k concurrent holds, each with a 10-minute TTL
Payments: ~50k authorisations over ~15 minutes = ~55/sec — easy, but each one
can fail, and each failure must return a seat.

The asymmetry to state early: 50,000 total writes against 2 million concurrent attempts. The system is not write-heavy; it is contention-heavy, and those need opposite solutions.

2. High-level architecture​

3. Component inventory​

ComponentConcrete choiceWhy this one
Waiting roomQueue-it / a custom token service at the edge (Envoy, Cloudflare)Must sit outside the application, or the application is what falls over
Lock storeRedis Cluster with Lua scriptsSingle-threaded per shard, so a script is atomic without distributed locking
Authoritative inventoryPostgres or CockroachDBRedis is the fast path; the database is the truth and the audit trail
Hold expiryRedis sorted set by expiry timestamp, swept by workersSimple, exact, and observable. SQS delay queues or RabbitMQ DLX also work
Seat mapCompact bitmap, CDN-cached with a short TTL, plus WebSocket deltas500k QPS of reads must never touch the booking core
PaymentsSaga with explicit compensationA failed charge must return the seat, always
Bot defenceDevice fingerprinting, behavioural scoring, proof-of-workFairness is a product requirement here, not a nice-to-have

4. The toughest parts​

4.1 The virtual waiting room: keeping the herd outside​

Why it's hard. If 2M users hit the booking service simultaneously, it does not matter how well the seat locking works — connection pools exhaust, the database's lock manager thrashes, latency climbs, retries multiply, and the system collapses before selling a single ticket. You cannot solve this inside the application. Autoscaling does not help either: you would need 40× capacity for 60 seconds.

Solution — admission control at the edge, with a cryptographic token and a governed drip rate.

// Edge: reject anything without a valid token BEFORE it costs you anything.
async function handler(request) {
const token = request.headers.get('x-admission');
if (!token || !(await verifyJWT(token, PUBLIC_KEY))) {
return Response.redirect('/waiting-room?event=123', 302);
}
return fetch(origin, request); // admitted: pass through
}

Three properties make this work. The token is signed and single-use, so it cannot be shared or replayed. The admission rate is dynamic, driven by the booking service's observed latency and error rate — a closed control loop rather than a fixed constant. And the queue position is assigned at arrival, which means being 50 ms faster does not help you and the waiting room becomes a fairness mechanism as well as a throttle.

Show users their position and an ETA. It costs nothing and it is the difference between a queue people tolerate and a refresh storm — which is itself a load-generating behaviour you are trying to prevent.

4.2 Holding a seat atomically​

Why it's hard. Two users click the same seat in the same millisecond. SELECT status FROM seats WHERE id=42 followed by UPDATE seats SET status='held' is the classic check-then-act race: both read available, both write held, both proceed to checkout, one of them is told at the venue that their seat does not exist. Pessimistic row locking (SELECT FOR UPDATE) is correct but serialises access to the hottest rows and, in a popular onsale, produces lock convoys and deadlocks across multi-seat selections.

Solution — hold in Redis with a single Lua script, then persist asynchronously.

-- hold_seats.lua
-- KEYS = seat keys; ARGV[1] = user_id, ARGV[2] = hold_ttl_seconds, ARGV[3] = hold_id
-- All-or-nothing: either every requested seat is held, or none is.
for i = 1, #KEYS do
if redis.call('EXISTS', KEYS[i]) == 1 then
return {0, i} -- already held: fail fast, hold nothing
end
end

for i = 1, #KEYS do
redis.call('SET', KEYS[i], ARGV[3], 'EX', tonumber(ARGV[2]))
end

-- Record the hold for expiry sweeping and for the user's cart view.
redis.call('ZADD', 'holds:expiry', tonumber(ARGV[4]), ARGV[3])
redis.call('HSET', 'hold:' .. ARGV[3], 'user', ARGV[1], 'seats', table.concat(KEYS, ','))
return {1, 0}
held, failed_idx = await redis.evalsha(
HOLD_SHA, keys=[f"seat:{eid}:{s}" for s in seats],
args=[user_id, 600, hold_id, int(time.time()) + 600])
if not held:
raise SeatUnavailable(seats[failed_idx - 1])

Why this is correct and fast: Redis executes the entire script on a single thread with no interleaving, so the check-then-act race cannot occur. Why it is atomic across multiple seats: the script checks all seats before setting any, so a party of four never ends up with two seats and a failure.

The multi-seat requirement forces all the event's seat keys onto one Redis slot, done with a hash tag: seat:{event:123}:A14. Everything in braces determines the slot, so the whole event lives on one shard — which is fine, because 50,000 keys and 50,000 writes is nothing for a single Redis shard, and it buys you script atomicity.

Redis is the fast path, not the truth. Write the hold to Postgres asynchronously and the purchase synchronously; if Redis is lost entirely, holds evaporate (annoying, recoverable) but sold seats do not (unacceptable, and protected by the database). Say that split explicitly — it is the sentence that shows you know where correctness lives.

4.3 Releasing abandoned holds within seconds​

Why it's hard. A user holds four seats and closes their laptop. Those seats are invisible to everyone else until the hold expires. If expiry is handled by a nightly job, inventory is destroyed. If it is handled by a per-hold timer thread, 50,000 timers is a resource problem and a crashed process loses them all. And expiry must be exactly once — releasing a seat that was actually purchased in the interim would oversell.

Solution — a sorted set keyed by expiry timestamp, swept by idempotent workers.

# Sweeper: runs every second on a few workers. Bounded work, crash-safe.
async def sweep():
now = time.time()
# Atomically claim a batch so two sweepers never process the same hold.
expired = await redis.eval("""
local ids = redis.call('ZRANGEBYSCORE', KEYS[1], 0, ARGV[1], 'LIMIT', 0, 500)
if #ids > 0 then redis.call('ZREM', KEYS[1], unpack(ids)) end
return ids
""", keys=["holds:expiry"], args=[now])

for hold_id in expired:
# Conditional release: only if the hold is still a hold. If checkout
# converted it to a purchase, the state check makes this a no-op.
await release_if_unpurchased(hold_id)

async def release_if_unpurchased(hold_id):
async with pg.transaction():
row = await pg.fetchrow(
"SELECT status, seats FROM holds WHERE hold_id=$1 FOR UPDATE", hold_id)
if row is None or row["status"] != "held":
return # purchased or already released
await pg.execute("UPDATE holds SET status='expired' WHERE hold_id=$1", hold_id)
await redis.delete(*[f"seat:{{event:{eid}}}:{s}" for s in row["seats"]])
await events.publish("seats.released", row["seats"]) # push to seat maps

Note the defence in depth: Redis TTLs already expire the seat keys on their own, and the sweeper exists to do the bookkeeping (mark the hold expired, notify watchers, restore the database view). Even if every sweeper dies, seats become available again through TTL alone. Two independent mechanisms for the same guarantee, with the cheap one as the backstop.

The conditional check inside a transaction is what makes it exactly-once-in-effect: a race between "hold expires" and "user completes purchase" resolves deterministically to whichever reaches the row lock first, and the loser becomes a no-op.

4.4 Bots, scalpers, and what "fair" means​

Why it's hard. A meaningful share of onsale traffic is automated. Bots are faster than humans at every step: they hold seats in milliseconds, solve simple challenges, and rotate through thousands of residential IPs and accounts. If you optimise purely for throughput, you optimise for bots. And unlike most abuse problems, the damage is highly visible — it becomes a news story and, in several jurisdictions, a regulatory matter.

Solution — layered friction that costs bots far more than humans, plus fairness mechanisms that make speed irrelevant.

LayerMechanismWhy it costs bots more
IdentityVerified phone/payment method, account age, purchase historyCreating thousands of verified identities is expensive
DeviceFingerprinting (canvas, fonts, TLS/JA4), headless-browser detectionBot farms reuse environments; entropy collapses
BehaviourMouse movement, dwell time, scroll patterns before clickingHumans are irregular; scripts are not
EconomicClient-side proof-of-work before admissionA few CPU-seconds is invisible to one human, ruinous at 10,000 parallel sessions
FairnessRandomised queue position among everyone present at onsaleSpeed stops mattering entirely
LimitsPer-identity and per-payment-instrument ticket caps, enforced atomicallyCaps the payoff even on success

The randomised queue is the strongest single measure and it is a design decision, not a detection problem: everyone who arrives during a pre-onsale registration window gets a random position. Being 10 ms faster gains nothing, so the entire incentive to build a fast bot evaporates. Registration also gives you a verified-identity funnel before the rush, when you have time to evaluate it.

Enforce ticket limits atomically inside the hold script, not as a separate check — otherwise the limit itself has a check-then-act race, which is exactly how bots exceed it:

local bought = tonumber(redis.call('HGET', 'limits:' .. ARGV[1], ARGV[5]) or '0')
if bought + #KEYS > tonumber(ARGV[6]) then return {0, -1} end -- over the cap
redis.call('HINCRBY', 'limits:' .. ARGV[1], ARGV[5], #KEYS) -- same atomic script

4.5 Serving the seat map to 500,000 people at once​

Why it's hard. Every waiting user stares at a seat map, and every one of them wants it live. Querying seat state per request means 500k QPS against the same Redis keys the booking core depends on — the read path starves the write path, and the writes are the ones that make money.

Solution — treat availability as a small, cacheable, broadcastable artifact rather than a query.

# The entire availability state for a 50k-seat venue is a 50k-bit bitmap: 6 KB.
# Regenerate it once per second and serve it to everyone from cache.
async def publish_seatmap(event_id):
bits = bitarray(venue_seat_count(event_id))
for i, seat in enumerate(all_seats(event_id)):
bits[i] = await redis.exists(f"seat:{{event:{event_id}}}:{seat}")
blob = gzip.compress(bits.tobytes()) # ~2 KB compressed
version = int(time.time())
await cdn.put(f"/maps/{event_id}/{version}.bin", blob, ttl=2)
await ws.broadcast(event_id, {"v": version}) # tell clients a new version exists

Three properties: the payload is tiny (a bitmap, not a JSON array of seat objects), it is identical for everyone so the CDN can serve it (unlike the per-user result sets in post search), and clients get a WebSocket nudge rather than polling.

For fine-grained updates, push deltas over the WebSocket — "seats A14, A15 now held" — so a client updates a handful of pixels rather than re-fetching. And accept that the map is up to two seconds stale: users will occasionally click a seat that was just taken, and the correct response is a fast, clear "that seat just went — here are three nearby" rather than a design that pretends to be real-time and cannot be.

4.6 When the payment fails after the seat is held​

Why it's hard. The user holds four seats and submits payment. The payment provider takes 8 seconds and returns a timeout — you do not know whether the charge succeeded. Release the seats and you may have charged someone for nothing. Keep them held and you may hold seats forever for a payment that never happened. Meanwhile the hold TTL is ticking.

Solution — a saga with explicit states, an idempotency key, and a reconciliation path for the unknown case.

async def checkout(hold_id, payment_method):
# 1. Extend the hold to cover payment processing. Never let it expire mid-charge.
await extend_hold(hold_id, seconds=180)

# 2. Create the order in 'pending_payment' — a durable record BEFORE charging,
# so a crash leaves evidence that a charge may exist.
order = await orders.create(hold_id, status="pending_payment",
idempotency_key=f"order:{hold_id}")
try:
result = await psp.charge(amount=order.total, method=payment_method,
idempotency_key=order.idempotency_key, timeout=15)
except (Timeout, ConnectionError):
# 3. UNKNOWN state — the worst case. Do NOT release, do NOT confirm.
await orders.mark(order.id, "payment_unknown")
await recon.enqueue(order.id, check_after=30) # poll the PSP for truth
return CheckoutPending(order.id)

if result.succeeded:
await confirm_purchase(hold_id, order.id) # seats -> sold, permanently
return CheckoutComplete(order.id)

await orders.mark(order.id, "payment_failed")
await release_hold(hold_id) # compensating action
return CheckoutFailed(result.decline_reason)

The payment_unknown branch is the one that separates a real answer from a textbook one. Never guess on a timeout. Poll the provider with the same idempotency key until you learn the truth, keep the seats held while you do, and if the charge did land, complete the order retroactively. The idempotency key is what makes a retry safe — the same key returns the original charge rather than creating a second one, the same discipline as the payment system.

Bound the unknown window: after a few minutes with no resolution, release the seats and record a claim for support. Holding a seat indefinitely for an unresolvable payment is worse than a rare manual refund.

4.7 General admission: the other inventory model​

Why it's hard. Reserved seating is a set of unique locks. General admission is a counter — "10,000 tickets, sell up to 10,000" — and counters have exactly the contention problem that seat locks avoid. UPDATE inventory SET remaining = remaining - 2 WHERE remaining >= 2 on one row at 33k requests/sec is a lock convoy on a single database row.

Solution — partition the counter, and fall back to a shared remainder as it drains.

-- gate_decrement.lua: shard the pool into N buckets to spread contention,
-- and only consult the shared remainder when a bucket is exhausted.
local bucket = KEYS[1] -- e.g. "ga:{event:123}:b7"
local shared = KEYS[2] -- "ga:{event:123}:shared"
local want = tonumber(ARGV[1])

local avail = tonumber(redis.call('GET', bucket) or '0')
if avail >= want then
redis.call('DECRBY', bucket, want)
return 1
end

-- Bucket empty: try the shared remainder (contended, but only late in the sale).
local rem = tonumber(redis.call('GET', shared) or '0')
if rem >= want then
redis.call('DECRBY', shared, want)
return 1
end
return 0

Split 10,000 tickets into 32 buckets of ~300 plus a shared remainder, and assign users to buckets by hash. Contention drops 32× for the large majority of the sale, and only the final tickets contend on the shared counter — which is fine, because by then the request rate has collapsed.

The trade-off to name: bucketing can leave a bucket empty while others have stock, so a user is told "sold out" while tickets remain. The shared remainder plus a rebalancing sweeper (periodically move stock from fat buckets to the shared pool) keeps that window small. Being upfront about this imperfection, and bounding it, is better than claiming a partitioned counter is exact.

5. What breaks first​

EventFirst failureMitigation
Onsale momentEverything, if there is no waiting roomEdge admission control with a governed drip rate
Seat map pollingRead traffic starves the booking coreBitmap on the CDN + WebSocket deltas; reads never touch booking Redis
Redis shard for the hot eventSingle-shard CPU on Lua evals50k ops is well within one shard; hash tags keep it deliberate, not accidental
Payment provider degradationHolds pile up in payment_unknownExtended holds, reconciliation polling, bounded unknown window
Sweeper outageExpired holds not released promptlyRedis TTL releases seats independently; sweeper only does bookkeeping
Bot swarmReal fans see "sold out" in 12 secondsRandomised queue position, verified registration, atomic per-identity caps
Postgres failoverPurchases pauseHolds continue in Redis; purchases queue and retry; never confirm without a durable write

6. Cheat sheet​

  • The asymmetry: 2M concurrent attempts, 50k total writes. Contention, not throughput.
  • Waiting room at the edge, signed single-use admission tokens, drip rate governed by booking-service health.
  • Holds: one Redis Lua script, all-or-nothing across a multi-seat request, hash-tagged to one slot per event.
  • Expiry: sorted set by timestamp swept by idempotent workers, with Redis TTL as an independent backstop.
  • Fairness: randomised queue position from a pre-registration window — it removes the incentive to build a bot.
  • Reads: availability as a 6 KB bitmap on the CDN plus WebSocket deltas. Two seconds stale, and say so.
  • Payment: saga with payment_unknown as a first-class state; never guess on a timeout; idempotency key end to end.
  • The one-liner: "Keep the crowd outside the building with a waiting room, make the hold a single atomic script so overselling is structurally impossible, and treat the seat map as a cached artifact rather than a query."