Skip to main content

Distributed Payment System (Stripe)

Sharpened prompt. Design a payment platform processing 10,000 transactions/sec across dozens of currencies and payment rails, where a network timeout never double-charges a customer, every cent is accounted for in an auditable ledger, a multi-step money movement either completes or fully unwinds, and the internal books reconcile against the bank's daily settlement file to the penny.

Everything on this page follows from one property that distinguishes payments from every other system in this playbook: you cannot undo a side effect that happened in someone else's system. Once the card network has moved money, no amount of clever retry logic takes it back.

1. Problem framing​

Functional requirements​

  • Charge a card: authorise, capture, refund, void.
  • Multi-party flows: platform fees, marketplace splits, payouts to sellers.
  • Multiple rails: card networks, ACH, SEPA, wallets — with different timing and failure semantics.
  • A queryable, immutable ledger for every movement.

Non-functional requirements​

PropertyTargetConsequence
Double-charge rateZeroIdempotency keys enforced at the storage layer, end to end
Ledger correctnessDebits equal credits, alwaysDouble-entry with a database-enforced invariant
Availability99.99% for authorisationMulti-region, but never at the cost of correctness
AuditabilityEvery state reconstructable years laterAppend-only events, no updates, no deletes
ReconciliationDaily, to the centAutomated matching with an exception queue

Back-of-the-envelope​

Volume: 10k TPS peak, ~3k TPS average = 260M transactions/day
Ledger: each transaction writes 4-8 ledger entries (double-entry, multi-party)
260M × 6 × 200 B = ~312 GB/day = 114 TB/year. Append-only, partitioned.
Idempotency: 260M keys/day × 24h retention (minimum) × 300 B = ~78 GB hot
Settlement: daily bank files, millions of rows, matched against internal entries
Latency: authorisation is dominated by the card network (200-2000ms). Your own
processing must be a small fraction of that — target under 50ms.

The latency point reframes the problem usefully: you are not optimising throughput, you are optimising correctness inside someone else's latency budget.

2. High-level architecture​

3. Component inventory​

ComponentConcrete choiceWhy this one
Idempotency storePostgres with a UNIQUE constraint on the keyThe database enforces it; application-level checks race
OrchestrationTemporal (or Step Functions)Durable execution: the workflow's state survives process death, which is what a saga needs
LedgerPostgres/Spanner append-only, or TigerBeetleTigerBeetle is purpose-built for double-entry at high throughput with enforced balance invariants
Token vaultIsolated service, HSM-backed, separate deploymentKeeps PCI scope to one small system instead of your whole estate
ReconciliationSpark batch over ledger + bank filesMillions of rows, fuzzy matching, runs daily
Event busKafka with a transactional outboxLedger write and event publish must not diverge

4. The toughest parts​

4.1 Idempotency when you don't know if the charge happened​

Why it's hard. The client sends a charge request. The network drops the response. The client cannot distinguish "the request never arrived" from "it succeeded and the acknowledgement was lost." Retrying risks a double charge; not retrying risks losing a legitimate payment. This ambiguity is unavoidable — it is the two-generals problem with money attached — and the only fix is to make the retry safe.

Solution — a client-supplied idempotency key, claimed by a unique constraint, storing the response.

CREATE TABLE idempotency_keys (
key TEXT PRIMARY KEY, -- client-supplied, e.g. a UUID
request_hash BYTEA NOT NULL, -- sha256 of the canonical request body
state TEXT NOT NULL, -- 'in_progress' | 'completed'
response_code INT,
response_body JSONB,
locked_at TIMESTAMPTZ,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
async def charge(key: str, req: ChargeRequest):
h = sha256(canonical(req))
try:
# Claim the key. The unique constraint is the mutual exclusion; there is
# no window between "check" and "act" because it is one statement.
await pg.execute(
"INSERT INTO idempotency_keys (key, request_hash, state, locked_at) "
"VALUES ($1, $2, 'in_progress', now())", key, h)
except UniqueViolation:
row = await pg.fetchrow("SELECT * FROM idempotency_keys WHERE key=$1", key)

# Same key, DIFFERENT body: a client bug. Reject loudly rather than
# silently returning the wrong charge's result.
if row["request_hash"] != h:
raise IdempotencyKeyReuse(key)

if row["state"] == "completed":
return Response(row["response_code"], row["response_body"]) # replay

# Still in progress: another request holds it. Tell the client to retry,
# rather than starting a second charge.
raise ConcurrentRequest(retry_after=1)

# We own the key. Do the work, then record the response in the SAME transaction
# that records the charge, so the two can never disagree.
result = await execute_charge(req, idem=key)
async with pg.transaction():
await persist_charge(result)
await pg.execute("UPDATE idempotency_keys SET state='completed', "
"response_code=$2, response_body=$3 WHERE key=$1",
key, result.code, result.body)
return result

Four details that make this correct rather than approximately correct:

  • The unique constraint does the work. SELECT then INSERT has a race; a single INSERT does not.
  • Hash the request body. Reusing a key with different parameters is a client bug that must fail loudly — silently returning the original response would charge the wrong amount and be nearly impossible to debug.
  • in_progress is a real state. Two concurrent requests with the same key must not both proceed; the second gets a retryable error.
  • Propagate the key downstream. Pass a deterministic reference to the card network too, so even if your guard fails, the network rejects the replay. Defence in depth across a trust boundary.

Retention matters: keep keys for at least 24 hours (Stripe's window), ideally longer for high-value flows. Expiring too early reopens the double-charge window for a client retrying after a long outage.

4.2 The ledger: double-entry, integers, and no deletes​

Why it's hard. A single balance column updated in place is the most common and most damaging design error in payments. It has no history, so you cannot explain how a balance got where it is; it is a contention hotspot; concurrent updates lose writes; and there is no structural invariant preventing money from being created or destroyed by a bug. Add floating-point arithmetic and you get rounding errors that accumulate across millions of transactions into real, unexplainable discrepancies.

Solution — an append-only double-entry ledger in integer minor units, with a balance invariant the database enforces.

-- Immutable. No UPDATE, no DELETE. Corrections are new, offsetting entries.
CREATE TABLE ledger_entries (
entry_id BIGSERIAL PRIMARY KEY,
transfer_id UUID NOT NULL, -- groups the legs of one movement
account_id BIGINT NOT NULL,
direction SMALLINT NOT NULL, -- +1 debit, -1 credit
amount BIGINT NOT NULL CHECK (amount > 0), -- MINOR UNITS. Never float.
currency CHAR(3) NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
metadata JSONB
);

-- The invariant: every transfer's legs must sum to zero, per currency.
CREATE OR REPLACE FUNCTION assert_balanced() RETURNS TRIGGER AS $$
BEGIN
IF (SELECT SUM(direction * amount) FROM ledger_entries
WHERE transfer_id = NEW.transfer_id) <> 0 THEN
RAISE EXCEPTION 'unbalanced transfer %', NEW.transfer_id;
END IF;
RETURN NULL;
END; $$ LANGUAGE plpgsql;

A single card payment with a platform fee is not one row but a set of balanced legs:

transfer_id = t_9f3a ("customer pays $100, platform keeps $3")

account dir amount currency
customer_receivable +1 10000 USD (debit: they owe us)
merchant_payable -1 9700 USD (credit: we owe them)
platform_revenue -1 300 USD (credit: our fee)
-----
sum(direction × amount) = 0 ✓

Three rules to state as non-negotiable. Integer minor units — 10000 cents, never 100.00 — because binary floating point cannot represent 0.10 exactly and the error compounds. Append-only — a refund is a new balanced transfer, not a mutation of the original; the original charge must remain visible forever for audit. Balance is a projection — SUM(direction × amount) over an account, materialised incrementally for speed but always recomputable from the entries, which is what makes the system auditable.

For extreme throughput, TigerBeetle is purpose-built for exactly this shape: double-entry, integer amounts, enforced invariants, and hundreds of thousands of transfers per second. Naming it — and explaining that it exists because general-purpose databases make this workload awkward — is a strong signal.

4.3 Sagas: money that moves through many hands​

Why it's hard. A marketplace payment is several steps across several systems: authorise the card, capture, credit the seller's balance, deduct the platform fee, schedule a payout. There is no distributed transaction spanning Visa, your ledger, and a bank. Any step can fail, and steps that already succeeded cannot be rolled back by a database — they must be compensated by a new, opposite action.

Solution — an orchestrated saga with an explicit compensation for every step, running on a durable execution engine.

@workflow.defn
class MarketplacePayment:
@workflow.run
async def run(self, req: PaymentRequest) -> Result:
completed = []
try:
auth = await workflow.execute_activity(
authorize_card, req,
start_to_close_timeout=timedelta(seconds=30),
retry_policy=RetryPolicy(maximum_attempts=3))
completed.append(("auth", auth))

cap = await workflow.execute_activity(capture, auth.id, ...)
completed.append(("capture", cap))

# Ledger legs are one atomic local transaction — no saga needed inside.
led = await workflow.execute_activity(post_transfer, req.legs, ...)
completed.append(("ledger", led))

# Payouts are scheduled, not immediate: they settle on a bank timetable.
await workflow.execute_activity(schedule_payout, req.seller_id, ...)
return Result.success(cap.id)

except Exception as e:
# Compensate in REVERSE order. Each compensation is itself idempotent
# and retried until it succeeds — a failed compensation is an incident.
for step, data in reversed(completed):
await workflow.execute_activity(COMPENSATE[step], data,
retry_policy=RetryPolicy(maximum_attempts=100))
raise

COMPENSATE = {
"auth": void_authorization, # release the hold
"capture": refund_capture, # a NEW movement, not an undo
"ledger": post_reversing_entries, # offsetting legs, original preserved
}

Four things worth saying about this:

  • Compensation is not rollback. Refunding a capture is a new transaction with its own fee and its own ledger entries. The customer's statement shows both. Sagas trade atomicity for eventual consistency with a visible history, and that history is a feature in finance.
  • Compensations must be idempotent and must eventually succeed. A failed compensation leaves money in limbo, which is why it retries ~100 times and then pages a human. This is the one place where "give up" is not an option.
  • Durable execution is what makes it practical. Temporal persists the workflow's state after every step, so a worker crash resumes mid-saga with all local state intact. Hand-rolling this means building a state machine, a scheduler, and a recovery mechanism — badly.
  • Not everything needs a saga. The ledger legs are a single local transaction; keep the saga boundary at the external systems and use ordinary ACID inside your own.

4.4 Reconciliation: proving the books are right​

Why it's hard. Your ledger says you moved $4,281,905.11 today. The bank's settlement file says $4,281,847.63. The $57.48 gap could be a fee you did not model, a transaction that settled a day later, a partially-refunded charge, a currency conversion difference, or a genuine bug that is silently losing money. Finding out which is a matching problem over millions of rows with no shared identifier — bank files often strip your reference and carry their own.

Solution — a tiered matching pipeline that resolves the easy cases automatically and escalates the rest.

def reconcile(day):
internal = ledger.settled_on(day) # our view
external = parse_bank_file(day) # NACHA / card network report

matched, unmatched_int, unmatched_ext = [], list(internal), list(external)

# Tier 1: exact match on reference + amount. Resolves the large majority.
matched += exact_match(unmatched_int, unmatched_ext,
key=lambda r: (r.reference, r.amount_minor, r.currency))

# Tier 2: amount + date within tolerance (settlement timing differs by rail).
matched += fuzzy_match(unmatched_int, unmatched_ext,
amount_tol=0, date_tol=timedelta(days=2))

# Tier 3: many-to-one. Banks batch: one deposit covers hundreds of charges.
matched += subset_sum_match(unmatched_int, unmatched_ext, max_group=500)

# Everything left is an exception. Classify it so humans see patterns, not rows.
for item in unmatched_int + unmatched_ext:
exceptions.enqueue(classify(item)) # missing_fee | timing | duplicate | unknown
return Report(matched, exceptions.count(), delta=sum_delta(internal, external))

Three points that show operational maturity. Classify exceptions rather than listing them — "412 rows unmatched" is unactionable; "412 rows, all Brazilian card transactions, all missing a 0.35% network fee we do not model" is a one-line fix. Set a materiality threshold and alarm on trend, not absolute — a small daily delta is normal, a growing one is a bug. And reconcile the ledger against itself too: a daily job asserting that every account's materialised balance equals the sum of its entries catches projection bugs before they reach a customer statement.

4.5 PCI scope: never touching the card number​

Why it's hard. Card data (the PAN) drags every system that touches it into PCI DSS scope — encryption, key management, network segmentation, quarterly scans, annual audits. If the PAN flows through your API servers, your logs, your queues, and your database, then all of them are in scope, and compliance cost scales with the number of systems, not the amount of data.

Solution — collect card data in an isolated boundary and let everything else handle tokens only.

Browser ──── card fields rendered in an IFRAME served by the VAULT domain ────► Vault
│ (your JavaScript can never read them — this is the whole point) │
│ tokenise │
│◄─────────────── token: tok_1J2k3l (no PAN, useless if stolen) ────────────────┤
│
└─► Your API (token only) ─► Ledger, orchestration, analytics: ALL OUT OF SCOPE
│
└─► Vault detokenises at the moment of
transmission to the card network only.

The critical property is that the PAN never enters your application's memory, logs, or storage. An iframe (or Stripe Elements / Apple Pay style flow) means the card fields belong to a different origin, so your own JavaScript cannot read them even accidentally. That single decision reduces PCI scope from "everything" to "the vault," which is the difference between a routine SAQ-A questionnaire and a full Level 1 audit of your entire estate.

Add network tokens where available: the card networks issue merchant-specific tokens that survive card reissuance, which both improves authorisation rates and means a breach of your token store yields nothing usable elsewhere.

4.6 Multi-currency without losing money in the cracks​

Why it's hard. A customer pays in EUR, the seller is paid in BRL, the platform's fee is in USD, and the FX rate moves between authorisation and settlement. Naively converting at each step and rounding to two decimals silently creates or destroys fractions of a cent on every transaction, and across millions of transactions those fractions become a real, unexplainable imbalance.

Solution — never mix currencies in one balanced transfer; model conversion as an explicit two-legged exchange through an FX position account.

WRONG — mixes currencies inside one transfer; the invariant cannot even be checked:
debit customer 10000 EUR
credit merchant 10800 USD -- sums to zero in no currency

RIGHT — two balanced transfers plus an explicit FX account that holds the position:
transfer 1 (EUR): debit customer 10000 EUR / credit fx_position 10000 EUR
transfer 2 (USD): debit fx_position 10800 USD / credit merchant 10800 USD

The fx_position account now holds +10000 EUR and -10800 USD. Its balance IS the
platform's currency exposure, visible and hedgeable, rather than hidden in rounding.

Three rules to state. Each transfer balances within a single currency — the invariant from 4.2 is per-currency, and a cross-currency transfer is two transfers, not one. Store the rate and its timestamp with the transfer, so the conversion is reproducible during a dispute years later. And round once, deterministically, with a documented rule (banker's rounding, or always-toward-the-platform), booking the residual to a dedicated rounding account so it is visible rather than lost.

The framing that lands: "FX isn't a conversion function, it's a position. If you model it as arithmetic you lose cents; if you model it as an account you get a hedgeable exposure number for free."

4.7 Failure modes that are unique to money​

Why it's hard. Ordinary distributed-systems failures (timeouts, partitions, retries) have consequences here that no amount of eventual consistency repairs: a customer charged twice, a seller paid twice, a refund issued for a charge that never happened.

Solution — enumerate the specific failure modes and give each a mechanism.

FailureNaive outcomeMechanism
Timeout on authorisationRetry → double chargeIdempotency key propagated to the network; the network rejects the replay
Crash after charging, before ledger writeMoney moved, books do not show itTransactional outbox: ledger write and event publish in one local transaction, delivered after
Duplicate webhook from the providerDouble-credit the sellerWebhooks are idempotent by event ID; store processed IDs
Compensation failsMoney strandedRetry ~100×, then page. Compensations may never be abandoned
Clock skew across regionsOut-of-order settlementLedger ordering by a monotonic sequence, never wall clock
Region failover mid-sagaWorkflow lostDurable execution replicated cross-region; the workflow resumes, not restarts
Partial refund arithmeticRefunds exceed the originalEnforce SUM(refunds) <= original as a ledger-level constraint, not application logic

The transactional outbox is worth spelling out because it is the standard fix for the second row and it comes up constantly:

BEGIN;
INSERT INTO ledger_entries (...) VALUES (...); -- the money movement
INSERT INTO outbox (topic, payload) VALUES ('payment.captured', '{...}');
COMMIT;
-- A separate relay polls `outbox` and publishes to Kafka, marking rows sent.
-- The ledger write and the event can never disagree, because they are one commit.

Without it you have a dual-write problem: write to the database and publish to Kafka as two operations, and a crash between them leaves the two systems permanently inconsistent — with money as the thing they disagree about.

5. What breaks first​

EventFirst failureMitigation
Card network latency spikeSaga activities time out, holds accumulateLong activity timeouts, payment_unknown state, provider polling
Merchant retry storm after an outageDuplicate charge attemptsIdempotency keys make them free; rate-limit per merchant
Ledger write contention on a hot accountPlatform revenue account becomes a bottleneckShard hot accounts into sub-accounts summed on read; or use TigerBeetle
Bank file arrives late or malformedReconciliation gap grows silentlyAlarm on file absence, not just on mismatch; hold payouts until reconciled
FX rate provider outageCross-currency payments blockedCached rate with a widened spread and an explicit staleness cap
Region loss mid-sagaIn-flight payments in unknown stateCross-region durable execution; recovery workflow that queries providers for truth

6. Cheat sheet​

  • Idempotency: client key + UNIQUE constraint + request-body hash + stored response + in_progress state, propagated to the network.
  • Ledger: append-only, double-entry, integer minor units, balance enforced per transfer per currency, balances are projections.
  • Sagas: orchestrated on durable execution, compensation for every step in reverse, compensations retry until they succeed.
  • Reconciliation: tiered matching (exact → fuzzy → subset-sum), classify exceptions, alarm on the trend.
  • PCI: iframe/vault boundary so the PAN never touches your systems; scope shrinks from everything to one service.
  • FX: two single-currency transfers through an explicit position account; store the rate; book rounding residue visibly.
  • Dual writes: transactional outbox, always.
  • The one-liner: "You cannot roll back the card network, so every design choice — idempotency keys, immutable double-entry, compensating sagas, daily reconciliation — exists to make an irreversible external side effect safe to retry and possible to prove."