Vibe Engines
YouTube
System Design

Design a Payment System

Step 1 / 9

Learn system design by building a payment system like Stripe step by step.

The numbers to beat1charge per keysafeto retrystoredresult

The whole design, in writing

Learn system design by building a payment system like Stripe step by step. An interactive guide covering the charge flow, idempotency to prevent double-charges, a double-entry ledger, card tokenization, async processing with webhooks, reconciliation against PSP statements, and consistency at scale.

Every step of the build above, written out: the problem each piece solves, the option that was taken and the ones that were not, the numbers, and how it fails in production.

The big idea

What makes payments hard?

Moving money looks like a simple write, but it’s the opposite: it touches external banks you don’t control, must never double-charge or lose a cent, has to survive every retry and crash, and must prove afterward that the books balance to the penny.

Merchant Appcharge $X
New in this step: Merchant App.

Treat a payment as a state machine recorded in an immutable double-entry ledger, made safe against retries by idempotency, run asynchronously around slow processors, and reconciled against the bank’s own records. Correctness first, everywhere.

What the new pieces do

Merchant Appclient
Wants to charge a customer and get a clear yes/no, without ever touching raw card data or risking a double-charge on a retry.

Step 1 · The flow

Authorize, then capture

“Charge this card $50” isn’t one step. The money has to be checked and held, then actually taken — and the request leaves your system entirely to reach the card networks, which can be slow or fail.

Payment APIidempotentPayment Workerauth + capturePSP / NetworksVisa · banks
New in this step: Payment API, Payment Worker, PSP / Networks.

"Charge this card $50" — but the request leaves your system to reach slow card networks. How do you model it?

  1. A boolean call can’t express "funds held but not yet taken," can’t handle the networks failing mid-flight, and gives you nowhere to retry from. Money needs explicit states.

  2. Assuming success loses money when the network declines or fails, and gives the merchant no real answer. You must track the charge to a confirmed outcome.

  3. Splitting auth from capture lets you reserve funds now and finalize later, and explicit states (created→authorized→captured→settled) let you reason about partial failures and retries.

The Payment API hands the charge to a Payment Worker that drives it through states: authorize (hold funds) then capture (take them), talking to the PSP / card networks. Splitting auth from capture lets you reserve now and finalize later.

Why this piece earns its place

The state machine only helps if every state it can reach is one you can recover from, and the nastiest one isn’t in the diagram: you called the PSP and never heard back. A timeout is not a decline. If the worker writes failed because a socket closed, you’ve declared dead a charge the card network may well have approved — and the customer can see the hold. So the way out of an in-flight authorize has to be a read, not a guess: send your own payment id along as the PSP’s reference on the way out, and when a call times out, look the charge up by that reference and let the answer pick the state. That shapes the rest of the flow — every external call is attempt, then confirm, and the worker is allowed to be unsure. The split also quietly assumes capture happens while the hold is still alive. Holds expire, often around seven days, so keep the expiry beside the state and either capture, re-authorize, or explicitly release — an abandoned hold still sits on the customer’s available balance.

What the new pieces do

Payment APIbackend
The entry point for charges and refunds. Validates the request, enforces idempotency, and orchestrates the payment without exposing card networks.
Payment Workerservice
Drives the charge through its states — authorize, then capture — talking to the card networks and recording every transition to the ledger.
PSP / Networksservice
The external processors and card networks that actually move money. Slow, occasionally flaky, and entirely outside your control — so isolate them.

Step 2 · Never charge twice

Idempotency

Networks time out. The merchant retries the charge, unsure if the first one went through. Naively, that’s two charges for one purchase — the cardinal sin of payments.

Merchant AppPayment APIPayment WorkerPSP / Networks
The system as it stands at this step. · swipe to pan the diagram

A charge times out and the merchant retries, unsure if the first went through. How do you avoid double-charging?

  1. A fuzzy "looks like a duplicate" check is racy and wrong — two legitimate $50 charges look identical, and a slow first charge may not be visible yet. You need an explicit key, not a guess.

  2. The first request does the work and stores the result under the key; any retry with the same key returns that same result without charging again. Make repeating a request unable to repeat its effect.

  3. A timeout is ambiguous — the charge may have succeeded — so "never retry" means merchants miss real charges or hang waiting. Retrying must be safe, which is the system’s job to guarantee.

Require an idempotency key per logical charge. The first request does the work and stores the result under that key; any retry with the same key returns the same result without charging again. Safe to retry, always.

Why this piece earns its place

The key only works if writing it is the same act as doing the work. Retries arrive while the first charge is still in flight — the result isn’t stored yet, so a naive check-then-act lets both through and hands you the double-charge you just designed away. Insert the key first, in the same transaction that records the intent, with a uniqueness constraint on it, so the second request loses the race and is told the charge is in progress rather than being handed a blank. The other half is the body. A key with no memory of what it was used for will happily return a stored $50 result to a request that now says something else, so a merchant who reuses a key across two orders gets silence instead of a charge. Store a hash of the request beside the key and reject a mismatch loudly rather than replaying. And keys need a lifetime: keep them at least as long as any client will still be retrying, because the moment a key expires the request it was protecting becomes chargeable all over again.

  • 1charge per key
  • safeto retry
  • storedresult

Back of the envelope

key per logical charge
1 charge per key, no matter how many retries
first request ⇒ do + store result
replays return the stored result
ambiguous timeout ⇒ safe retry
repeating a request can’t repeat its effect

Step 3 · Track every cent

The double-entry ledger

A single “balance” field that you increment is impossible to audit and easy to corrupt under concurrency. When money goes missing, you need to know exactly where — and prove it.

Payment Workerstate machineLedgerdouble-entry
New in this step: Ledger.

How do you record balances so money is auditable and can’t be silently corrupted under concurrency?

  1. A mutable balance is impossible to audit (no history of why it changed) and easy to corrupt under concurrent updates. When money goes missing you can’t prove where it went.

  2. A side log that drifts from the authoritative balance gives you two disagreeing sources of truth. The record of movements must BE the source of balances, not a copy of it.

  3. Every transaction debits one account and credits another by equal amounts, append-only, never edited. Debits must equal credits so the books always sum to zero — a continuous built-in correctness check.

Record every movement in an immutable, double-entry Ledger: each transaction debits one account and credits another by equal amounts, and rows are never edited — only new entries appended. Balances are derived by summing entries.

Why this piece earns its place

Deriving a balance by summing is what keeps this honest, and it’s also the part that gets slower every day: a merchant’s balance is a scan of their whole history, and across millions of payments that isn’t a query you want on a live path. The knob is checkpointing — periodically fold the entries up to a point into a stored balance, then sum only the tail. Checkpoint too rarely and reads crawl; checkpoint too eagerly and you spend most of the work recomputing balances nobody asked for. The real failure is going one step further and letting the checkpoint become the answer. The moment anything writes to it directly, or trusts it after a correcting entry landed behind it, you’ve rebuilt the mutable balance field this step exists to avoid. Hold the line that a checkpoint is a cache of a fold: always reproducible from the entries, re-verified by recomputation on a schedule, and when it disagrees with the sum it’s the checkpoint that’s wrong, never the entries.

What the new pieces do

Ledgerstore
The immutable, double-entry record of every movement of money. The single source of truth for balances — append-only, never edited.

Back of the envelope

every txn: debit = credit
the books always sum to zero
append-only, never edited
a full auditable history of every cent
balance = Σ entries
derived, not a mutable field

Step 4 · Don’t hold the cards

Tokenization & PCI

Storing raw card numbers makes you a giant target and drags your entire system into the strictest PCI compliance scope. One leak is catastrophic — and most of your services never need the real number anyway.

Payment APIPayment WorkerPSP / NetworksLedgerToken Vault
New in this step: Token Vault. · swipe to pan the diagram

You need to charge cards repeatedly but storing card numbers makes you a giant breach target. What do you do?

  1. Even encrypted, raw card data in your database drags every service into the strictest PCI scope and is catastrophic if leaked. Most of your system never needs the real number.

  2. Card data goes straight into a hardened vault that returns a token; your API, workers and ledger handle only meaningless tokens. Raw PANs live solely in the vault and at the network — shrinking PCI scope and blast radius.

  3. Pushing card storage onto every merchant multiplies breach targets and PCI burdens, and still sends raw cards across your system. Centralize sensitive data in one vault, not everywhere.

Send card data straight into a PCI-compliant Token Vault, which returns a token. Everything else — your API, workers, ledger — stores and passes only tokens; raw card numbers exist solely inside the vault and at the network.

Why this piece earns its place

The vault shrinks your compliance scope only under an assumption worth saying out loud: the real card number never touches anything you run. If your checkout form posts the card to your own API and your API forwards it to the vault, that number has been through your load balancer, your web tier, your request logs and whatever traced the call — and all of it is back in the strictest scope, tokens or not. So the collection path is the design, not the storage: the card field belongs to the vault, rendered in an iframe or hosted field that posts straight to it, and what reaches your API is already a token. The day the assumption stops holding is rarely a decision anyone announces. It’s a support tool that lets an agent type a card in to fix an order, a debug endpoint that echoes the request body, a mobile client that rolls its own form because the hosted one looked wrong. Each is a small, reasonable change that silently drags the whole system back inside the scope you paid to leave.

What the new pieces do

Token Vaultstore
Stores card details in a PCI-compliant vault and returns a token. The rest of the system handles only tokens, never raw PANs — shrinking PCI scope.

Step 5 · Slow banks, fast API

Async events & webhooks

Authorization and settlement can take seconds — or arrive minutes later (a bank confirmation, a delayed failure). Blocking the merchant’s request until everything finishes is both slow and impossible.

WebhooksPayment Events
New in this step: Webhooks, Payment Events. · swipe to pan the diagram

Authorization and settlement can take seconds or arrive minutes later. How does the API respond fast?

  1. Settlement can take minutes or arrive later as a delayed confirmation — you can’t hold an HTTP request that long, and you don’t control bank latency. Blocking couples your response to the slowest external party.

  2. Polling wastes requests and still misses results that arrive much later (delayed failures, bank confirmations). Push the outcome when it happens, don’t have the merchant keep asking.

  3. The API returns the current state immediately; a durable event stream drives webhooks that notify the merchant of the final result (with retries) once it arrives. Slow external steps run off the request path.

Emit each state change to a durable Payment Events stream. The API responds quickly with the current state; Webhooks later notify the merchant of the final outcome (with retries). Slow, external steps run asynchronously off the request path.

Why this piece earns its place

Pushing the outcome means handing it to an endpoint you don’t own, and that endpoint will be down, slow, or behind. Retries make delivery at-least-once, so the merchant will see the same event twice, and because a retry of an old event interleaves with fresh ones they can see captured arrive before authorized. Put an event id and the payment’s state version in every payload, and say plainly in the docs that a receiver must ignore an event it has already handled and one older than the state it holds. Sign the payload too, or a webhook is an unauthenticated POST claiming money moved. The harder discipline is on your side: a payment isn’t captured because the webhook was delivered, and it isn’t un-captured because it wasn’t. Delivery is a separate object with its own attempts and its own dead letter, and once those attempts run out the merchant needs a way to catch up unaided — a read API over payments and events, so a day of failed deliveries is recoverable without anyone re-sending anything by hand.

What the new pieces do

Webhooksservice
Pushes final payment status (succeeded/failed/refunded) to the merchant asynchronously, with retries, since results often arrive after the initial response.
Payment Eventsbus
A durable stream of payment state transitions that drives webhooks, ledger updates, analytics and reconciliation — decoupling slow steps from the API.

Step 6 · Prove it balances

Reconciliation

Your ledger says one thing; the processor’s settlement may say another, because of fees, declines, chargebacks, or a missed event. Over millions of payments, tiny discrepancies hide real lost money.

Reconciliationmatch PSPPSP Statementssettlement files
New in this step: Reconciliation, PSP Statements.

Your ledger and the processor’s settlement can disagree (fees, declines, chargebacks, a missed event). How do you catch lost money?

  1. The live path can miss events, and fees/declines/chargebacks alter what actually settled. Over millions of payments, tiny undetected discrepancies hide real lost money. "We think it’s right" isn’t enough for money.

  2. A daily job compares your ledger line-by-line against the bank’s settlement statements — external truth — and flags every mismatch for investigation. It catches whatever the live path missed.

  3. Silently overwriting your ledger destroys the audit trail and hides the bug causing the drift. Mismatches must be flagged and investigated, corrections recorded as new entries — never edits.

A daily Reconciliation job compares your Ledger against the PSP’s settlement statements, line by line, and flags every mismatch for investigation. The bank’s record is treated as external truth your internal books must agree with.

Why this piece earns its place

A line-by-line diff of two systems that disagree by design will flag nearly everything on its first run. The statement nets fees out, settles a capture a day or more after you recorded it, converts currency at its own rate, and holds a reserve back — none of that is a missing cent, and all of it looks like one. So reconciliation is mostly a matching rule: which of your entries is expected in which statement window, and which differences are explained rather than lost. That window is the knob. Tighten it and today’s captures come up unmatched every single morning, people learn the report is noise, and the one real break scrolls past unread. Widen it and genuinely lost money sits inside the tolerance for days before anyone looks. Set it from how the processor actually settles, then treat whatever survives as work rather than output: every unmatched item gets an owner, an age, and a resolution recorded as its own ledger entries — so the exception queue is something that drains, not a page that only ever grows.

What the new pieces do

Reconciliationworker
Compares your ledger against the PSP’s settlement statements daily, flagging any mismatch — the audit that proves the books actually balance.
PSP Statementsstore
The processor’s record of what truly settled and when. Reconciliation treats this as external truth to verify your internal ledger against.

Back of the envelope

ledger vs PSP statement, daily
line-by-line external verification
flag every mismatch
fees, declines, chargebacks, missed events
external truth, then investigate
proven correct, not assumed

Step 7 · Consistency & scale

The sharp edges

A capture succeeds at the PSP but your ledger write fails — now your records and the bank disagree. And at scale, hot merchants and refunds/chargebacks add states and load that strain the system.

Merchant AppPayment APIPayment WorkerPSP / NetworksLedgerToken VaultReconciliationPSP StatementsWebhooksPayment Events
The system as it stands at this step. · swipe to pan the diagram

Make the worker a durable, retryable state machine (outbox pattern: persist intent, then act, then mark done) so it always converges even after crashes. Shard by merchant/account, and model refunds and chargebacks as first-class ledger transactions, never edits.

Why this piece earns its place

The outbox’s whole value is that it re-drives calls until the external side agrees — which means this step retries exactly the operations step 2 only protected at your own front door. Step 1’s answer to a timeout was to look the charge up by your own reference, and that only works once the processor can see it — the outbox re-drives in exactly the window where it can’t, the call that died in the network or died in your process before it ever left. The merchant’s key stopped a duplicate request becoming a duplicate charge; neither of those makes the retry itself safe. If each attempt builds a fresh request, a capture that timed out and quietly succeeded gets captured again on retry, and you’ve shipped the failure this whole page opened on. So the key you send the processor has to be derived from the payment and its attempt, persisted with the intent before the first call leaves, and reused byte for byte on every retry — including after a crash and a restart on a different host, which is precisely the moment a value held only in memory would be regenerated. Treat it as part of the durable record, not something a client library invents per call. That’s what makes the reconciler relentless rather than dangerous.

You did it

You just designed a payment system.

Merchant AppPayment APIPayment WorkerPSP / NetworksLedgerToken VaultReconciliationPSP StatementsWebhooksPayment Events
The finished design, end to end. · swipe to pan the diagram

Everything you assembled, in order

  • A payment modeled as a state machine: authorize then capture via the PSP.
  • Idempotency keys make every charge safe to retry — exactly one charge per key.
  • An immutable double-entry ledger keeps every cent auditable and balanced.
  • Tokenization confines card data to a PCI vault, shrinking the blast radius.
  • Async events + webhooks decouple the fast API from slow, external banks.
  • Daily reconciliation against PSP statements proves the books actually agree.
  • A durable, retryable worker plus refunds-as-ledger-entries handle failure and scale.

Where an interviewer pokes next

Getting the boxes right is the easy half. These are the questions that separate a candidate who drew the diagram from one who has run the thing. Answer each one out loud before you open it.

  1. How do you get atomicity across your ledger and the external PSP?

    You can’t — there’s no distributed transaction spanning your database and Visa. Use the outbox/saga pattern: persist intent durably, perform the external call, then record the result, with a reconciler that retries until both sides agree. The ledger is the immutable history; eventual consistency plus relentless reconciliation replaces the atomicity you can’t have.

  2. How are refunds and chargebacks modeled?

    As new, first-class ledger transactions — never edits to the original. A refund is a reverse entry (credit the customer, debit your revenue); a chargeback adds its own entries plus fees. Because the ledger is append-only, the full history (charge → refund → chargeback) stays intact and balances re-derive correctly by summing.

  3. Why are authorize and capture separate, and what goes wrong between them?

    Auth places a hold (capture later, e.g. when goods ship); capture takes the money. The gap enables ship-then-charge and amount adjustments, but auths expire (often ~7 days) and the hold can be lost, so you must capture in time or re-authorize. Partial and over-captures are their own states the machine handles.

  4. How do you stop a single hot merchant overwhelming the system?

    Shard by merchant/account so load and ledger writes spread across nodes, rate-limit per merchant, and run state transitions through the durable event stream so spikes queue rather than topple the API. Per-account sharding keeps a hot merchant’s contention off everyone else’s ledger.

  5. Why double-entry instead of just logging transactions?

    A plain log records what you did; double-entry enforces an invariant — every movement debits one account and credits another equally, so the system sums to zero by construction. That makes corruption detectable (the books won’t balance), lets you derive any balance from history, and is exactly what auditors and reconciliation rely on. It’s a correctness mechanism, not just a record.

Check yourself — the answers, and why

Eight steps in, these are the calls you should be able to make cold. Pick one, then read why.

  1. A payment is modeled as a state machine (authorize→capture→settled) so that…

    Explicit states, not a boolean "paid?", let you handle the networks being slow or failing mid-flight.

  2. Idempotency keys ensure that…

    Timeouts are ambiguous — the key makes repeating a request unable to repeat its effect.

  3. A double-entry ledger keeps money honest because…

    Append-only entries with debits=credits give a continuous, auditable correctness check.

  4. Tokenization (a PCI vault) exists to…

    The rest of the system handles meaningless tokens — slashing both PCI scope and breach damage.

  5. Reconciliation against PSP statements turns…

    Distributed systems drift; money can’t — the bank’s record is external truth your ledger must match.

How you’d open this design in an interview

Before any boxes: agree what it must do, pin the qualities that shape everything, then build — naming each trade-off as you make it. The walkthrough above is that exact order.

What it must do

Agree on these before drawing a single box.

  • Charge a card: authorize (hold) then capture (take) the amount, returning a clear succeeded/failed.
  • Never double-charge: an idempotency key makes a retried charge produce exactly one charge.
  • Track every cent: record each movement in an immutable double-entry ledger; balances are derived sums.
  • Handle cards safely: tokenize via a PCI vault so the system passes only tokens, never raw card numbers.
  • Notify asynchronously: webhook the merchant the final outcome — succeeded, failed, refunded — off the request path.

The qualities that shape everything

Each one names the mechanism that buys it.

Reason about partial failures
Model a payment as an explicit state machine — created → authorized → captured → settled — so a network failing mid-flight has a state to retry from, not a lost boolean.
Retries can’t repeat the effect
An idempotency key per logical charge: the first request does the work and stores the result; any retry with the same key returns that stored result without charging again.
Money is auditable, can’t silently corrupt
An immutable, append-only double-entry ledger where every debit equals a credit, so the books always sum to zero — a continuous built-in correctness check.
Shrink the breach blast radius
Tokenization confines raw card data to one PCI-compliant vault; the API, workers and ledger handle only meaningless tokens.
Fast API over slow, flaky banks
Emit state changes to a durable event stream and webhook the final outcome with retries, so the response never blocks on bank latency you don’t control.
Prove the books actually agree
A daily reconciliation job compares the ledger line-by-line against the PSP’s settlement statements and flags every mismatch.

The trade-offs you say out loud

Senior signal isn’t the boxes — it’s naming what you gave up and why it was the right price.

A state machine (authorize→capture) over one synchronous paid/not-paid call

A boolean call can’t express “funds held but not yet taken,” can’t survive the card networks failing mid-flight, and leaves nowhere to retry from. Explicit states let you reason about partial failure.

An idempotency key over a fuzzy “looks like a duplicate” check

Two legitimate identical charges look the same and a slow first charge may not be visible yet, so guessing is racy. An explicit key makes repeating a request unable to repeat its effect.

An immutable double-entry ledger over a mutable balance field

A single incremented balance has no history and corrupts under concurrency. Append-only entries with debits = credits sum to zero, so corruption shows up as books that don’t balance.

Tokenize in a PCI vault over encrypting cards in your own database

Even encrypted, raw card data drags every service into the strictest PCI scope and is catastrophic if leaked. A vault confines PANs to one hardened component and slashes the blast radius.

Async events + webhooks over blocking the API until the bank settles

Settlement can take minutes or arrive later as a delayed confirmation, and you don’t control bank latency — blocking couples your response to the slowest external party.

The answer, out loud

What a strong answer to “Design a Payment System” sounds like, first question to last trade-off. It is about 7 minutes of talking; the whiteboard and the interviewer fill the rest of the 45. Read it aloud once, then close the page and give it yourself.

  1. 0–4 min

    What actually makes this hard

    Before I draw any boxes I want to agree on what makes this hard, because it isn’t throughput. We’re moving money through banks we don’t own, on a request that can fail anywhere along the way, and there are two things we can never do: charge someone twice, and lose track of a cent. So I’m going to optimise for correctness and for being able to prove it afterwards, and I’ll happily trade latency and some moving parts to get there. I’ll assume merchants call us to charge a card, and that they need an answer clear enough to show a customer.

  2. 4–10 min

    A charge is two things, not one

    The first thing I’d push back on is treating a charge as one call that returns paid or not paid. Underneath there are two distinct operations: authorising, which checks the card and holds the money, and capturing, which actually takes it. Keeping them separate is what lets a shop hold funds now and charge when the goods ship. More importantly it gives me states — created, authorised, captured, settled, plus the failure branches — so when the card network goes quiet halfway through I have something to resume from instead of a boolean that lost its meaning. The API takes the request, and a worker drives the payment through those states and does the talking to the processors.

    Built in step 1: Authorize, then capture
  3. 10–16 min

    Retries have to be free

    Then the obvious question: the merchant’s request times out, they don’t know whether it went through, and they retry. I won’t try to detect that. Two identical charges for the same amount are a completely normal thing for a shop to take, so guessing is guaranteed to be wrong in one direction or the other. Instead I’d require an idempotency key on every charge, chosen by the merchant, one per logical purchase. The first request does the work and stores its outcome under that key; any later request carrying the same key gets that same outcome back without a card being touched. Retrying stops being a risk, which is what lets everything downstream retry too.

    Built in step 2: Idempotency
  4. 16–23 min

    The books

    For the money itself I’d keep an immutable double-entry ledger rather than a balance we update. Every movement writes two entries, a debit and an equal credit, nothing is ever edited, and a balance is the sum of the entries. That buys two things. There’s a full history, so when someone asks where a payment went I can answer rather than theorise. And because debits and credits have to net to zero, the books check themselves — corruption shows up as an imbalance instead of quietly becoming the new balance. Refunds and chargebacks go in as new entries too; I never reach back and change the original.

    Built in step 3: The double-entry ledger
  5. 23–28 min

    Keep the cards away from us

    Cards I want as far away from our systems as possible. Anything holding real card numbers falls into the strictest compliance scope and is worth a lot to whoever breaks in. So card data goes straight into a PCI vault, which hands back a token, and our API, workers and ledger only ever see that token. The real number lives in the vault and at the network, nowhere else. I’d frame that as a blast-radius decision more than a compliance one — I’d rather have one hardened thing to defend than a dozen services that all technically could be holding a card.

    Built in step 4: Tokenization & PCI
  6. 28–34 min

    Don’t make the merchant wait on a bank

    Now latency. Authorisation takes seconds and settlement can be minutes or the next day, so I’m not holding the merchant’s HTTP request open until a bank has finished. The worker emits each state change onto a durable event stream, the API answers immediately with the current state, and webhooks push the final outcome to the merchant once it exists, with retries. The same stream then feeds the ledger, analytics and reconciliation, so none of those need another synchronous hop on the charge path.

    Built in step 5: Async events & webhooks
  7. 34–39 min

    Proving it rather than believing it

    None of that proves we’re right, so the next piece is reconciliation. Every day we compare our ledger against the processor’s settlement statements and flag anything that doesn’t line up. The statement is external truth: it knows what really settled, what got declined, what was charged back and what fees came out. A mismatch is a bug report, not something to paper over — I’d never edit the ledger to agree with it. I’d record a correcting entry and go find out why we drifted in the first place.

    Built in step 6: Reconciliation
  8. 39–43 min

    The edges, and scale

    The sharpest edge is that there’s no transaction spanning our database and Visa. A capture can succeed at the processor while our own write fails. I’d handle that with an outbox: persist the intent, then act, then record what came back, and let a worker keep re-driving anything unfinished until both sides agree. For scale I’d shard by merchant account so one busy shop’s writes stay off everyone else’s path, and I’d keep every unusual outcome — a partial capture, an expired hold, a chargeback — as a state in the machine rather than a special case bolted on somewhere.

    Built in step 7: The sharp edges
  9. 43–45 min

    What I traded away

    So the trade I’m making throughout is latency and moving parts in exchange for being able to prove the books. A fully synchronous design is much easier to explain and falls apart the first time a bank goes quiet. With more time I’d work through refunds and chargebacks properly — the fee entries, the dispute timeline, who ends up carrying the loss — and how we’d onboard a second processor without the ledger ever caring which one settled a given payment.

What this teaches

Learn system design by building a payment system like Stripe step by step. An interactive guide covering the charge flow, idempotency to prevent double-charges, a double-entry ledger, card tokenization, async processing with webhooks, reconciliation against PSP statements, and consistency at scale.

Key takeaways

  • A payment modeled as a state machine: authorize then capture via the PSP.
  • Idempotency keys make every charge safe to retry — exactly one charge per key.
  • An immutable double-entry ledger keeps every cent auditable and balanced.
  • Tokenization confines card data to a PCI vault, shrinking the blast radius.
  • Async events + webhooks decouple the fast API from slow, external banks.
  • Daily reconciliation against PSP statements proves the books actually agree.
  • A durable, retryable worker plus refunds-as-ledger-entries handle failure and scale.

Concepts covered

  • What makes payments hard?
  • Authorize, then capture
  • Idempotency
  • The double-entry ledger
  • Tokenization & PCI
  • Async events & webhooks
  • Reconciliation
  • The sharp edges
built to balance, not memorized — make the calls, drop the PSP, run the gauntlet.
Finished this one? 0 / 65 System Designs done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More System Designs