</>Oriol Martí
← Blog

Designing a payment system that never charges twice

· 6 min read

In most systems, a bug is fixed with a deploy. Not in payments. If you charge a customer twice, no deploy gives them their money back. You need a support ticket, a refund and a chunk of their trust.

So I design payment systems around an uncomfortable premise: the network will fail at the worst possible moment. The request will reach the provider and the response will get lost on the way back. The user will hit “Pay” twice. The webhook will arrive before the synchronous response, or three times, or never.

This post is the diagram of a checkout flow that survives all of that. It doesn’t describe any specific system: these are the principles I’d apply today to any platform that charges through one or more providers.

The diagram

  Client ──POST /payments──▶ ┌──────────────────────┐
  (Idempotency-Key)          │ Payments API         │
                             │  └─ key registry     │
                             │     (DB)             │
                             └──────────┬───────────┘
                                        ▼
                             ┌──────────────────────┐      ┌────────────┐
                             │ Orchestrator         │─────▶│ Provider A │
                             │  └─ payment state    │      └────────────┘
                             │     machine          │      ┌────────────┐
                             │                      │─────▶│ Provider B │
                             └──────────▲───────────┘      └─────┬──────┘
                                        │                        │
                             ┌──────────┴───────────┐  webhooks  │
                             │ Webhook receiver     │◀───────────┘
                             │  └─ deduplication    │
                             └──────────────────────┘

  Nightly: Reconciliation ── compares ──▶ our DB ⇄ provider reports

Four pieces, each there to absorb a different failure:

Piece Failure it absorbs
Idempotency key The client retries the same operation.
State machine Responses that arrive out of order or contradict each other.
Webhook receiver Duplicate, late or missing notifications.
Reconciliation Whatever slips past the other three.

Idempotency: same request, same result

An idempotent operation has the same effect whether you run it once or ten times. A charge isn’t idempotent by nature, so you have to make it so.

The standard technique is the idempotency key. The client generates a unique identifier per payment attempt (not per HTTP request) and sends it with every retry. The server stores the key together with the result. If it sees the same key again, it returns the stored result instead of charging again.

async function createPayment(key: string, input: PaymentInput) {
  // The UNIQUE constraint on the key is what prevents the race,
  // not the if: two concurrent requests can't both insert it.
  const claimed = await db.idempotency.insertIfAbsent({
    key,
    requestHash: hash(input),
    status: 'in_progress',
  });

  if (!claimed) {
    const existing = await db.idempotency.get(key);
    if (existing.requestHash !== hash(input)) throw new Conflict('Same key, different request');
    if (existing.status === 'in_progress') throw new Conflict('Payment in progress, retry later');
    return existing.response; // already processed: same result, no second charge
  }

  const response = await orchestrator.charge(input, { idempotencyKey: key });
  await db.idempotency.complete(key, response);
  return response;
}

Three details in that code are easy to miss.

The database enforces exclusivity. “Does the key exist?” followed by “insert it” is two operations, and another request fits in between. A UNIQUE constraint closes that gap; application code doesn’t.

The key is bound to the payload. The same key with a different amount is a client bug. Reject it instead of returning the previous result.

The key travels to the provider too. Most payment providers accept their own idempotency key. If the orchestrator retries after a timeout, the provider recognises the retry and doesn’t charge twice. Without this, idempotency stops at your boundary and the risk is still there one hop further.

The state machine

A payment isn’t just “done” or “not done”. It moves through states, and what matters most is defining which transitions are allowed and which aren’t.

                 ┌──────────┐
                 │ created  │
                 └────┬─────┘
                      ▼
                 ┌──────────┐   timeout / no response     ┌──────────┐
                 │ pending  │────────────────────────────▶│ unknown  │
                 └─┬──────┬─┘                             └────┬─────┘
       authorised  │      │  declined               ask the    │
                   ▼      ▼                         provider   │
          ┌────────────┐ ┌──────────┐                          │
          │ authorized │ │  failed  │◀─────────────────────────┤
          └─────┬──────┘ └──────────┘                          │
                ▼                                              │
          ┌────────────┐                                       │
          │  captured  │◀──────────────────────────────────────┘
          └─────┬──────┘
                ▼
          ┌────────────┐
          │  refunded  │
          └────────────┘

The key state is unknown. It’s the honest one: if the provider didn’t respond, you don’t know whether it charged. Marking the payment failed invites the user to retry, and that’s your double charge. Marking it captured might ship something nobody paid for. unknown forces you to resolve it by asking the provider before deciding.

The other rule is that transitions only move forward. If a payment is already captured and a stale pending event shows up, it’s ignored. You enforce that with a conditional update (UPDATE ... WHERE status = 'pending'), not a separate read and write. If the update touches zero rows, another process got there first, and that’s not an error.

Webhooks: late, duplicated or missing

The provider confirms the final outcome asynchronously, via webhook. Design the receiver assuming webhooks:

  • Arrive duplicated. The provider retries until it gets a 2xx. Every event carries an ID; storing it under a unique constraint turns the second processing into a no-op.
  • Arrive out of order. A refunded can land before its captured. A forward-only state machine absorbs it.
  • Never arrive. A deploy, a network blip, an expired certificate. That’s why they can’t be the only source of truth: payments stuck in pending or unknown for too long get actively polled from the provider.
  • Can be forged. Verify the signature before processing anything.

And one operational rule: the receiver answers fast and processes later. Verify the signature, store the event, return 200 and process it from a queue. If processing takes longer than the provider’s timeout, the provider retries and you create the duplicates yourself.

Reconciliation: the safety net

With all of the above, the vast majority of cases are covered. Reconciliation exists for the rest.

Every day, payments in our database are compared against the provider’s reports. Mismatches fall into three groups:

Mismatch What it means
Charged at the provider, not in our DB A lost webhook or an unresolved unknown. Fulfil or refund.
Charged in our DB, not at the provider A bug on our side. Serious: we’re treating something as paid that isn’t.
Different amounts Currencies, fees or partial captures modelled wrong.

Reconciliation should find almost nothing. If it finds a lot, it isn’t the solution: it’s the sign that one of the earlier pieces is broken.

Several providers: the orchestrator

With more than one provider comes the temptation of automatic failover: if provider A fails, retry with B. It’s the easiest way to charge twice. “Failed” rarely means “didn’t charge”: a timeout from A may well be a charge that went through.

The orchestrator should only switch providers when the error is definitive and happened before the charge: the provider was down before accepting the request, the payment method isn’t supported, or there was an explicit decline. On unknown, resolve it with A before touching B.

Common mistakes

  • Generating the idempotency key on the server. Then it protects nothing: every client retry would create a new key.
  • Expiring keys too early. They must outlive the client’s longest retry window.
  • Treating a timeout as a failure. It’s an unknown.
  • Processing the webhook inside the HTTP request. The provider’s retries create duplicates.
  • Trusting webhooks as the only source of truth. You need active polling and reconciliation.
  • Not storing the provider’s raw response. When a dispute comes, it’s the only evidence of what happened.

None of these pieces is hard on its own. What makes the system robust is assuming, from the design onwards, that any call can half-fail, and that the right question isn’t “did it work?” but “what do I know for sure about the state of this payment?”.