ID

Home·Writings

Redis Distributed Locks for Payment Processing

By ·

Use Redis locks around payment work without trusting them too much: SET NX PX, Lua release, TTLs, fencing tokens, and database guarantees.

On a payments platform I worked on, the scary bug was not that a queue retried a failed transfer. Retries are normal. The scary part was two workers picking up work for the same transfer reference close enough together that both believed they were the only one responsible.

That kind of bug usually starts innocently. A mobile client times out, an API writes a job, the queue redelivers it, a worker crashes after calling a provider, and suddenly you have duplicate work moving through the system. A Redis lock can reduce the chance of two workers doing the same thing at the same time, but it must never be the only thing protecting money.

This post uses Ledgerline, a fictional digital bank, to show the pattern I use: Redis as a coordination tool, PostgreSQL as the source of correctness, and fencing tokens when stale workers can still cause damage.

What the lock is for

The lock protects a short critical section. For example, one worker should be allowed to process a transfer reference at a time:

const lockKey = `lock:transfer:${transferId}`;
const token = randomUUID();
const ttlMs = 30_000;

const acquired = await redis.set(lockKey, token, 'NX', 'PX', ttlMs);
if (acquired !== 'OK') {
  throw new ConflictException('Transfer is already being processed');
}

NX means "only set this if the key does not exist". PX gives the lock an expiry in milliseconds, so a crashed worker does not hold it forever.

This is useful, but it is only a lease. The worker has the right to proceed for roughly 30 seconds. It does not have a permanent claim, and Redis cannot tell the rest of your system that this worker is still alive.

Release with a token, not DEL

The worker that releases a lock must prove it still owns it. A plain DEL lock:transfer:123 is dangerous:

  1. Worker A acquires the lock for 30 seconds.
  2. Worker A pauses for 40 seconds because of GC, network delay, or a slow provider.
  3. The lock expires.
  4. Worker B acquires the same key.
  5. Worker A resumes and deletes Worker B's lock.

The safe release is a small Lua script that checks the token before deleting:

const RELEASE_LOCK = `
if redis.call("GET", KEYS[1]) == ARGV[1] then
  return redis.call("DEL", KEYS[1])
else
  return 0
end
`;

await redis.eval(RELEASE_LOCK, 1, lockKey, token);

Now Worker A can only release the lock it actually created.

Keep the database in charge

Here is the more important rule: the lock is not the guarantee. The database is.

For Ledgerline, every transfer has a stable reference and a state machine:

CREATE TABLE transfers (
  id uuid PRIMARY KEY,
  user_id uuid NOT NULL,
  amount bigint NOT NULL,
  currency text NOT NULL,
  provider_reference text NOT NULL UNIQUE,
  status text NOT NULL CHECK (status IN (
    'pending', 'processing', 'sent', 'failed', 'reconciling'
  )),
  version bigint NOT NULL DEFAULT 0,
  updated_at timestamptz NOT NULL DEFAULT now()
);

Before a worker sends money, it claims the row with a conditional update:

UPDATE transfers
SET status = 'processing',
    version = version + 1,
    updated_at = now()
WHERE id = $1
  AND status IN ('pending', 'reconciling')
RETURNING id, provider_reference, version;

If this returns no rows, another worker already moved the transfer forward. Stop. Do not call the provider.

That one SQL statement does more for correctness than the Redis lock. The lock reduces duplicate pressure on the database and provider. The row update decides who is allowed to proceed.

I use the same thinking for API idempotency: Redis is a fast path, but Postgres carries the real invariant. I covered that in idempotent payments with Redis and PostgreSQL.

Set the TTL from the work, then make the work smaller

A common mistake is picking a lock TTL that "feels safe", like 5 minutes. That hides design problems.

The TTL should cover the critical section, with enough margin for normal jitter. If processing usually takes 2 seconds and the provider timeout is 8 seconds, a 30 second TTL is reasonable. If the job can take 10 minutes, do not hold one Redis lock across the whole job. Split the work:

  • claim the database row;
  • create or reuse the provider reference;
  • call the provider with a strict timeout;
  • store the result or mark the row for reconciliation;
  • release the lock.

Long-running work should be resumable. A lock should not be the thing keeping your process honest for several minutes.

Use fencing tokens for stale workers

The hardest failure mode is a stale worker. It acquired the lock, got delayed past the TTL, and then continued doing work after another worker had taken over.

This is where a fencing token helps. Every successful claim increments a monotonic version in the database. Downstream writes must include that version and only apply if it is still current.

const row = await db.oneOrNone<{
  provider_reference: string;
  version: number;
}>(`
  UPDATE transfers
  SET status = 'processing',
      version = version + 1,
      updated_at = now()
  WHERE id = $1
    AND status IN ('pending', 'reconciling')
  RETURNING provider_reference, version
`, [transferId]);

if (!row) return;

const result = await provider.send({
  reference: row.provider_reference,
  amount,
  currency,
});

await db.none(`
  UPDATE transfers
  SET status = $2,
      updated_at = now()
  WHERE id = $1
    AND version = $3
`, [transferId, result.ok ? 'sent' : 'failed', row.version]);

If Worker A resumes late with version 8, but Worker B has already moved the row to version 9, Worker A's final update affects zero rows. It cannot overwrite newer state.

When the external provider supports idempotency keys, send provider_reference too. That protects you even if two calls escape your system. If the provider does not support idempotency, reconciliation becomes mandatory because you need a way to detect what actually happened.

What about Redlock?

Redis has a distributed lock algorithm called Redlock. It tries to acquire locks across multiple independent Redis nodes so a single node failure does not decide ownership.

I do not start there for payment processing. The debate around Redlock is not the main issue. The main issue is that a distributed lock still cannot replace application-level correctness. Network partitions, pauses, expired leases and retries still exist.

For most product teams, the better order is:

  • one well-operated Redis deployment for coordination;
  • database constraints and conditional updates for correctness;
  • stable provider references for external idempotency;
  • reconciliation for uncertain provider outcomes;
  • fencing tokens for stale workers.

If you truly need a stronger distributed coordination system, look at something designed for it, like etcd or ZooKeeper. But even then, keep the money invariant in your database and ledger.

Failure modes to design for

Redis is down. The system should either fall back to the database claim path or temporarily reject new processing with a retryable error. It should not process money without the database guard.

The worker crashes after calling the provider. Mark the transfer as reconciling when the provider result is unknown. A background job should query the provider by provider_reference and finish the transfer.

The provider times out. Timeout does not mean failure. It means "unknown". Do not let a second worker immediately send another transfer unless the provider reference makes that retry idempotent.

The lock expires while the provider call is still running. Fencing tokens stop stale state writes. Provider idempotency stops duplicate external movement. The Redis lock alone stops neither.

Takeaways

PieceJob
SET key token NX PX ttlAcquires a short-lived lease
Lua release scriptPrevents one worker from deleting another worker's lock
Conditional SQL updateThe real claim on the transfer
version fencing tokenBlocks stale workers from overwriting newer state
Provider referenceGives the external system a deduplication key
Reconciliation jobResolves unknown outcomes after timeouts and crashes

Redis locks are useful, but only when they are treated as coordination, not correctness. For payment processing, correctness comes from database constraints, explicit state transitions, idempotent provider references and reconciliation. The lock is just there to keep the system calmer while those stronger guarantees do the real work.