I Had Duplicate Transactions in Production — Here's How Redis Locks Fixed It
Duplicate transactions don’t always happen because a user double-clicks a button. A request can reach the server, the operation can finish, and then the response can get lost. From the client’s point of view, it looks like nothing happened—so it retries.
I ran into this class of problem while working on production transaction flows. The tricky part was that every individual request looked valid. The issue showed up when two requests for the same operation arrived close together and both tried to move the transaction forward.
A Redis lock helped us stop those requests from processing the same transaction at the same time. But the lock wasn’t the whole solution. We also needed idempotency and a database constraint, because retries and process failures can happen even when a lock is in place.
How the race happens
Imagine a transaction in a pending state. Two requests arrive almost together, possibly at different Node.js instances:
Request A: read status = pending
Request B: read status = pending
Request A: apply the transaction
Request B: apply the transaction again
Both requests made their decision using the same old state. If the application only checks “is this pending?” and then updates it in a separate step, the check and the update can race.
This can happen because of a retry after a timeout, a double submit, two workers receiving the same job, or traffic being handled by separate application instances. A JavaScript variable or in-memory mutex only protects one process. It can’t coordinate the other Node.js instances.
Why we used a Redis lock
We needed a quick way to make one instance at a time own the short critical section for a particular transaction. Redis lets us create a lock key only if it doesn’t already exist, and attach an expiry to it in the same operation:
SET lock:transaction:<transaction-id> <unique-token> NX PX 10000
NX means “set this only if the key does not already exist.” PX gives the lock a time-to-live in milliseconds. The first request gets the lock; another request for the same transaction sees that it is already being processed.
The exact lock duration depends on the work being protected. The 10-second value above is just an example, not a value to copy blindly.
Acquire and release the lock safely
Each attempt needs its own random token. When releasing a lock, the process should check that the token still matches. Otherwise, a slow request could accidentally delete a lock that expired and was acquired by another process.
Here’s a simplified example using Node.js and node-redis:
import { randomUUID } from "node:crypto";
const acquireLockScript = `
if redis.call("SET", KEYS[1], ARGV[1], "NX", "PX", ARGV[2]) then
return 1
end
return 0
`;
const releaseLockScript = `
if redis.call("GET", KEYS[1]) == ARGV[1] then
return redis.call("DEL", KEYS[1])
end
return 0
`;
async function processTransaction(transactionId, redis) {
const lockKey = "lock:transaction:" + transactionId;
const lockToken = randomUUID();
const lockTtlMs = 10_000; // Example only; choose based on the critical section.
const acquired = await redis.eval(acquireLockScript, {
keys: [lockKey],
arguments: [lockToken, String(lockTtlMs)],
});
if (Number(acquired) !== 1) {
// Another request is handling it. Check the durable transaction state
// or return an "in progress" response; don't apply the operation again.
return { status: "in_progress" };
}
try {
return await applyTransactionWithIdempotencyCheck(transactionId);
} finally {
await redis.eval(releaseLockScript, {
keys: [lockKey],
arguments: [lockToken],
});
}
}
The Lua release script checks the token and deletes the key as one atomic operation. A plain DEL lockKey in finally is unsafe: by the time it runs, the original lock may have expired and another request may own the key.
The transaction function above is intentionally left as an application-specific step. The lock controls who enters the critical section; it should not contain the business rules for changing balances, orders, or payment state.
The lock reduced overlap. Idempotency handled retries.
A lock helps when requests overlap. It does not recognize that a retry arriving later is the same user action. For that, the client should send an idempotency key that stays the same for every retry of the same logical operation.
The server can store that key with the transaction and return the original result if it has already been processed. In MongoDB, a unique compound index can enforce one record for a given account and idempotency key:
db.transactions.createIndex(
{ accountId: 1, idempotencyKey: 1 },
{ unique: true }
);
The names here are examples; use the fields that identify one logical transaction in your own schema. The database constraint is important because it still protects the invariant if the application has a bug or two requests get past the lock logic.
For state changes, I also want the update itself to be conditional. For example, transition a record from pending to processing only if it is still pending, and treat “no document matched” as another request having already changed it. The check and state change should not be separate unprotected steps.
What if the process crashes or the lock expires?
The expiry prevents a lock from remaining forever if a process dies. It also creates an edge case: the original process might still be running when its lock expires. A second process could then acquire the lock.
That’s why a Redis lock should protect a short, bounded critical section. Keep the TTL longer than the normal work, but don’t assume that a TTL alone guarantees only one process can ever act. For long-running work, use durable state transitions and, where needed, fencing tokens so an older worker can’t overwrite work from a newer one.
If a transaction calls an external payment provider, don’t rely on a Redis lock held across a long network call to guarantee exactly-once behavior. Pass a stable idempotency key to the provider when supported, store the provider’s transaction reference, and make retries safe on your side too. A timeout can leave the caller unsure whether the provider completed the operation.
What I’d keep in place
For a transaction flow, I think about these as separate safeguards:
- A Redis lock to reduce simultaneous processing of the same transaction.
- An idempotency key so retries return the existing result instead of creating another one.
- A unique database constraint to enforce the invariant durably.
- A conditional state transition so stale reads can’t authorize the same change twice.
-
A recovery path for operations left in
processingafter a crash or timeout.
The Redis lock made the race much less likely in the hot path. The durable checks are what made the transaction flow safe when requests were retried, workers stopped, or timing didn’t go the way we expected.
That was the main lesson for me: a lock is useful coordination, but it isn’t a substitute for idempotency or a database rule. When money or other important state is involved, I want the source of truth to be able to reject a duplicate even if the lock layer fails.
Top comments (0)