DEV Community

Jesse Muuya
Jesse Muuya

Posted on

What Happens After the Fifth Retry? Escalating Stuck Payments Instead of Looping Forever

A week ago I wrote about a payment that got stuck in PENDING, and the fix I shipped: a reconciliation job that only marks a payment settled after confirming the customer actually got what they paid for, not before. A reader, Mihai Perdum, left a comment on that post that I didn't have a good answer to at the time:

Is there a retry-count or age threshold that pages someone if fulfillment repeatedly fails structurally rather than transiently?

There wasn't. The job I'd shipped would retry a failing payment every five minutes, forever, with no distinction between "the database had a blip, try again" and "this event was deleted and will never exist again, stop trying." A payment stuck for a structural reason would loop quietly until it aged out of the sweep window and disappeared with nobody told.

This is the fix, and what testing it taught me about my own first attempt.

The gap

Here's the shape of the job from the last post:

try {
  await onPaymentSettled(payment);
  await db.payment.update({ where: { id: payment.id }, data: { status: 'SUCCESS' } });
} catch (err) {
  // reverts to PENDING, retries next tick
}
Enter fullscreen mode Exit fullscreen mode

Every failure looked the same to this code: log it, leave it PENDING, try again in five minutes. That's correct for a transient failure — a timeout, a dropped connection. It's wrong for anything that will never succeed on retry: a deleted event, a sold-out ticket tier, a foreign key that no longer resolves. Those don't need patience, they need a human.

The fix: count, classify, escalate

Two additions. An attempt counter, and a classifier that distinguishes "try again" from "stop and tell someone."

} catch (err) {
  const attempts = payment.fulfilmentAttempts + 1;
  const structural = isStructuralFailure(err);
  const escalate = structural || attempts >= MAX_FULFILMENT_ATTEMPTS;

  await db.payment.update({
    where: { id: payment.id },
    data: {
      fulfilmentAttempts: attempts,
      lastFulfilmentError: String(err.message).slice(0, 500),
      ...(escalate && { needsAttention: true }),
    },
  });

  if (escalate) {
    await notifyAdmins(payment, err);
  }
}
Enter fullscreen mode Exit fullscreen mode

The sweep query excludes anything flagged needsAttention, so an escalated payment stops being silently retried and instead sits somewhere a human will actually see it.

Where my first version got it wrong

The classifier's job is deciding which failures are structural. My first pass:

function isStructuralFailure(err) {
  return err.name === 'PrismaClientValidationError' || err.code === 'P2025';
}
Enter fullscreen mode Exit fullscreen mode

I tested it against a real synthetic failure — a payment referencing an event that didn't exist — expecting immediate escalation. Instead:

❌ Fulfilment failed for payment OTV-TEST-RETRYCAP-A (attempt 1, transient) — reverting to PENDING for retry
Enter fullscreen mode Exit fullscreen mode

Four more identical failures later:

❌ Fulfilment failed for payment OTV-TEST-RETRYCAP-A (attempt 5, transient) — ESCALATING, needs manual attention
Enter fullscreen mode Exit fullscreen mode

It escalated, but only after burning through the whole retry budget on something that could never have succeeded. The actual error was a plain 404 thrown by application code, not a Prisma error at all, so my classifier didn't recognize it. Fixed:

function isStructuralFailure(err) {
  const httpStatus = typeof err.getStatus === 'function' ? err.getStatus() : 0;
  return (
    (httpStatus >= 400 && httpStatus < 500) ||
    err.name === 'PrismaClientValidationError' ||
    err.code === 'P2025' || err.code === 'P2003' || err.code === 'P2002'
  );
}
Enter fullscreen mode Exit fullscreen mode

Same test, rebuilt:

❌ Fulfilment failed for payment OTV-TEST-RETRYCAP-A (attempt 1, structural) — ESCALATING, needs manual attention
Enter fullscreen mode Exit fullscreen mode

One attempt instead of five. The lesson: a classifier is a claim about what your errors look like, and the only way to know if the claim is true is to throw a real one at it and watch.

Making sure the alert doesn't page twice

Webhook deliveries and retries can both re-enter the failure path after a payment is already flagged. Without a guard, an already-escalated payment would generate a fresh alert on every subsequent retry attempt — not wrong, exactly, but noisy enough that people learn to ignore it. The fix is checking whether the flag was already set before sending anything:

if (escalate && !payment.needsAttention) {
  await notifyAdmins(payment, err);
}
Enter fullscreen mode Exit fullscreen mode

Tested by escalating a payment once, confirming one notification row, then sending the identical failing delivery again and confirming the notification count stayed at one. It did.

Shipping the same discipline in the sold version

I maintain a small payments engine (African Payment Gateways Engine) that includes a reconciliation script with the same fulfil-before-settle shape from the last post. Auditing it after building this fix found it had the exact gap Mihai described — no counter, no classifier, retry forever. Same fix, ported: an attempt counter and needsAttention column on the payment row, the same structural-vs-transient split, the same exclusion from the sweep once flagged.

Tested the same way — a synthetic payment referencing a nonexistent event:

❌ Reconcile failed for TEST-REF-2 (attempt 1, structural) — ESCALATING, needs manual attention: event not found
Enter fullscreen mode Exit fullscreen mode

And confirmed an escalated row stops being swept while ordinary pending payments keep being checked normally:

📡 Found 0 PENDING records awaiting verification polling.
Enter fullscreen mode Exit fullscreen mode

(The escalated row, still PENDING in the database, correctly excluded from that batch.)

What this doesn't solve

The counter only fires when something actually calls the fulfilment hook and it throws. If the job that's supposed to be running stops running entirely — crashes, gets undeployed, a scheduler misconfiguration — a paid-but-unfulfilled payment sits there with nothing watching it, because nothing ever entered the catch block to start counting. Closing that gap needs a second, independent check: something that looks at "how old is the oldest payment the provider says succeeded but we haven't marked settled," regardless of whether the reconciliation job ran at all. That's not built yet. It's the next thing on the list.

Thanks to Mihai for the comment that started this. If your reconciliation job treats every failure the same way, it's worth checking whether that's actually true — the fastest way I know is the same one that caught my own mistake: throw the exact error you expect at it, and read what comes back.


If you're building this yourself: African Payment Gateways Engine ships the fulfil-before-settle reconciliation pattern with this retry-cap and escalation built in. The free africa-payments-utils package covers phone normalization and webhook signature verification if that's all you need.

Top comments (0)