DEV Community

Cover image for The Happy Path Is Lying to You: What Production Systems Taught Me About Idempotency, Reversals, and Financial Truth
Williams Ashibuogwu
Williams Ashibuogwu

Posted on

The Happy Path Is Lying to You: What Production Systems Taught Me About Idempotency, Reversals, and Financial Truth

There is a point where a software project stops being "an app."

The UI might still look the same.

There are forms, tables, buttons, APIs, and database rows.

But suddenly a bug does not mean:

The page displayed the wrong thing.

It means:

The system charged twice.

Or:

The accounting entry was duplicated.

Or:

Someone reversed a transaction they were not authorized to touch.

Or:

The balance shown to the user was stale, but the system treated it as authoritative.

That changes the way you have to build software.

Over the last few weeks, I have been working on two production systems from very different domains.

One handles business operations such as sales, inventory, accounting, reporting, staff permissions, and branch activity.

The other handles accounts, balances, transfers, pricing, reconciliation, and asynchronous financial events.

Different products.

Same engineering problem.

Once software starts carrying real consequences, the main job is no longer creating features.

The job is preserving truth when the system is under pressure.

This post is about some of the patterns that become important at that point.


CRUD works until your rows start meaning something

A typical application starts simply:

Create
Read
Update
Delete
Enter fullscreen mode Exit fullscreen mode

You build a few resources.

Then you add authentication.

Then permissions.

Then reports.

Then background jobs.

Then integrations.

Eventually, one of those database rows starts representing something that matters outside the database.

A stock movement affects whether an item can be sold.

A payment affects receivables.

A journal entry changes financial statements.

A transfer affects someone's available money.

At that point, this becomes dangerous:

UPDATE transactions
SET amount = 125000
WHERE id = 42;
Enter fullscreen mode Exit fullscreen mode

Why?

Because you may have just changed history.

That transaction might already have affected:

  • an account balance
  • a journal
  • a report
  • a notification
  • an external transfer
  • a reconciliation process

The system now needs something stronger than CRUD.

It needs invariants.


Invariants are the real feature

For a manual accounting entry, the screen can be simple.

Account Debit Credit
Cash 50,000 0
Sales Revenue 0 50,000

The hard part is not rendering those inputs.

The hard part is guaranteeing that the data can never become financially impossible.

For example:

function validateJournal(lines: JournalLine[]) {
  if (lines.length < 2) {
    throw new Error("A journal requires at least two lines");
  }

  let debit = 0n;
  let credit = 0n;

  for (const line of lines) {
    const hasDebit = line.debit > 0n;
    const hasCredit = line.credit > 0n;

    if (hasDebit === hasCredit) {
      throw new Error(
        "Each line must contain either a debit or a credit"
      );
    }

    debit += line.debit;
    credit += line.credit;
  }

  if (debit === 0n || debit !== credit) {
    throw new Error("Journal is not balanced");
  }
}
Enter fullscreen mode Exit fullscreen mode

That validation is useful.

It is not enough.

The system still needs to know:

  • Is the account active?
  • Can this account receive manual postings?
  • Is the accounting period open?
  • Does the user belong to this branch?
  • Has this exact operation already been posted?
  • Did the account state change after the form loaded?

That last question is where things get interesting.


Validate again at commit time

Imagine a finance user opens a posting screen at 2:00 PM.

The accounting period is open.

They prepare a long journal.

At 2:05 PM, an administrator closes the accounting period.

At 2:07 PM, the user clicks Post.

If the application only validated the period when the form loaded, you now have a posting inside a closed period.

The rule needs to be enforced inside the transaction that performs the write.

Conceptually:

await db.transaction(async tx => {
  const period = await tx.accountingPeriods
    .where({ id: input.periodId })
    .forUpdate()
    .first();

  if (!period || period.status !== "OPEN") {
    throw new Error("Accounting period is closed");
  }

  const accountIds = [
    ...new Set(input.lines.map(line => line.accountId))
  ].sort();

  const accounts = await tx.accounts
    .whereIn("id", accountIds)
    .orderBy("id")
    .forUpdate();

  for (const account of accounts) {
    if (!account.active || !account.manualPostingAllowed) {
      throw new Error("Account cannot accept this posting");
    }
  }

  await createJournal(tx, input);
});
Enter fullscreen mode Exit fullscreen mode

The important idea is not the syntax.

It is this:

A rule that matters financially should be checked as close as possible to the write that makes the operation permanent.


The check-then-insert race

Here is one of my favorite production bugs because the code often looks completely reasonable.

const existing = await db.journals.findFirst({
  where: {
    sourceType,
    sourceId,
    sourceEvent
  }
});

if (existing) {
  return existing;
}

return db.journals.create({
  data: {
    sourceType,
    sourceId,
    sourceEvent
  }
});
Enter fullscreen mode Exit fullscreen mode

Looks fine.

Until two requests arrive together.

Request A

Checks.

Nothing exists.

Request B

Checks.

Nothing exists.

Request A

Inserts.

Request B

Also tries to insert.

If the database has no uniqueness constraint, congratulations:

You now have two financial effects.

If the database does have a uniqueness constraint, that is much better.

But Request B may still receive an ugly database exception even though the operation actually succeeded.

The stronger implementation treats the constraint as part of the control flow.

try {
  return await db.journals.create({
    data: {
      sourceType,
      sourceId,
      sourceEvent
    }
  });
} catch (error) {
  if (!isUniqueViolation(error)) {
    throw error;
  }

  const winner = await db.journals.findFirst({
    where: {
      sourceType,
      sourceId,
      sourceEvent
    }
  });

  if (!winner) {
    throw error;
  }

  return winner;
}
Enter fullscreen mode Exit fullscreen mode

And at the database layer:

CREATE UNIQUE INDEX journal_source_identity_unique
ON journals (
  source_type,
  source_id,
  source_event
);
Enter fullscreen mode Exit fullscreen mode

Now the database decides who wins.

The application knows how to handle the loser.

Important: Application checks can reduce unnecessary conflicts. Database constraints enforce reality.


Idempotency is not a frontend feature

A common mistake is thinking this solves duplicate operations:

button.disabled = true;
Enter fullscreen mode Exit fullscreen mode

It helps UX.

It does not solve duplication.

A user can still:

  • open two browser tabs
  • resend an HTTP request
  • refresh after a timeout
  • trigger a mobile retry
  • cause a queue redelivery
  • receive the same webhook twice

A consequential operation should have a stable identity.

For example:

{
  "operation_id": "01K7H4N9S9P1D9JYC0H7...",
  "amount_minor": 250000,
  "currency": "NGN"
}
Enter fullscreen mode Exit fullscreen mode

Then persist that identity.

CREATE UNIQUE INDEX transfer_operation_unique
ON transfers(operation_id);
Enter fullscreen mode Exit fullscreen mode

Retries can now become safe:

const previous = await transfers.findByOperationId(operationId);

if (previous) {
  if (previous.requestHash !== hash(input)) {
    throw new ConflictError(
      "Operation ID was already used with different data"
    );
  }

  return previous;
}
Enter fullscreen mode Exit fullscreen mode

This gives you two useful behaviors:

same key + same payload
=> return previous result

same key + different payload
=> reject
Enter fullscreen mode Exit fullscreen mode

That is much stronger than blindly returning:

Already exists
Enter fullscreen mode Exit fullscreen mode

Reversals are better than rewriting history

Suppose a journal entry is wrong.

The tempting solution:

UPDATE journal_lines
SET debit = 0, credit = 100000
WHERE id = 9001;
Enter fullscreen mode Exit fullscreen mode

Now you have a clean-looking current state.

But you lost historical truth.

A better model:

Journal #123
Debit Cash                 100,000
Credit Revenue             100,000

Reversal #456
Credit Cash                100,000
Debit Revenue              100,000
Enter fullscreen mode Exit fullscreen mode

The original event still exists.

The correction exists separately.

They are linked.

type Journal = {
  id: string;
  reversedBy?: string;
  reverses?: string;
};
Enter fullscreen mode Exit fullscreen mode

And the reversal command itself should be idempotent.

async function reverseJournal(id: string) {
  return db.transaction(async tx => {
    const original = await tx.journals
      .where({ id })
      .forUpdate()
      .first();

    if (!original) {
      throw new NotFoundError();
    }

    if (original.reversedBy) {
      return tx.journals.find(original.reversedBy);
    }

    const reversal = await createOppositeJournal(
      tx,
      original
    );

    await tx.journals.update(original.id, {
      reversedBy: reversal.id
    });

    return reversal;
  });
}
Enter fullscreen mode Exit fullscreen mode

Retrying the reversal should not produce:

reverse
reverse
reverse
reverse
Enter fullscreen mode Exit fullscreen mode

It should return the one reversal that already exists.


Financial history should be append-heavy

This pattern is useful far beyond journals.

Payments, refunds, voids, settlement corrections, and adjustments can all be represented as related events.

That makes the system much easier to explain later.

Why does this balance equal ₦42,300?

Instead of saying:

> Because that is the value currently stored in the balance column.

The system can explain:

Opening balance         +100,000
Purchase                 -35,000
Transfer                 -20,000
Fee                       -1,500
Refund                    +5,000
Adjustment                -6,200
--------------------------------
Current balance           42,300
Enter fullscreen mode Exit fullscreen mode

That explainability becomes extremely important once support, finance teams, auditors, and users start asking questions.


Your browser should not decide how much money moves

Suppose a transfer has:

Amount:       10,000
Network fee:     100
Service fee:      50
Total debit:  10,150
Enter fullscreen mode Exit fullscreen mode

Do not let the client invent the final debit.

This is weak:

const total = amount + networkFee + serviceFee;

await api.post("/transfer", {
  amount,
  total
});
Enter fullscreen mode Exit fullscreen mode

Why trust total?

The client is not authoritative.

A stronger flow looks like this.

Step 1: Create a quote

POST /transfer-quotes
Content-Type: application/json
Enter fullscreen mode Exit fullscreen mode
{
  "amount_minor": 1000000,
  "destination_id": "acct_123"
}
Enter fullscreen mode Exit fullscreen mode

Server response:

{
  "quote_id": "quote_abc",
  "amount_minor": 1000000,
  "fee_minor": 15000,
  "total_debit_minor": 1015000,
  "expires_at": "2026-10-07T20:30:00Z"
}
Enter fullscreen mode Exit fullscreen mode

Step 2: Display the quote

The UI renders what the server calculated.

Step 3: Authorize that exact debit

POST /transfers
Content-Type: application/json
Enter fullscreen mode Exit fullscreen mode
{
  "quote_id": "quote_abc",
  "authorization": {
    "type": "pin",
    "proof": "..."
  }
}
Enter fullscreen mode Exit fullscreen mode

The backend verifies the quote itself.

Not a client-provided total.

The server should always be able to explain exactly what amount the user authorized.


Minor units deserve boring code

Money bugs are often surprisingly boring.

Consider:

const minimum = 50000;
Enter fullscreen mode Exit fullscreen mode

What does that mean?

₦50,000?

₦500?

50,000 kobo?

Without a convention, the code is ambiguous.

I prefer names like:

const minimumMinor = 50_000n;
Enter fullscreen mode Exit fullscreen mode

Formatting happens only at the presentation edge.

function formatMoney(minor: bigint): string {
  return new Intl.NumberFormat("en-NG", {
    style: "currency",
    currency: "NGN"
  }).format(Number(minor) / 100);
}
Enter fullscreen mode Exit fullscreen mode

Arithmetic stays integer-based:

const totalDebitMinor =
  amountMinor +
  serviceFeeMinor +
  networkFeeMinor;
Enter fullscreen mode Exit fullscreen mode

Not:

const total = 100.1 + 20.2;
Enter fullscreen mode Exit fullscreen mode

You do not want floating-point surprises anywhere near money.


Reconciliation is not cleanup

External systems introduce an uncomfortable state.

You send a transfer request.

Then:

Your request times out.
Enter fullscreen mode Exit fullscreen mode

What happened?

You do not know.

That is different from failure.

This can be dangerous:

try {
  await sendTransfer(input);
} catch {
  transfer.status = "FAILED";
}
Enter fullscreen mode Exit fullscreen mode

The external system may have accepted the transaction before the connection disappeared.

Now your local system says:

FAILED
Enter fullscreen mode Exit fullscreen mode

while the external system later says:

SUCCESS
Enter fullscreen mode Exit fullscreen mode

Then the user retries.

The original transfer eventually completes.

The retry completes too.

A better state model includes uncertainty.

type TransferStatus =
  | "PENDING"
  | "SUBMITTED"
  | "UNKNOWN"
  | "CONFIRMED"
  | "FAILED";
Enter fullscreen mode Exit fullscreen mode

After an ambiguous timeout:

transfer.status = "UNKNOWN";
Enter fullscreen mode Exit fullscreen mode

Then reconciliation decides the final state.

const external = await fetchTransferStatus(
  transfer.externalReference
);

switch (external.status) {
  case "successful":
    await settleExactlyOnce(transfer, external);
    break;

  case "failed":
    await markFailed(transfer);
    break;

  default:
    await scheduleAnotherCheck(transfer);
}
Enter fullscreen mode Exit fullscreen mode

UNKNOWN is a valid state.

Pretending uncertainty does not exist is not.


Webhooks can arrive twice

Actually, assume they will.

This is dangerous:

app.post("/webhook", async req => {
  await creditAccount(req.body);
});
Enter fullscreen mode Exit fullscreen mode

A webhook sender may retry because:

  • your server returned 500
  • the connection closed too early
  • they did not receive your acknowledgement
  • delivery is intentionally at-least-once

A safer approach:

await db.transaction(async tx => {
  const exists = await tx.webhookEvents.findUnique({
    where: {
      eventId: event.id
    }
  });

  if (exists) {
    return;
  }

  await applyFinancialEffect(tx, event);

  await tx.webhookEvents.create({
    data: {
      eventId: event.id,
      processedAt: new Date()
    }
  });
});
Enter fullscreen mode Exit fullscreen mode

Plus a database constraint:

CREATE UNIQUE INDEX webhook_event_unique
ON webhook_events(event_id);
Enter fullscreen mode Exit fullscreen mode

Again, the database is part of your correctness model.


Stale data can be worse than no data

One particularly interesting class of bugs happens when the application already has a value.

For example:

{
  "balance_minor": 0
}
Enter fullscreen mode Exit fullscreen mode

The application sees that and assumes:

Balance is zero.
Enter fullscreen mode Exit fullscreen mode

But that value may have been captured hours ago.

Meanwhile, a more recent confirmed event says funds exist.

The local value is not missing.

It is stale.

That can be more dangerous because it looks authoritative.

A better representation carries metadata.

type BalanceSnapshot = {
  amountMinor: bigint;
  observedAt: Date;
  source: "ledger" | "external-confirmation" | "cache";
};
Enter fullscreen mode Exit fullscreen mode

Now the application can reason about evidence.

function chooseBalance(
  a: BalanceSnapshot,
  b: BalanceSnapshot
) {
  if (
    a.source === "external-confirmation" &&
    b.source === "cache"
  ) {
    return a;
  }

  return a.observedAt > b.observedAt ? a : b;
}
Enter fullscreen mode Exit fullscreen mode

The real implementation may be more complicated.

The important lesson is simpler:

Presence does not imply authority.


Authorization belongs inside the domain

A UI might hide a reversal button.

That is not security.

This:

{user.canReverse && (
  <button>Reverse</button>
)}
Enter fullscreen mode Exit fullscreen mode

is purely presentation.

The real check must happen server-side.

async function reverseEntry(user, entryId) {
  const entry = await entries.find(entryId);

  if (!user.permissions.includes("journal.reverse")) {
    throw new ForbiddenError();
  }

  if (entry.branchId !== user.branchId) {
    throw new ForbiddenError();
  }

  return reversalService.reverse(entry);
}
Enter fullscreen mode Exit fullscreen mode

Notice that there are two separate questions:

Can this user reverse entries?

Can this user reverse this specific entry?
Enter fullscreen mode Exit fullscreen mode

The second one is where many authorization bugs appear.


Lock rows in a deterministic order

Concurrency can also create deadlocks.

Suppose one transaction locks:

Account 8
then Account 12
Enter fullscreen mode Exit fullscreen mode

while another locks:

Account 12
then Account 8
Enter fullscreen mode Exit fullscreen mode

Both can wait on each other.

One useful mitigation is deterministic ordering.

const ids = [...new Set(accountIds)]
  .sort((a, b) => a.localeCompare(b));

const accounts = await tx.accounts
  .whereIn("id", ids)
  .orderBy("id")
  .forUpdate();
Enter fullscreen mode Exit fullscreen mode

Every transaction tries to acquire the same resources in the same order.

Not every deadlock disappears.

But you remove an entire class of avoidable ones.


UI correctness is still correctness

Not every production bug is deep distributed-systems theory.

One issue looked approximately like this:

<div class="product-code text-xs">
  ABC-123
</div>
Enter fullscreen mode Exit fullscreen mode

The data existed.

The backend returned it.

The test found it in the HTML.

But an old CSS rule contained:

.product-row .text-xs {
  display: none;
}
Enter fullscreen mode Exit fullscreen mode

The automated test said:

PASS
Enter fullscreen mode Exit fullscreen mode

The user said:

Where is the product code?
Enter fullscreen mode Exit fullscreen mode

Both observations were technically correct.

That is why this test:

expect(html).toContain("ABC-123");
Enter fullscreen mode Exit fullscreen mode

is not equivalent to:

await expect(
  page.getByText("ABC-123")
).toBeVisible();
Enter fullscreen mode Exit fullscreen mode

Sometimes the correct layer to test is the browser.


A passing request does not always mean a successful operation

Another pattern worth separating is transport success from business success.

Suppose an API call returns HTTP 200:

{
  "status": "accepted"
}
Enter fullscreen mode Exit fullscreen mode

That might only mean:

The request was accepted for processing.

It may not mean:

The financial operation is complete.

This distinction should exist in your own models too.

type OperationState =
  | "CREATED"
  | "AUTHORIZED"
  | "SUBMITTED"
  | "PROCESSING"
  | "SETTLED"
  | "FAILED";
Enter fullscreen mode Exit fullscreen mode

Collapsing everything into:

success = true
Enter fullscreen mode Exit fullscreen mode

removes useful information.

For asynchronous operations, state machines are usually more honest than booleans.


Think in commands, not just endpoints

A useful shift for consequential systems is to stop thinking only in HTTP routes.

Instead of:

POST /journal
POST /transfer
POST /refund
Enter fullscreen mode Exit fullscreen mode

think:

PostJournal
SubmitTransfer
ReverseJournal
RefundPayment
SettleTransfer
Enter fullscreen mode Exit fullscreen mode

A command can have:

  • an operation identity
  • authorization rules
  • invariants
  • state preconditions
  • side effects
  • an idempotency policy
  • an audit trail

For example:

type PostJournalCommand = {
  operationId: string;
  actorId: string;
  branchId: string;
  postingDate: string;
  lines: JournalLine[];
};
Enter fullscreen mode Exit fullscreen mode

Then the command handler becomes the boundary that protects the operation.

async function handlePostJournal(
  command: PostJournalCommand
) {
  authorize(command);
  validate(command);

  return db.transaction(async tx => {
    return postJournalExactlyOnce(tx, command);
  });
}
Enter fullscreen mode Exit fullscreen mode

The HTTP controller becomes much less interesting.

That is usually a good sign.


A useful retry rule

I increasingly think about operations using this simple question:

If I execute this again, what should happen?

There are only a few acceptable answers.

Read operation

Execute again
=> return current data
Enter fullscreen mode Exit fullscreen mode

Idempotent command

Execute again with same identity
=> return previous result
Enter fullscreen mode Exit fullscreen mode

Conflicting retry

Same identity + different payload
=> reject
Enter fullscreen mode Exit fullscreen mode

Explicitly repeatable command

Execute again
=> intentionally create another operation
Enter fullscreen mode Exit fullscreen mode

The dangerous answer is:

Execute again
=> maybe duplicate something, depends on timing
Enter fullscreen mode Exit fullscreen mode

What I now test by default

For consequential features, I increasingly start with hostile questions.

What if the request arrives twice?

What if two workers run this simultaneously?

What if the external API succeeds but our request times out?

What if the webhook is redelivered?

What if the accounting period closes while the form is open?

What if the same idempotency key is reused with different data?

What if the user belongs to another branch?

What if the client modifies the fee?

What if stale local state disagrees with a newer confirmed event?

What if the correction command is executed twice?

What if the data exists but the UI hides it?

These are not exotic edge cases.

They are normal production behavior.


A practical checklist

When building something that affects money or operational truth, this is the checklist I increasingly reach for:

[ ] Is there a stable identity for this operation?

[ ] Can the same request safely execute twice?

[ ] Is uniqueness enforced in the database?

[ ] Are important state checks repeated at commit time?

[ ] Are concurrent writes considered?

[ ] Are locks acquired consistently?

[ ] Can historical financial data be edited?

[ ] Should correction be represented as a reversal instead?

[ ] Is the server authoritative for financial calculations?

[ ] Are money values stored using exact minor units?

[ ] Can asynchronous events be delivered more than once?

[ ] Is UNKNOWN represented separately from FAILED?

[ ] Is there a reconciliation path?

[ ] Can stale state override fresher evidence?

[ ] Are permissions checked server-side?

[ ] Is object, branch, or account scope enforced?

[ ] Can the system explain how the current balance was produced?

[ ] Are user-visible properties tested in the browser?
Enter fullscreen mode Exit fullscreen mode

If several answers are uncomfortable, that is usually where the interesting engineering work begins.


The real feature is preserving truth

I used to think production hardening happened after the feature.

Build the feature.

Then make it reliable.

I increasingly think that is the wrong mental model.

For consequential systems, reliability is part of what the feature means.

A transfer that succeeds twice is not a working transfer feature.

A journal entry that can be posted into a closed period is not a working journal feature.

A balance that cannot explain its own history is not a trustworthy balance.

A reversal that rewrites history instead of recording correction is not a complete accounting workflow.

A permission system that only hides buttons is not authorization.

The implementation is not finished simply because the happy path works.

It is finished when the system can survive:

retries
timeouts
concurrency
duplicate delivery
stale state
partial failure
correction
unauthorized requests
Enter fullscreen mode Exit fullscreen mode

without inventing a new version of the truth.

That is the part of production engineering I find increasingly interesting.

Not just making the software work.

Making sure it continues telling the truth when everything around it gets messy.

Top comments (0)