DEV Community

Cover image for Why My Background Job Processor Deadlocked Under Real Load (And How CancellationToken Saved It)
Imran Ahmed
Imran Ahmed

Posted on

Why My Background Job Processor Deadlocked Under Real Load (And How CancellationToken Saved It)

Why My Background Job Processor Deadlocked Under Real Load (And How CancellationToken Saved It)

The Setup

A recurring inventory-sync job ran nightly in production, updating stock levels across multiple services. In development, it worked flawlessly — single-threaded, no contention. In production, it froze entirely. Monitoring showed blocked EF Core queries piling up.

The Root Cause: A Deadlock Born From Lock Mismatches

Two independent background services were responsible for updating inventory data:

  • One service synced orders into inventory records.
  • Another synced supplier data into the same set of inventory rows.

Both used EF Core transactions. But critically, they did not acquire row locks in the same order. When both tried to update overlapping rows simultaneously, SQL Server detected a deadlock and chose one as the victim — repeatedly failing whichever job happened to lose.

This wasn’t a transient failure — it was a structural problem caused by inconsistent lock ordering.

Why Retries Made It Worse

Initially, the system retried failed jobs automatically. But retries only amplified contention. Two jobs would collide again, deadlock again, fail again — burning CPU and extending downtime.

The Fix: Cancellation, Not Retries

Instead of retrying indefinitely, I introduced a shared CancellationToken that flowed through both services. When one job detected it was losing a race — either via timeout or explicit coordination — it would cancel its own operation gracefully.

public async Task SyncInventoryAsync(CancellationToken ct)
{
    await foreach (var batch in GetBatches(ct))
    {
        ct.ThrowIfCancellationRequested();
        await UpdateBatch(batch, ct);
    }
}
Enter fullscreen mode Exit fullscreen mode

This way, the second job didn’t wait on a lock forever — it exited cleanly, leaving room for the winner to finish.

Making Cancellation Safe: Idempotent Upserts

Cancellation meant interruption — but interruptions aren’t safe unless operations can resume without side effects. To ensure that, every write operation became an idempotent upsert, keyed uniquely on item identifiers.

context.InventoryItems.Upsert(batch);
Enter fullscreen mode Exit fullscreen mode

This ensured re-running or resuming a partial job wouldn’t duplicate entries or overwrite newer values.

Key Takeaways

  1. Deadlocks often come from inconsistent lock ordering, especially under concurrent background loads.
  2. Don’t retry blindly — retries can worsen contention and delay resolution.
  3. Use cancellation tokens to enable graceful shutdown/resumption during contention.
  4. Idempotency ensures reliability when jobs may be interrupted and resumed.
  5. Background jobs aren’t fire-and-forget — they must be designed for concurrency from day one.

Top comments (0)