When Your Background Job Retries Are Silently Stepping on Each Other
In .NET backend development, we rely heavily on background jobs for asynchronous processing: sending emails, generating reports, or processing orders. We use libraries like Hangfire, Quartz, or custom IHostedService implementations to ensure resilience.
A common pattern is to wrap database operations in a try-catch block that catches transient exceptions (like SqlException for deadlocks or timeouts) and retries the operation. The assumption is usually that the code is idempotent—running it twice shouldn’t change the result.
But what happens when "idempotent" isn't quite true? What happens when two different jobs, retrying due to transient failures, both attempt to modify the same shared resource?
The Silent Overwrite
Let’s look at a scenario. We have an OrderService that processes customer orders. When an order is placed, a background job is triggered to update the Inventory table.
Suppose two orders are placed almost simultaneously:
- Order 1 for Product X (Quantity 5).
- Order 2 for Product X (Quantity 3).
Both jobs start almost at the same time.
- Job 1 reads
Inventoryrow for Product X. Current stock: 10. - Job 2 reads
Inventoryrow for Product X. Current stock: 10. - Job 1 calculates new stock: 10 - 5 = 5.
- Job 2 calculates new stock: 10 - 3 = 7.
- Job 1 tries to write
5to the DB. It hits a transient network blip and the write fails. - The retry mechanism kicks in for Job 1. It re-executes the logic. It reads the inventory again? No. In many implementations, the entity was already loaded into memory. The retry just re-sends the
UPDATEstatement with the stale value it calculated earlier (5). - Job 2 successfully writes
7to the DB. - Job 1’s retry finally succeeds and writes
5to the DB.
Result: The stock is now 5. But it should be 2 (10 - 5 - 3). Job 1’s retry silently overwrote Job 2’s valid change. This is a "lost update" caused by retry mechanics stepping on each other.
Why Tests Miss This
Standard unit tests usually mock the database or test in isolation. You might test:
- "Given stock 10, order 5, result is 5." (Pass)
- "Given stock 10, order 3, result is 7." (Pass)
But you rarely test: "Given two concurrent jobs with a transient failure on the first, does the second job’s data survive the first job’s retry?"
Simulating this in integration tests is difficult. You’d need to inject a delay or a specific exception during the save of the first job while the second job is proceeding. Most developers skip this complexity, leaving the bug to production.
The Fix: Optimistic Concurrency
The robust solution is to use optimistic concurrency control. Instead of blindly writing, we ask the database: "Did anything else change this row while I was away?"
In SQL Server, we use the rowversion data type. It’s a system-generated number that increments every time the row is updated.
Step 1: Add a Row Version Column
In your EF Core model:
public class InventoryItem
{
public int Id { get; set; }
public string ProductName { get; set; }
public int Stock { get; set; }
// Use RowVersion or Timestamp
[Timestamp]
public byte[] RowVersion { get; set; }
}
In your SQL Server database, ensure you have a rowversion column. If you’re using SQL Server, you can just add ROWVERSION to the table definition.
Step 2: Implement Idempotent Retry Guards
When you retry a job, you must ensure that the state you’re writing is still valid. With rowversion, EF Core automatically adds a WHERE clause to your UPDATE statements.
// In your service
public async Task ProcessOrderAsync(int orderId, int quantity)
{
using var context = new AppDbContext();
var item = await context.InventoryItems.FindAsync(productId);
if (item == null) return;
item.Stock -= quantity;
try
{
await context.SaveChangesAsync();
}
catch (DbUpdateConcurrencyException ex)
{
// This is thrown if the RowVersion has changed.
// It means another job updated this row.
// Instead of overwriting, we should fetch the latest state and re-calculate.
// Or, in the context of a simple deduction, we might just let it fail
// and let the queue re-enqueue the job.
var entry = ex.Entries.Single();
var databaseValues = await entry.OriginalValues.StoreGeneratedValuesAsync();
// Logic to handle the conflict.
// For inventory, you might need to re-read and check if the stock
// is still sufficient, or if the order has already been processed.
throw; // Let the retry mechanism handle it, or log and fail fast.
}
}
By catching DbUpdateConcurrencyException, you prevent the silent overwrite. If Job 1 retries and finds that Job 2 has already updated the rowversion, the UPDATE affects zero rows. The exception is thrown, and you know the state is stale.
Step 3: Make the Business Logic Idempotent
Retries alone aren’t enough. Your code must handle the "already processed" state. If Job 1 is retried and the rowversion matches (meaning no one else touched it), it’s safe to re-apply. But if it doesn’t match, you must check: "Has this specific order already been deducted?"
Often, this means adding a ProcessedOrderIds table or a flag on the order itself, and checking it before performing the calculation.
Practical Takeaway
Concurrency bugs are not about your code being "wrong" in the happy path. They’re about your code failing to handle the overlap between retries and other active writers.
- Never assume that a retry after a transient error is safe without checking the current database state.
- Use
rowversion(or similar mechanisms likeupdated_atwith application-level checks) for any shared mutable resources. - Design for conflict. When a concurrency exception is thrown, your job should either re-read the state and retry the business logic, or fail fast so the queue can re-process it with fresh data.
Your background jobs should be resilient to failures, but they shouldn’t be destructive to data integrity.
Top comments (0)