Retry Storms: How a 502 Fix Created a Thousand Duplicate Orders
In modern distributed systems, transient failures are inevitable. Timeouts, circuit breakers, and retries are standard defense mechanisms. But when idempotency—the guarantee that processing the same request multiple times has the same effect as processing it once—is skipped, retries turn into a destructive force.
This isn’t a theoretical scenario. It’s a real incident that unfolded in a production e-commerce system, where a single upstream timeout cascaded into thousands of duplicate orders.
The Incident
The system in question was an ASP.NET Core backend handling order processing for an online retailer. Orders flowed through a REST API endpoint into a SQL Server database. To improve reliability, the payment gateway upstream implemented automatic retries on transient failures like HTTP 502s.
Everything worked fine under normal conditions. Then came a brief network blip—an upstream timeout caused the first attempt to fail with a 502. The gateway retried.
But the retry mechanism had no notion of idempotency. Each retry was treated as a brand-new request. From the backend’s perspective, it was as if several customers simultaneously submitted identical orders.
Within minutes, hundreds of duplicate order records appeared in the database. The issue compounded as more gateway retries fired, ultimately resulting in thousands of unintended duplicate entries. Customer service lines lit up.
Root Cause Analysis
The root cause wasn't the retry logic itself. Retries are essential for resilience. The problem was the absence of idempotency enforcement—there was nothing preventing identical requests from producing multiple side effects (i.e., inserting multiple rows).
Without idempotency:
- A retry isn’t “trying again”—it’s “trying fresh.”
- Downstream systems treat every request as independent.
- Side effects multiply uncontrollably.
The Fix: Lightweight ASP.NET Core Idempotency Middleware
Rather than modifying all endpoints or disabling retries (which would hurt reliability), we introduced a global idempotency middleware that enforces idempotency at the HTTP layer.
Here’s the core idea:
- Clients must include an
Idempotency-Keyheader with each request. - On receiving the request, compute a deterministic key (e.g., hash of
{TenantId}:{IdempotencyKey}). - Check a fast cache (Redis) for existing results:
- If found → return cached result immediately.
- If not found → proceed and cache outcome (success/failure) for future retries.
- Enforce this behavior only on stateful operations (
POST,PUT, etc.).
This approach ensures:
- Same operation + same key → same result.
- Retries don’t trigger redundant work.
- Failures still propagate correctly.
Basic Implementation Sketch
public class IdempotencyMiddleware
{
private readonly RequestDelegate _next;
public IdempotencyMiddleware(RequestDelegate next)
{
_next = next;
}
public async Task InvokeAsync(HttpContext context, IIdempotencyStore store)
{
if (!IsStatefulRequest(context) || !HasIdempotencyKey(context))
{
await _next(context);
return;
}
var key = ComputeDeterministicKey(context);
var cachedResult = await store.GetAsync(key);
if (cachedResult != null)
{
await ApplyCachedResponse(context, cachedResult);
return;
}
// Capture response stream to intercept and cache
using var buffer = new MemoryStream();
var originalBodyStream = context.Response.Body;
context.Response.Body = buffer;
await _next(context);
buffer.Seek(0, SeekOrigin.Begin);
var responseBytes = buffer.ToArray();
var statusCode = context.Response.StatusCode;
await store.SetAsync(key, new CachedResponse(statusCode, responseBytes));
await buffer.CopyAsync(originalBodyStream);
}
private static bool IsStatefulRequest(HttpContext ctx) =>
ctx.Request.Method == HttpMethods.Post || ctx.Request.Method == HttpMethods.Put;
private static bool HasIdempotencyKey(HttpContext ctx) =>
!string.IsNullOrEmpty(ctx.Request.Headers["Idempotency-Key"]);
private static string ComputeDeterministicKey(HttpContext ctx)
{
var rawKey = ctx.Request.Headers["Idempotency-Key"];
var tenantId = ctx.User?.FindFirst("tenant_id")?.Value ?? "default";
using var sha = SHA256.Create();
var input = $"{tenantId}:{rawKey}";
var bytes = Encoding.UTF8.GetBytes(input);
var hashBytes = sha.ComputeHash(bytes);
return Convert.ToBase64String(hashBytes); // Deterministic cache key
}
private static async Task ApplyCachedResponse(HttpContext context, CachedResponse cached)
{
context.Response.StatusCode = cached.StatusCode;
context.Response.ContentLength = cached.Body?.Length;
await context.Response.Body.WriteAsync(cached.Body ?? Array.Empty<byte>());
}
}
Use Redis or an in-memory dictionary backed by persistence for production-grade scenarios.
Ponytail note: In-memory storage works locally but lacks durability. For multi-instance deployments, Redis or another shared cache is required.
Practical Takeaway
Every HTTP-based integration assumes transient failures will happen—but few plan for repeated attempts safely. Idempotency protects against unintended duplication and builds confidence in retry-heavy architectures.
Implement it early, ideally through middleware so it applies globally without touching business logic. With minimal overhead and clear semantics, it becomes the “seatbelt” for HTTP integrations.
Don’t wait until your logs fill up with duplicate orders—and your customers start asking why they were billed five times.
Summary Checklist
✅ Require Idempotency-Key on all mutating requests
✅ Use deterministic key derivation based on user tenant/context
✅ Cache final responses (including status codes) for reuse
✅ Store results temporarily (Redis TTL) to avoid unbounded growth
✅ Short-circuit known keys before hitting application logic
Start small—just one endpoint—and grow from there. Once adopted, idempotency removes entire classes of subtle bugs and operational headaches.
Top comments (0)