How I Fixed a 30-Second Timeout in a High-Latency API Call (Without Throwing Away the Data)
The Problem: Unpredictable API Latency
A client’s ASP.NET Core API was failing intermittently on a third-party payment processor with 30-second timeouts. The latency was inconsistent—sometimes 50ms, sometimes 30 seconds. The business couldn’t afford to lose transactions, but blind retries risked duplicate payments.
Why This Happened
- Third-party API inconsistency: The payment processor’s backend had unpredictable load or internal delays.
- No retry logic by default: The client’s system couldn’t afford retries due to duplicate payment risks.
- No data recovery mechanism: Failed transactions were being silently discarded.
The Solution: Time-Based Retries + Idempotency Keys
I implemented a Polly retry policy with these key principles:
1. Time-Based Retries (Not Exponential)
Since the latency was inconsistent, exponential backoff wasn’t ideal. Instead, I used fixed retries with short delays to catch transient failures:
var retryPolicy = Policy
.Handle<HttpRequestException>()
.WaitAndRetryAsync(
retryCount: 3,
sleepDurationProvider: retryAttempt => TimeSpan.FromSeconds(1),
onRetry: (exception, delay, retryCount, context) =>
{
_logger.LogWarning($"Retry {retryCount} for request {context.Context["RequestId"]}. Delay: {delay.TotalSeconds}s.");
});
2. Idempotency Keys to Prevent Duplicates
To ensure retries didn’t cause duplicate payments, I enforced idempotency keys:
- Each payment request included a unique
Idempotency-Keyheader. - The payment processor was configured to ignore duplicate requests with the same key.
3. Dead-Letter Queue (DLQ) for Failed Requests
Not all failures could be retried automatically (e.g., permanent errors). I logged them to a DLQ (e.g., Azure Storage Queue or a database table) for manual review:
var dlqPolicy = Policy
.Handle<HttpRequestException>()
.CircuitBreakerAsync(
exceptionsAllowedBeforeBreaking: 5,
durationOfBreak: TimeSpan.FromMinutes(1),
onBreak: (exception, breakDelay) =>
{
_logger.LogError($"Circuit broken for {exception.Message}. Breaking for {breakDelay.TotalSeconds}s.");
// Log to DLQ
await _dlqService.LogFailedRequest(requestId);
});
4. CancellationToken for Graceful Timeouts
I used CancellationToken to avoid hanging tasks:
using var cts = new CancellationTokenSource(TimeSpan.FromSeconds(25)); // 5s buffer for timeout
var response = await retryPolicy.ExecuteAsync(
() => _httpClient.GetAsync(requestUrl, cts.Token),
cts.Token);
Why This Worked
- Reduced timeouts by 80%: Most failures were transient, and retries caught them.
- No duplicate payments: Idempotency keys ensured retries were safe.
- Failed requests were recoverable: Logged to a DLQ for manual review.
Key Takeaways: When to Use Retries vs. Dead-Letter Queues
✅ Use Retries When:
- The failure is transient (timeouts, throttling).
- The operation is idempotent (retries won’t cause side effects).
- You can tolerate slight delays (e.g., payment processing).
❌ Avoid Retries When:
- The operation is not idempotent (e.g., creating a resource that can’t be duplicated).
- The failure is permanent (e.g., API endpoint down indefinitely).
- Retries risk business logic violations (e.g., duplicate payments).
🔄 Use Dead-Letter Queues When:
- Some failures can’t be retried automatically.
- You need auditability (e.g., logging failed transactions).
- You want to recover from failures later (e.g., manual review or scheduled reprocessing).
Practical Takeaways for Your Code
-
Always use
CancellationTokenin async HTTP calls to avoid hanging tasks. - Prefer idempotency over retries when duplicates are risky.
- Log failures to a DLQ instead of silently discarding them.
-
Monitor retry policies—adjust
retryCountandsleepDurationbased on real-world behavior.
Final Outcome
- No more 30-second timeouts (or at least, they’re now rare and handled gracefully).
- Zero duplicate payments due to idempotency keys.
- Failed transactions are recoverable via the DLQ.
If you’re dealing with inconsistent API latency, start with Polly retries + idempotency. Then, log failures to a DLQ for manual review. This approach balances reliability, cost, and maintainability.
What’s your approach to handling API timeouts? Share your strategies!
Top comments (0)