DEV Community

Cover image for Building ResilienceLab: retries, circuit breakers, and tracing in .NET 10
Leopoldo Benavente Cadena
Leopoldo Benavente Cadena

Posted on AI-assisted

Building ResilienceLab: retries, circuit breakers, and tracing in .NET 10

I wanted to build something useful and ended up creating ResilienceLab 0.5.0, an experimental resilience library for .NET 10 with its own API and an MIT license. I developed it with coding agents, reviewing the contracts and checking the results.

It includes retries for exceptions and generic results, cooperative timeouts, circuit breakers, fallback, rate limiting, hedging, and an HTTP adapter integrated with IHttpClientFactory. Optional observability uses logs, metrics, and ActivitySource traces.

Strategy order changes the behavior

In the catalog API, I composed these layers from outside to inside:

Fallback → Circuit breaker → Retry → Timeout → HTTP
Enter fullscreen mode Exit fullscreen mode

The circuit counts one failed call after its retries are exhausted, while each attempt has its own timeout. Two additional retries allow up to three HTTP requests. The HTTP adapter has its retry budget disabled to avoid multiplying attempts.

Fallback returns only the last valid catalog snapshot, within a bounded age. If the upstream service fails and there is no usable cache, the API returns 503. It does not invent a successful response.

What I verified

The library has 194 passing Release tests. The consumer API adds 12 reproducible HTTP scenarios, covering recovery, timeouts, circuit opening and recovery, cold and expired caches, and consumer cancellation without pending upstream handlers. I also made a real external request and verified that an OpenTelemetry Collector received traces and metrics over OTLP.

These are bounded checks, not proof of production performance or long-term stability. A cooperative timeout requests cancellation and waits for the operation to finish. Hedging cancels and awaits losing attempts, which can delay completion if they ignore cancellation and can increase upstream load. The library is experimental and does not offer Polly API compatibility.

Try it

dotnet add package ResilienceLab --version 0.5.0
dotnet add package ResilienceLab.Http --version 0.5.0
Enter fullscreen mode Exit fullscreen mode

Create a .NET 10 console app, install the core package, and replace Program.cs with this example:

using System.Net.Http;
using ResilienceLab;

var attempts = 0;
var result = await Retry.ExecuteAsync(
    ct =>
    {
        ct.ThrowIfCancellationRequested();
        attempts++;
        return attempts < 3
            ? Task.FromException<string>(new HttpRequestException())
            : Task.FromResult("Recovered");
    },
    new RetryOptions
    {
        MaxRetries = 2,
        BaseDelay = TimeSpan.FromMilliseconds(100),
        MaxDelay = TimeSpan.FromSeconds(1),
        ShouldRetry = ex => ex is HttpRequestException
    });

Console.WriteLine($"{result}; attempts: {attempts}");
Enter fullscreen mode Exit fullscreen mode

This simulates two transient failures and prints Recovered; attempts: 3. The default full jitter randomizes the bounded delay; cancellation is never retried.

Which retry, cancellation, or fallback behavior has been hardest to verify in your applications?

Top comments (0)