DEV Community

Said Olano
Said Olano

Posted on

Circuit Breaker Pattern: Preventing Cascading Failures

Circuit Breaker Pattern: Prevent Cascading Failures in Distributed Systems

Introduction

Your microservice calls an external API. The API is slow. Your service waits... and waits... and waits.

Meanwhile, your thread pool is exhausted. New requests pile up. Your entire service becomes unresponsive.

This is the cascading failure problem, and it's one of the most dangerous patterns in distributed systems.

The Circuit Breaker pattern is your safety net.

It's a simple but powerful mechanism that prevents your application from making requests to failing services, giving them time to recover while protecting your own resources.

What is the Circuit Breaker Pattern?

A Circuit Breaker monitors the health of calls to a remote service. When the failure rate exceeds a threshold, it "trips" and stops sending requests to that service, returning errors immediately without attempting the call.

Think of it like an electrical circuit breaker in your home: when there's too much current (failures), it flips open and stops the flow (requests) before damage occurs.

Three States

CLOSED (Normal)
    ↓
  [failures exceed threshold]
    ↓
OPEN (Failing)
    ↓
  [after timeout period]
    ↓
HALF-OPEN (Testing)
    ↓
  [test request succeeds]
    ↓
CLOSED (Recovery)
Enter fullscreen mode Exit fullscreen mode

CLOSED State:

  • Circuit is functioning normally
  • All requests pass through to the service
  • Failures are counted
  • When failures exceed threshold → transition to OPEN

OPEN State:

  • Circuit has tripped
  • All requests immediately return error (fail-fast)
  • No requests sent to the failing service
  • Service gets time to recover
  • After timeout → transition to HALF-OPEN

HALF-OPEN State:

  • Testing if service has recovered
  • Limited requests allowed (usually 1)
  • If request succeeds → transition to CLOSED
  • If request fails → transition back to OPEN

Failure Detection Mechanisms

1. Failure Count Threshold

circuit_breaker:
  failure_threshold: 5  # Open after 5 failures
  window_duration: 60s  # Count failures in 60s window
Enter fullscreen mode Exit fullscreen mode

Simple approach: count consecutive failures, trip after threshold.

Pros:

  • Easy to understand and implement
  • Low computational overhead

Cons:

  • Sensitive to temporary blips
  • Doesn't account for success ratio

2. Failure Rate Threshold

circuit_breaker:
  failure_rate_threshold: 50%  # Open if 50% fail
  window_duration: 60s
  minimum_requests: 10  # Need at least 10 requests to measure
Enter fullscreen mode Exit fullscreen mode

More sophisticated: trip if failure rate exceeds threshold over window.

Example:

Last 60 seconds:
- 20 requests total
- 12 failures
- Failure rate: 60% (exceeds 50% threshold)
- Action: Trip circuit
Enter fullscreen mode Exit fullscreen mode

Pros:

  • Accounts for overall health
  • More stable than count-based

Cons:

  • Requires statistical tracking
  • Need minimum request volume for accuracy

3. Response Time Threshold

circuit_breaker:
  slow_call_duration_threshold: 2000ms
  slow_call_rate_threshold: 50%
Enter fullscreen mode Exit fullscreen mode

Trip if too many calls are slow (timeout).

Pros:

  • Catches degraded performance
  • Prevents resource exhaustion

Cons:

  • Requires latency percentile tracking
  • May need tuning per service

4. Custom Health Check

circuit_breaker:
  health_check:
    url: https://service/health
    interval: 5s
    expected_status: 200
Enter fullscreen mode Exit fullscreen mode

Actively probe service health.

Pros:

  • Most accurate health detection
  • Can provide detailed diagnostics

Cons:

  • Additional network overhead
  • Complexity in implementation

Real-World Implementation

Basic Circuit Breaker (Java)

public class CircuitBreaker {
    private State state = State.CLOSED;
    private int failureCount = 0;
    private int successCount = 0;
    private long lastFailureTime = 0;

    private static final int FAILURE_THRESHOLD = 5;
    private static final int TIMEOUT = 60000; // 60 seconds
    private static final int SUCCESS_THRESHOLD = 2;

    public <T> T execute(Supplier<T> operation) {
        if (state == State.OPEN) {
            if (System.currentTimeMillis() - lastFailureTime > TIMEOUT) {
                state = State.HALF_OPEN;
                successCount = 0;
            } else {
                throw new CircuitBreakerOpenException("Circuit breaker is OPEN");
            }
        }

        try {
            T result = operation.get();
            onSuccess();
            return result;
        } catch (Exception e) {
            onFailure();
            throw e;
        }
    }

    private void onSuccess() {
        failureCount = 0;

        if (state == State.HALF_OPEN) {
            successCount++;
            if (successCount >= SUCCESS_THRESHOLD) {
                state = State.CLOSED;
                successCount = 0;
            }
        }
    }

    private void onFailure() {
        lastFailureTime = System.currentTimeMillis();
        failureCount++;

        if (state == State.HALF_OPEN) {
            state = State.OPEN;
        } else if (failureCount >= FAILURE_THRESHOLD) {
            state = State.OPEN;
        }
    }

    enum State { CLOSED, OPEN, HALF_OPEN }
}
Enter fullscreen mode Exit fullscreen mode

Usage

CircuitBreaker breaker = new CircuitBreaker();

public User getUser(String userId) {
    return breaker.execute(() -> 
        userService.fetchUser(userId)
    );
}
Enter fullscreen mode Exit fullscreen mode

Integration with Retry and Fallback

Circuit Breaker works best with other patterns:

Request
  ↓
Circuit Breaker (First line of defense)
  ├─ CLOSED: Forward to service
  │    ↓
  │  Retry (Second line: retry on transient failures)
  │    ├─ Retry failed
  │    └─ Fallback (Third line: graceful degradation)
  │
  └─ OPEN: Skip service entirely
       ↓
     Fallback (Return cached/default data)
Enter fullscreen mode Exit fullscreen mode

Complete Pattern

public User getUser(String userId) {
    try {
        return circuitBreaker.execute(() -> {
            try {
                return userService.fetchUser(userId);
            } catch (TemporaryException e) {
                throw e; // Allow retry
            } catch (PermanentException e) {
                throw new CircuitBreakerException(e); // Trip circuit
            }
        });
    } catch (CircuitBreakerOpenException | TemporaryException e) {
        // Fallback: return cached user
        return cache.get(userId)
            .orElse(new User(userId, "Unknown", "cached"));
    }
}
Enter fullscreen mode Exit fullscreen mode

Metrics & Monitoring

Track these metrics for each circuit breaker:

Primary Metrics:
├─ state (CLOSED, OPEN, HALF_OPEN)
├─ success_count (requests that succeeded)
├─ failure_count (requests that failed)
├─ rejection_count (requests rejected by open circuit)
├─ total_time (total time spent in each state)
└─ last_failure_time (when last failure occurred)

Derived Metrics:
├─ success_rate = success / (success + failure)
├─ failure_rate = failure / (success + failure)
├─ rejection_rate = rejection / (success + failure + rejection)
└─ time_in_open = time_in_open_state / total_time
Enter fullscreen mode Exit fullscreen mode

Example Dashboard

User Service Circuit Breaker
├─ State: CLOSED ✅
├─ Success Rate: 99.5%
├─ Failure Rate: 0.5% (1/200)
├─ Last Failure: 2 minutes ago
├─ Avg Response Time: 45ms
└─ Total Requests: 200

Order Service Circuit Breaker
├─ State: OPEN ⚠️
├─ Success Rate: 20%
├─ Failure Rate: 80% (16/20)
├─ Last Failure: 15 seconds ago
├─ Avg Response Time: 5000ms (timeout)
├─ Time Until Recovery: 45 seconds
└─ Total Requests: 20
Enter fullscreen mode Exit fullscreen mode

Configuration Best Practices

1. Per-Service Configuration

Different services have different characteristics:

circuit_breakers:
  user_service:
    failure_threshold: 5
    timeout: 30s

  payment_service:
    failure_threshold: 2  # More aggressive
    timeout: 10s

  recommendation_service:
    failure_threshold: 10  # More lenient
    timeout: 60s
Enter fullscreen mode Exit fullscreen mode

2. Tuning Parameters

Parameter Purpose Typical Value
Failure Threshold When to trip 5-10 failures
Timeout Duration Recovery wait 30-60 seconds
Success Threshold When to close from HALF-OPEN 1-2 successes
Window Size Measurement window 60 seconds
Minimum Requests Before calculating rate 10-20

3. Avoid These Mistakes

❌ Too Aggressive:

failure_threshold: 1  # Trip after 1 failure
timeout: 1s
Enter fullscreen mode Exit fullscreen mode

→ Trips on every temporary network hiccup

✅ Balanced:

failure_threshold: 5
timeout: 30s
minimum_requests: 20
failure_rate: 50%
Enter fullscreen mode Exit fullscreen mode

❌ Too Lenient:

failure_threshold: 100
timeout: 300s
Enter fullscreen mode Exit fullscreen mode

→ Doesn't protect system from cascading failures

Popular Circuit Breaker Libraries

Java

Library Features Best For
Resilience4j Lightweight, functional New projects
Hystrix Feature-rich, battle-tested Production systems
Polly Comprehensive (for .NET) .NET ecosystem

Example with Resilience4j

CircuitBreakerRegistry registry = CircuitBreakerRegistry.ofDefaults();

CircuitBreaker breaker = registry.circuitBreaker("userService",
    CircuitBreakerConfig.custom()
        .failureRateThreshold(50.0f)
        .waitDurationInOpenState(Duration.ofSeconds(30))
        .slowCallRateThreshold(50.0f)
        .slowCallDurationThreshold(Duration.ofSeconds(2))
        .minimumNumberOfCalls(10)
        .build());

userService = Decorators.ofSupplier(() -> service.getUser(id))
    .withCircuitBreaker(breaker)
    .withRetry(Retry.ofDefaults("userService"))
    .withFallback(List.of(
        new TimeoutException(),
        new CircuitBreakerOpenException()),
        e -> cachedUser)
    .decorate();
Enter fullscreen mode Exit fullscreen mode

Cascading Failures: Before & After

Without Circuit Breaker

User Service 99% availability
Order Service 99% availability
Payment Service 99% availability
Recommendation Service 99% availability

Combined uptime:
0.99 × 0.99 × 0.99 × 0.99 = 96.06%

Single Order Service failure:
├─ User Service waits for Order Service timeout (30s)
├─ Orders pile up, thread pool exhausted
├─ User Service becomes unresponsive
└─ Entire system cascade fails
    (Total downtime: hours)
Enter fullscreen mode Exit fullscreen mode

With Circuit Breaker

Order Service fails
    ↓
After 5 failures, circuit trips (2-3 seconds)
    ↓
Circuit Breaker opens
    ↓
Subsequent requests fail immediately (no timeout wait)
    ↓
Thread pool recovers
    ↓
User Service continues to serve requests
    ↓
After 30 seconds, retry Order Service (HALF-OPEN)
    ↓
Order Service recovered
    ↓
Circuit closes
    ↓
Full system recovery (2-3 minutes)
Enter fullscreen mode Exit fullscreen mode

Advanced Patterns

1. Bulkhead Isolation

Separate thread pools per service prevent one failure from affecting others:

Thread Pool A (User Service)
Thread Pool B (Order Service)
Thread Pool C (Payment Service)

Order Service fails:
├─ Pool B exhausted
└─ Pools A & C continue normally
Enter fullscreen mode Exit fullscreen mode

2. Adaptive Thresholds

Adjust parameters based on time of day:

circuit_breaker:
  peak_hours: (9:00 - 17:00)
    failure_threshold: 10
    timeout: 60s

  off_peak: (17:00 - 9:00)
    failure_threshold: 5
    timeout: 30s
Enter fullscreen mode Exit fullscreen mode

3. Service Mesh Integration

Tools like Istio implement Circuit Breaker at infrastructure level:

apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
  name: order-service
spec:
  host: order-service
  trafficPolicy:
    outlierDetection:
      consecutiveErrors: 5
      interval: 30s
      baseEjectionTime: 30s
      maxEjectionPercent: 50
Enter fullscreen mode Exit fullscreen mode

Anti-Patterns to Avoid

1. Silent Failures

❌ Bad:

try {
    return circuitBreaker.execute(() -> service.call());
} catch (CircuitBreakerOpenException e) {
    return null; // Silently fail
}
Enter fullscreen mode Exit fullscreen mode

✅ Good:

try {
    return circuitBreaker.execute(() -> service.call());
} catch (CircuitBreakerOpenException e) {
    log.warn("Circuit open for service", e);
    return getFallbackValue();
}
Enter fullscreen mode Exit fullscreen mode

2. Ignoring the State

❌ Bad:

// No monitoring
circuitBreaker.execute(operation);
Enter fullscreen mode Exit fullscreen mode

✅ Good:

// Monitor state transitions
circuitBreaker.onOpen(state -> 
    metrics.increment("circuit.open", 
        Tags.of("service", "order"))
);
circuitBreaker.onClose(state -> 
    metrics.increment("circuit.close")
);
Enter fullscreen mode Exit fullscreen mode

3. Generic Configuration

❌ Bad:

# One size fits all
global_circuit_breaker:
  threshold: 5
  timeout: 30s
Enter fullscreen mode Exit fullscreen mode

✅ Good:

# Per-service tuning
services:
  critical_payment:
    threshold: 2
    timeout: 10s
  cache_layer:
    threshold: 20
    timeout: 60s
Enter fullscreen mode Exit fullscreen mode

Conclusion

The Circuit Breaker pattern is essential for building resilient distributed systems.

It provides:
✅ Fail-fast: Errors return immediately instead of timing out
✅ Resource protection: Prevents thread pool exhaustion
✅ Recovery time: Gives failing services time to heal
✅ Cascading failure prevention: Stops failure propagation
✅ Observability: Metrics on service health

A well-configured Circuit Breaker can mean the difference between a brief service hiccup and a system-wide outage.

Implement it today, tune it carefully, and monitor it relentlessly.

Top comments (0)