Circuit Breaker Pattern: Prevent Cascading Failures in Distributed Systems
Introduction
Your microservice calls an external API. The API is slow. Your service waits... and waits... and waits.
Meanwhile, your thread pool is exhausted. New requests pile up. Your entire service becomes unresponsive.
This is the cascading failure problem, and it's one of the most dangerous patterns in distributed systems.
The Circuit Breaker pattern is your safety net.
It's a simple but powerful mechanism that prevents your application from making requests to failing services, giving them time to recover while protecting your own resources.
What is the Circuit Breaker Pattern?
A Circuit Breaker monitors the health of calls to a remote service. When the failure rate exceeds a threshold, it "trips" and stops sending requests to that service, returning errors immediately without attempting the call.
Think of it like an electrical circuit breaker in your home: when there's too much current (failures), it flips open and stops the flow (requests) before damage occurs.
Three States
CLOSED (Normal)
↓
[failures exceed threshold]
↓
OPEN (Failing)
↓
[after timeout period]
↓
HALF-OPEN (Testing)
↓
[test request succeeds]
↓
CLOSED (Recovery)
CLOSED State:
- Circuit is functioning normally
- All requests pass through to the service
- Failures are counted
- When failures exceed threshold → transition to OPEN
OPEN State:
- Circuit has tripped
- All requests immediately return error (fail-fast)
- No requests sent to the failing service
- Service gets time to recover
- After timeout → transition to HALF-OPEN
HALF-OPEN State:
- Testing if service has recovered
- Limited requests allowed (usually 1)
- If request succeeds → transition to CLOSED
- If request fails → transition back to OPEN
Failure Detection Mechanisms
1. Failure Count Threshold
circuit_breaker:
failure_threshold: 5 # Open after 5 failures
window_duration: 60s # Count failures in 60s window
Simple approach: count consecutive failures, trip after threshold.
Pros:
- Easy to understand and implement
- Low computational overhead
Cons:
- Sensitive to temporary blips
- Doesn't account for success ratio
2. Failure Rate Threshold
circuit_breaker:
failure_rate_threshold: 50% # Open if 50% fail
window_duration: 60s
minimum_requests: 10 # Need at least 10 requests to measure
More sophisticated: trip if failure rate exceeds threshold over window.
Example:
Last 60 seconds:
- 20 requests total
- 12 failures
- Failure rate: 60% (exceeds 50% threshold)
- Action: Trip circuit
Pros:
- Accounts for overall health
- More stable than count-based
Cons:
- Requires statistical tracking
- Need minimum request volume for accuracy
3. Response Time Threshold
circuit_breaker:
slow_call_duration_threshold: 2000ms
slow_call_rate_threshold: 50%
Trip if too many calls are slow (timeout).
Pros:
- Catches degraded performance
- Prevents resource exhaustion
Cons:
- Requires latency percentile tracking
- May need tuning per service
4. Custom Health Check
circuit_breaker:
health_check:
url: https://service/health
interval: 5s
expected_status: 200
Actively probe service health.
Pros:
- Most accurate health detection
- Can provide detailed diagnostics
Cons:
- Additional network overhead
- Complexity in implementation
Real-World Implementation
Basic Circuit Breaker (Java)
public class CircuitBreaker {
private State state = State.CLOSED;
private int failureCount = 0;
private int successCount = 0;
private long lastFailureTime = 0;
private static final int FAILURE_THRESHOLD = 5;
private static final int TIMEOUT = 60000; // 60 seconds
private static final int SUCCESS_THRESHOLD = 2;
public <T> T execute(Supplier<T> operation) {
if (state == State.OPEN) {
if (System.currentTimeMillis() - lastFailureTime > TIMEOUT) {
state = State.HALF_OPEN;
successCount = 0;
} else {
throw new CircuitBreakerOpenException("Circuit breaker is OPEN");
}
}
try {
T result = operation.get();
onSuccess();
return result;
} catch (Exception e) {
onFailure();
throw e;
}
}
private void onSuccess() {
failureCount = 0;
if (state == State.HALF_OPEN) {
successCount++;
if (successCount >= SUCCESS_THRESHOLD) {
state = State.CLOSED;
successCount = 0;
}
}
}
private void onFailure() {
lastFailureTime = System.currentTimeMillis();
failureCount++;
if (state == State.HALF_OPEN) {
state = State.OPEN;
} else if (failureCount >= FAILURE_THRESHOLD) {
state = State.OPEN;
}
}
enum State { CLOSED, OPEN, HALF_OPEN }
}
Usage
CircuitBreaker breaker = new CircuitBreaker();
public User getUser(String userId) {
return breaker.execute(() ->
userService.fetchUser(userId)
);
}
Integration with Retry and Fallback
Circuit Breaker works best with other patterns:
Request
↓
Circuit Breaker (First line of defense)
├─ CLOSED: Forward to service
│ ↓
│ Retry (Second line: retry on transient failures)
│ ├─ Retry failed
│ └─ Fallback (Third line: graceful degradation)
│
└─ OPEN: Skip service entirely
↓
Fallback (Return cached/default data)
Complete Pattern
public User getUser(String userId) {
try {
return circuitBreaker.execute(() -> {
try {
return userService.fetchUser(userId);
} catch (TemporaryException e) {
throw e; // Allow retry
} catch (PermanentException e) {
throw new CircuitBreakerException(e); // Trip circuit
}
});
} catch (CircuitBreakerOpenException | TemporaryException e) {
// Fallback: return cached user
return cache.get(userId)
.orElse(new User(userId, "Unknown", "cached"));
}
}
Metrics & Monitoring
Track these metrics for each circuit breaker:
Primary Metrics:
├─ state (CLOSED, OPEN, HALF_OPEN)
├─ success_count (requests that succeeded)
├─ failure_count (requests that failed)
├─ rejection_count (requests rejected by open circuit)
├─ total_time (total time spent in each state)
└─ last_failure_time (when last failure occurred)
Derived Metrics:
├─ success_rate = success / (success + failure)
├─ failure_rate = failure / (success + failure)
├─ rejection_rate = rejection / (success + failure + rejection)
└─ time_in_open = time_in_open_state / total_time
Example Dashboard
User Service Circuit Breaker
├─ State: CLOSED ✅
├─ Success Rate: 99.5%
├─ Failure Rate: 0.5% (1/200)
├─ Last Failure: 2 minutes ago
├─ Avg Response Time: 45ms
└─ Total Requests: 200
Order Service Circuit Breaker
├─ State: OPEN ⚠️
├─ Success Rate: 20%
├─ Failure Rate: 80% (16/20)
├─ Last Failure: 15 seconds ago
├─ Avg Response Time: 5000ms (timeout)
├─ Time Until Recovery: 45 seconds
└─ Total Requests: 20
Configuration Best Practices
1. Per-Service Configuration
Different services have different characteristics:
circuit_breakers:
user_service:
failure_threshold: 5
timeout: 30s
payment_service:
failure_threshold: 2 # More aggressive
timeout: 10s
recommendation_service:
failure_threshold: 10 # More lenient
timeout: 60s
2. Tuning Parameters
| Parameter | Purpose | Typical Value |
|---|---|---|
| Failure Threshold | When to trip | 5-10 failures |
| Timeout Duration | Recovery wait | 30-60 seconds |
| Success Threshold | When to close from HALF-OPEN | 1-2 successes |
| Window Size | Measurement window | 60 seconds |
| Minimum Requests | Before calculating rate | 10-20 |
3. Avoid These Mistakes
❌ Too Aggressive:
failure_threshold: 1 # Trip after 1 failure
timeout: 1s
→ Trips on every temporary network hiccup
✅ Balanced:
failure_threshold: 5
timeout: 30s
minimum_requests: 20
failure_rate: 50%
❌ Too Lenient:
failure_threshold: 100
timeout: 300s
→ Doesn't protect system from cascading failures
Popular Circuit Breaker Libraries
Java
| Library | Features | Best For |
|---|---|---|
| Resilience4j | Lightweight, functional | New projects |
| Hystrix | Feature-rich, battle-tested | Production systems |
| Polly | Comprehensive (for .NET) | .NET ecosystem |
Example with Resilience4j
CircuitBreakerRegistry registry = CircuitBreakerRegistry.ofDefaults();
CircuitBreaker breaker = registry.circuitBreaker("userService",
CircuitBreakerConfig.custom()
.failureRateThreshold(50.0f)
.waitDurationInOpenState(Duration.ofSeconds(30))
.slowCallRateThreshold(50.0f)
.slowCallDurationThreshold(Duration.ofSeconds(2))
.minimumNumberOfCalls(10)
.build());
userService = Decorators.ofSupplier(() -> service.getUser(id))
.withCircuitBreaker(breaker)
.withRetry(Retry.ofDefaults("userService"))
.withFallback(List.of(
new TimeoutException(),
new CircuitBreakerOpenException()),
e -> cachedUser)
.decorate();
Cascading Failures: Before & After
Without Circuit Breaker
User Service 99% availability
Order Service 99% availability
Payment Service 99% availability
Recommendation Service 99% availability
Combined uptime:
0.99 × 0.99 × 0.99 × 0.99 = 96.06%
Single Order Service failure:
├─ User Service waits for Order Service timeout (30s)
├─ Orders pile up, thread pool exhausted
├─ User Service becomes unresponsive
└─ Entire system cascade fails
(Total downtime: hours)
With Circuit Breaker
Order Service fails
↓
After 5 failures, circuit trips (2-3 seconds)
↓
Circuit Breaker opens
↓
Subsequent requests fail immediately (no timeout wait)
↓
Thread pool recovers
↓
User Service continues to serve requests
↓
After 30 seconds, retry Order Service (HALF-OPEN)
↓
Order Service recovered
↓
Circuit closes
↓
Full system recovery (2-3 minutes)
Advanced Patterns
1. Bulkhead Isolation
Separate thread pools per service prevent one failure from affecting others:
Thread Pool A (User Service)
Thread Pool B (Order Service)
Thread Pool C (Payment Service)
Order Service fails:
├─ Pool B exhausted
└─ Pools A & C continue normally
2. Adaptive Thresholds
Adjust parameters based on time of day:
circuit_breaker:
peak_hours: (9:00 - 17:00)
failure_threshold: 10
timeout: 60s
off_peak: (17:00 - 9:00)
failure_threshold: 5
timeout: 30s
3. Service Mesh Integration
Tools like Istio implement Circuit Breaker at infrastructure level:
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
name: order-service
spec:
host: order-service
trafficPolicy:
outlierDetection:
consecutiveErrors: 5
interval: 30s
baseEjectionTime: 30s
maxEjectionPercent: 50
Anti-Patterns to Avoid
1. Silent Failures
❌ Bad:
try {
return circuitBreaker.execute(() -> service.call());
} catch (CircuitBreakerOpenException e) {
return null; // Silently fail
}
✅ Good:
try {
return circuitBreaker.execute(() -> service.call());
} catch (CircuitBreakerOpenException e) {
log.warn("Circuit open for service", e);
return getFallbackValue();
}
2. Ignoring the State
❌ Bad:
// No monitoring
circuitBreaker.execute(operation);
✅ Good:
// Monitor state transitions
circuitBreaker.onOpen(state ->
metrics.increment("circuit.open",
Tags.of("service", "order"))
);
circuitBreaker.onClose(state ->
metrics.increment("circuit.close")
);
3. Generic Configuration
❌ Bad:
# One size fits all
global_circuit_breaker:
threshold: 5
timeout: 30s
✅ Good:
# Per-service tuning
services:
critical_payment:
threshold: 2
timeout: 10s
cache_layer:
threshold: 20
timeout: 60s
Conclusion
The Circuit Breaker pattern is essential for building resilient distributed systems.
It provides:
✅ Fail-fast: Errors return immediately instead of timing out
✅ Resource protection: Prevents thread pool exhaustion
✅ Recovery time: Gives failing services time to heal
✅ Cascading failure prevention: Stops failure propagation
✅ Observability: Metrics on service health
A well-configured Circuit Breaker can mean the difference between a brief service hiccup and a system-wide outage.
Implement it today, tune it carefully, and monitor it relentlessly.
Top comments (0)