DEV Community

Cover image for Half-Open Circuit Breaker Probes in Serverless Workers
Raylabs
Raylabs

Posted on Originally published at raylabs.app

Half-Open Circuit Breaker Probes in Serverless Workers

Managing dependency recovery in serverless architectures requires coordinating multiple concurrent invocations to prevent a thundering-herd problem when a service heals. Quick solution: Coordinate the half-open lease through a single SQLite-backed Durable Object per dependency, keep the network probe outside the transaction, and defer competing callers.

async function acquireProbeLease(storage: DurableObjectStorage, now: number): Promise<boolean> {
  return await storage.transaction(async (txn) => {
    const state = await txn.get("breaker_state");
    if (state && state.status === "open" && state.cooldownUntil > now) {
      return false;
    }
    const lease = await txn.get("probe_lease");
    if (lease && lease.expiresAt > now) {
      return false;
    }
    await txn.put("probe_lease", { expiresAt: now + 5000 });
    return true;
  });
}
Enter fullscreen mode Exit fullscreen mode

A circuit breaker protects downstream dependencies by failing fast while an upstream service is unhealthy. The recovery window introduces a new risk when independent request handlers attempt recovery simultaneously. This race condition happens whenever multiple concurrent invocations share a dependency and coordinate through shared state. A scheduled invocation runs at its configured time rather than automatically fanning out into multiple triggers, but a sudden influx of traffic can create the same competitive pressure.

Understanding the Circuit Breaker States

A reliable serverless circuit breaker relies on three distinct operational states:

  • Closed: normal traffic passes directly to a healthy dependency.
  • Open: calls fail fast or defer while the dependency remains unhealthy.
  • Half-open: after a cooldown period, the system admits a controlled probe instead of immediately restoring all traffic.

A successful probe closes the circuit breaker. A failed or timed-out probe reopens it and schedules another recovery attempt. A cooldown period represents a probe opportunity rather than a guarantee of full recovery.

Choosing a Coordination Boundary

Choosing the right coordination backend is essential for safe probe admission. Eventually consistent storage layers cannot provide the atomic guarantees required for single-probe coordination.

Option When to Use It Trade-off or Failure Mode Recommendation
Workers KV Global read-heavy configurations Eventually consistent, lacks atomic read-modify-write Do not use for transactional lease locks.
SQLite-backed Durable Object Single-region state consistency and transactions Bound to a single location, requires routing Recommended for single coordination points.
External Relational Database Existing enterprise database clusters Adds network latency to serverless invocations Avoid if serverless cold-start and latency matter.

A SQLite-backed Durable Object provides strong consistency and transactional guarantees for a single coordination point. Use a stable Durable Object identity for each protected dependency to manage lease admission in a short transaction.

Managing Recovery and Failure Handling

Leases must include an expiration timestamp so a crashed or timed-out claimant does not permanently block recovery. Callers that do not receive the lease should return a retryable or deferred outcome rather than spinning in a tight retry loop.

On a successful probe, close the breaker and clear the active lease. On a failure or timeout, reopen the breaker with an appropriate delay and jitter. For further reading on operational safety mechanisms, see Self-Expiring Kill Switches and Human-Readable Alerts.

Verification Checklist

  • Issue five simultaneous lease requests against the Durable Object and verify that exactly one caller receives permission to probe.
  • Confirm that expired leases allow a new claimant to acquire the probe slot after the timeout window elapses.
  • Verify that late probe results from an expired lease do not overwrite newer breaker states in storage.
  • Check that callers without a lease receive a deferred response rather than executing unauthorized network traffic.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

Official Platform Update

Security protocols have been updated for all developer accounts.

  • tr.ee/dev-to