DEV Community

Nainik Mehta
Nainik Mehta

Posted on

Colocate Adaptive Concurrency at the Bottleneck (2026)

TL;DR

Put adaptive concurrency limits where the resource is actually saturated, not just at the edge. Colocated, latency-driven admission control (Vegas/Gradient2/AIMD/PID patterns) gives faster, more accurate backpressure, prevents queue growth and retry storms, and preserves core paths during overloads.

Why colocate admission control?

Traditional edge rate limits and global token buckets shape traffic but often miss the real signal of collapse: queue growth and resource pressure at a storage or RPC hotspot. If the limiter only sees edge arrivals, it can “admit” work faster than a particular node can handle. By the time the edge notices elevated latencies, the node is already overloaded.

Colocating admission control with the hot resource gives you three advantages:

  • Earliest, most accurate signal: measure minRTT, current latency, goroutine counts, memory PSI or CPU throttle from the node itself.
  • Local, cheap decisions: no global coordination or cross-node consensus required for immediate backpressure.
  • Better outcomes: avoid long queues, reduce p99s, and prevent retry storms that amplify failures.

Real-world systems have converged on this approach: Uber’s Cinnamon, GitLab’s Gitaly adaptive limits, Netflix’s concurrency libraries, and databases like CockroachDB all colocate some form of admission or concurrency control near the bottleneck.

How the control loop works (concept)

Think TCP congestion control but for request admission. The control loop compares a reference latency (minRTT) to a live sample (curRTT). If curRTT grows, that’s evidence a queue is forming and the limiter should shrink the allowed in-flight requests. If curRTT is near minRTT, the limiter can gently grow.

Common patterns:

  • Vegas-style: increment/decrement the in-flight limit based on the ratio between sampleRTT and minRTT.
  • Gradient2/EMA: compare short and long-window latency averages to smooth bursts.
  • AIMD / PID: additive increase for normal operation, multiplicative decrease on backoff events. A PID regulator can target queue length directly.
  • Bring-your-own-signal: fuse latency with local signals (runnable goroutines, PSI, memory ratio) for robust decisions.

Concrete example: moving the limiter next to storage

The problem I observed: checkout throughput fell to 5% during a slow DB query because the API gateway kept admitting requests — it only saw edge traffic, not the growing queues at storage nodes. We moved the limiter from the gateway into the storage worker process. Under a spike, the in-flight cap fell from ~60 to ~12 when curRTT exceeded 1.5× minRTT. That local decision prevented queue growth, avoided a retry storm, and kept core paths healthy while degraded paths shed load.

Lessons learned from large deployments:

  • Uber/Cinnamon: priority-aware queues, Vegas-derived auto-tuner, a PID-based rejector, and a “bring your own signal” model for node metrics.
  • GitLab/Gitaly: combine cgroup signals with AIMD-style adjustments; provide min/max bounds and per-RPC scoping.
  • CockroachDB: slot/token models and monitoring runnable goroutines to adjust admission for LSM compaction pressure.

Minimal, illustrative control loop (Go)

This tiny loop shows the core idea: maintain minRTT, sample RTTs, and tweak an in-flight limit. In production you’ll want smoothing, bounds, covariances, and safety guards.

// illustrative only
var (
    limit    = 20
    minRTT   = time.Millisecond * 5
    alpha    = 1.5 // backoff trigger
    mu       sync.Mutex
)

func controlLoop() {
    for range time.Tick(time.Second) {
        curRTT := sampleRTT()   // short-window p95/p99 sample
        mu.Lock()
        if curRTT > time.Duration(float64(minRTT)*alpha) {
            if limit > 1 { limit = int(math.Max(1, float64(limit-1))) }
        } else {
            limit = limit + 1
        }
        // optional: cap by observed runnable goroutines or memory pressure
        limit = int(math.Min(float64(limit), observedMaxLimit()))
        mu.Unlock()
    }
}

func admitRequest() bool {
    mu.Lock()
    defer mu.Unlock()
    if inFlight >= limit { return false }
    inFlight++
    return true
}
Enter fullscreen mode Exit fullscreen mode

This is intentionally small — real systems add smoothing windows for minRTT, backoff multipliers (multiply-by-0.75), PID regulators for queue length, and bounds for min/max limits.

Signals to fuse (bring-your-own-signal)

Pick the signals closest to the resource:

  • Latency: minRTT and short-window p95/p99.
  • Concurrency state: actual in-flight requests, runnable goroutine counts (normalized by CPU), etc.
  • System pressure: cgroup memory ratio, PSI, I/O stall metrics, CPU throttle percent.
  • Application hints: cost scores per-RPC, priority tiers (interactive vs background).

Fusing signals avoids false positives (spurious latency spikes) and false negatives (high memory pressure but normal latency because the queue is still building).

Practical rollout & observability

Start small and incremental:

  1. Instrument: expose minRTT, sampleRTT, in-flight counts, queue sizes, and node pressure metrics to Prometheus/your telemetry.
  2. Localize a simple limiter to one service or partition and run in observe-only mode (log decisions, don’t reject) while comparing outcomes.
  3. Add rejection metadata: include priority, reason, and Retry-After. Structured rejections help clients back off instead of retrying immediately.
  4. Enable adaptive mode with conservative min/max bounds, and monitor SLOs (p50/p95/p99, error rates, queue sizes).
  5. Gradually widen rollout and tune the auto-tuner (window sizes, covariance checks) and priority tiers.

Tradeoffs and caveats

  • Local decisions reduce coordination but can be inconsistent across replicas. Use partition floors or global fallback for rare cross-node imbalances.
  • Adaptive limits react to observed signals — poorly chosen signals or windows can oscillate. Use PID or covariance checks to stabilize.
  • Don’t replace edge shaping: keep edge rate limits for coarse shaping and abuse protection. Colocated admission is the last line before resource exhaustion.

Conclusion

Adaptive concurrency limits colocated with the saturated resource are a high-leverage, low-coordination way to protect services during spikes. The ideas come from decades of congestion control (Vegas, AIMD, PID) and recent production frameworks (Uber Cinnamon, GitLab Gitaly, Netflix libraries, CockroachDB). Start with a simple latency-driven loop, fuse the right node-level signals, add small priority tiers, and roll it out observably. You’ll often find the local controller prevents collapse far earlier than any edge-layer heuristic ever could.

Have you moved admission control to the bottleneck? Share what it saved you from and which signals surprised you most.

Top comments (0)