DEV Community

Davi Orlandi
Davi Orlandi

Posted on

Rate Limiting Is Not Back Pressure: Layering Protection in Go Services

I used to treat rate limiting and back pressure as synonyms people swapped in design reviews for variety. Then an afternoon that should have been quiet melted a service I cared about, and the difference stopped being academic.

We had an ingestion API capped at 5,000 requests per second. That number came from an old benchmark with margin. A migration slowed the database from roughly 4ms writes to 90ms. The rate limiter kept admitting 5,000 rps because nothing about the database was visible to it. The internal queue between HTTP and the writer grew. Latency climbed from tens of milliseconds to tens of seconds. The pod died on memory. The rate limiter did its job. Its job was never this.

That story is why I write this essay. Rate limiting enforces a fixed ceiling chosen in advance. Back pressure is a signal travelling backwards from the component that is actually struggling. Load shedding is what you do when acceptance would make everything worse. The thing that defeats all three is an unbounded queue: it converts a fast rejection into unbounded latency and then a memory failure.

Two mechanisms, two failures

Rate limiting answers: is this caller allowed to send this much? The number comes from a contract, a tier, or last year's load test.

Back pressure answers: can this system absorb more right now? The answer moves when a dependency slows, a node disappears, or a noisy neighbor steals CPU.

The direction of the arrow is the whole distinction. Rate limiting pushes a constraint at the entrance. Back pressure pushes information upstream from wherever the bottleneck currently lives. You usually want both. Pretending one replaces the other is how you get green policy dashboards next to a dying process.

The unbounded queue is what makes overload fatal

jobs := make(chan Job) // unbounded in practice if you keep spawning senders

func Submit(j Job) error {
    select {
    case jobs <- j:
        return nil
    default:
        return ErrOverloaded // only works if the channel is buffered and full
    }
}

jobs = make(chan Job, 256) // capacity is an explicit decision
Enter fullscreen mode Exit fullscreen mode

An unbounded buffer does not remove overload. It hides it until latency and memory become the outage. Export queue depth everywhere you have a queue. Depth usually moves before the error rate does. That single habit has saved me more times than any clever token-bucket tweak.

A practical layering order in Go

At the HTTP edge I want a boring sequence:

  1. Identify the caller (limits need a key)
  2. Rate limit per API key, IP, or tenant
  3. Cap in-flight work with a semaphore or buffered channel
  4. Run the handler with deadlines
  5. Bound concurrency on outbound clients too
type Gate struct {
    lim   *rate.Limiter
    slots chan struct{}
}

func (g *Gate) Middleware(next http.Handler) http.Handler {
    return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
        if !g.lim.Allow() {
            http.Error(w, "rate limit", http.StatusTooManyRequests)
            return
        }
        select {
        case g.slots <- struct{}{}:
            defer func() { <-g.slots }()
            next.ServeHTTP(w, r)
        default:
            w.Header().Set("Retry-After", "1")
            http.Error(w, "overloaded", http.StatusServiceUnavailable)
        }
    })
}
Enter fullscreen mode Exit fullscreen mode

Use 429 for policy. Use 503 for overload. Clients should treat them differently: 429 means slow to the agreed rate; 503 means back off harder, with jitter, and consider opening a circuit. Collapsing both into one generic "retry later" response trains clients to stampede.

Load shedding is back pressure with a decision attached

Reject at the entrance before work starts. Prefer dropping low-regret traffic first: batch jobs, scrapers, expensive exports. Protect login and checkout longer. Check deadlines on queued work so you do not spend capacity on results nobody will read. Shedding is not cruelty. It is choosing which failures you can explain.

What I look for in reviews

Every queue should have a bound and a full policy. 429 and 503 should be intentional, not interchangeable. Shed rate should be on a dashboard, not only error rate. Retries should be budgeted at each layer. Clients should honor Retry-After. If those five lines are missing, the service is polite until it is catastrophic.

Closing

Rate limiting is a configured ceiling. Back pressure is a live capacity signal. One encodes what you believed last year. The other encodes what is true this minute. Bound every queue, fail fast when full, layer a token bucket in front of a concurrency gate, and teach clients to listen. That is not a research project. It is the difference between a graceful slow-down and a chain-reaction outage that starts with a number you once thought was safe.

Top comments (0)