DEV Community

Cover image for Building a small reverse proxy in Rust: three bugs a circuit breaker taught me
Bipin C
Bipin C

Posted on

Building a small reverse proxy in Rust: three bugs a circuit breaker taught me

Cloudflare's Pingora story, where NGINX was replaced by a Rust proxy that uses far less CPU, made me want to understand what a reverse proxy actually has to get right. So I wrote a small one that you can read in one sitting: ferryman.

cargo install --locked ferryman
ferryman --config config.toml
Enter fullscreen mode Exit fullscreen mode
# config.toml
health_interval_secs = 5
default_cooldown_secs = 30
failure_threshold = 3

[[routes]]
prefix = "/svc-a"
upstream = "http://localhost:8001"

[[routes]]
prefix = "/svc-b"
upstream = "http://localhost:8002"
Enter fullscreen mode Exit fullscreen mode

It routes by path prefix and runs a circuit breaker for each upstream. It also does active health checks, hot-reloads its config, terminates TLS, streams bodies over HTTP/1.1 and HTTP/2, and exposes Prometheus metrics. The whole thing is two crates: ferryman-core holds the breaker, routing, config and health checks, and ferryman holds the hyper server.

The feature list is the boring part. What I actually learned came from the bugs, so here are three.

1. A dead upstream is a 503, not a fallback

My first routing table did what a lot of sketches do. It walked the routes longest-prefix-first and returned the first one that matched and whose upstream was healthy:

self.rules.iter()
    .find(|(prefix, up)| path.starts_with(prefix) && up.is_routable())
Enter fullscreen mode Exit fullscreen mode

That looks reasonable until you have /api/v1 and /api pointing at different services. When /api/v1 goes down, its requests quietly fall through to /api, and a completely different service receives them. That's much worse than an error.

There was a second bug on the same line: starts_with means /svc-a also matches /svc-ab.

The fix was to make lookup ignore health entirely and match only on segment boundaries:

fn prefix_matches(prefix: &str, path: &str) -> bool {
    if prefix.ends_with('/') {
        return path.starts_with(prefix);
    }
    if !path.starts_with(prefix) {
        return false;
    }
    matches!(path.as_bytes().get(prefix.len()), None | Some(b'/'))
}
Enter fullscreen mode Exit fullscreen mode

The handler then asks the matched upstream's breaker for permission. If the circuit is open, the client gets a 503 upstream unavailable, which is honest and easy to alert on.

2. Late results break a lock-free circuit breaker

The breaker sits on the hot path, so everything in it is atomics: the state, a count of consecutive failures, and an opened_at timestamp in monotonic milliseconds. There are no locks.

The classic state machine goes like this. After N failures the circuit opens. Once the cooldown passes it goes half-open and lets exactly one probe request through. If the probe succeeds the circuit closes; if it fails, it opens again.

Picking exactly one probe is a compare-and-swap on the timestamp, so whoever wins the CAS gets the probe:

fn claim_probe_slot(&self) -> bool {
    let opened_at = self.opened_at_millis.load(Ordering::Acquire);
    let now = now_millis();
    if now.saturating_sub(opened_at) < self.cooldown_millis.load(Ordering::Relaxed) {
        return false;
    }
    self.opened_at_millis
        .compare_exchange(opened_at, now, Ordering::AcqRel, Ordering::Acquire)
        .is_ok()
}
Enter fullscreen mode Exit fullscreen mode

The subtle bug was late results. Suppose a slow request is admitted while the circuit is closed, the circuit then trips, and the slow request finishes successfully afterwards. A naive record_success() closes the circuit, which skips the cooldown and the single-probe rule completely. The opposite also happens: a request that times out late can reopen the circuit while the real probe is still in flight.

The fix is to give every admission a ticket:

pub enum Admission { Normal, Probe }

pub fn try_acquire(&self) -> Option<Admission>;
pub fn record_success(&self, a: Admission);
pub fn record_failure(&self, a: Admission);
Enter fullscreen mode Exit fullscreen mode
  • Normal results only count while the circuit is still closed. Otherwise they're stale and get ignored.
  • Probe results are authoritative: a success closes the circuit, and a failure opens it again.
  • Active health checks report as Probe. A healthy answer therefore closes an open circuit straight away, which is what gives fast failover recovery. But a healthy /health response while the circuit is closed does nothing. Otherwise an upstream whose /health is fine but whose real endpoints fail would keep having its failure count reset.

Two smaller details mattered too:

  • Stamp opened_at before publishing Open. Another thread must never see Open next to a stale timestamp, or it immediately "wins" a probe.
  • Keep the breaker across config reloads. Once I rebuilt breakers on every reload, in-flight requests reported their results to an orphaned breaker. Now the rebuilt table reuses the same Arc<Breaker> for each host:port and only retunes its cooldown and threshold.

3. hyper-util's auto builder has a slowloris-shaped hole

I serve HTTP/1.1 and HTTP/2 on the same port with hyper_util::server::conn::auto::Builder. I set header_read_timeout, felt safe, and moved on.

A code review pointed out two problems:

  1. header_read_timeout is silently ignored unless you also give the builder a timer (.timer(TokioTimer::new())).
  2. Even with a timer, the auto builder first sniffs the first bytes to decide between HTTP/1 and HTTP/2, and it keeps reading for as long as the bytes still look like the HTTP/2 preface PRI * HTTP/2.0…. That sniff happens before hyper's HTTP/1 connection exists, so no header timeout applies. A client that sends PRI and then goes quiet holds a file descriptor forever. A thousand of them and accept() starts returning EMFILE.

The fix is a first-request deadline on every connection. A flag is set the first time the service is called:

let first_request_deadline = async {
    tokio::time::sleep(HANDSHAKE_TIMEOUT).await;
    if seen_request.load(Ordering::Relaxed) {
        std::future::pending::<()>().await; // a request arrived; never fire
    }
};
tokio::select! {
    res = &mut conn => { /* normal end of connection */ }
    _ = first_request_deadline => { /* nothing in flight yet: drop it */ }
}
Enter fullscreen mode Exit fullscreen mode

This one timer also covers clients that finish a TLS handshake and then idle. There's a regression test that sends PRI and waits for the proxy to hang up.

Smaller things I'd tell my past self

  • A client failing is not an upstream failing. If a client aborts mid-upload, hyper reports it as a user error (hyper::Error::is_user()), so answer 400 and leave the breaker alone. Otherwise three bad clients can open the circuit for everyone.
  • HTTP/2 to HTTP/1.1 needs translating. Force the version to 1.1, set Host from :authority, and join HTTP/2's split cookie headers with "; " (RFC 9113 §8.2.3).
  • cargo deny is worth running from day one. Default features had pulled in aws-lc-rs, which needs a C toolchain the musl build didn't have, plus a crate under a license I hadn't allowed. Turning those features off produced a 9 MB FROM scratch image with only ring for crypto.

Numbers, honestly

From a CI smoke run (wrk2 at 5k rps through Docker, with an http-echo upstream on a GitHub runner):

metric result
errors 0
p50 ~1.7 ms
p99 ~3–4 ms
idle memory 7.6 MB
recovery after an upstream restarts 2.3 s (the next health check)

The real target is 50k rps with p99 under 5 ms, and it hasn't been measured yet because that needs quiet, dedicated hardware. I'll post those numbers when I have them.

Try it and read it

Next up are mTLS, JWT auth and WebSocket passthrough. Issues and "you got this wrong" comments are very welcome.

Top comments (0)