Cloudflare's Pingora story, where NGINX was replaced by a Rust proxy that uses far less CPU, made me want to understand what a reverse proxy actually has to get right. So I wrote a small one that you can read in one sitting: ferryman.
cargo install --locked ferryman
ferryman --config config.toml
# config.toml
health_interval_secs = 5
default_cooldown_secs = 30
failure_threshold = 3
[[routes]]
prefix = "/svc-a"
upstream = "http://localhost:8001"
[[routes]]
prefix = "/svc-b"
upstream = "http://localhost:8002"
It routes by path prefix and runs a circuit breaker for each upstream. It also does active health checks, hot-reloads its config, terminates TLS, streams bodies over HTTP/1.1 and HTTP/2, and exposes Prometheus metrics. The whole thing is two crates: ferryman-core holds the breaker, routing, config and health checks, and ferryman holds the hyper server.
The feature list is the boring part. What I actually learned came from the bugs, so here are three.
1. A dead upstream is a 503, not a fallback
My first routing table did what a lot of sketches do. It walked the routes longest-prefix-first and returned the first one that matched and whose upstream was healthy:
self.rules.iter()
.find(|(prefix, up)| path.starts_with(prefix) && up.is_routable())
That looks reasonable until you have /api/v1 and /api pointing at different services. When /api/v1 goes down, its requests quietly fall through to /api, and a completely different service receives them. That's much worse than an error.
There was a second bug on the same line: starts_with means /svc-a also matches /svc-ab.
The fix was to make lookup ignore health entirely and match only on segment boundaries:
fn prefix_matches(prefix: &str, path: &str) -> bool {
if prefix.ends_with('/') {
return path.starts_with(prefix);
}
if !path.starts_with(prefix) {
return false;
}
matches!(path.as_bytes().get(prefix.len()), None | Some(b'/'))
}
The handler then asks the matched upstream's breaker for permission. If the circuit is open, the client gets a 503 upstream unavailable, which is honest and easy to alert on.
2. Late results break a lock-free circuit breaker
The breaker sits on the hot path, so everything in it is atomics: the state, a count of consecutive failures, and an opened_at timestamp in monotonic milliseconds. There are no locks.
The classic state machine goes like this. After N failures the circuit opens. Once the cooldown passes it goes half-open and lets exactly one probe request through. If the probe succeeds the circuit closes; if it fails, it opens again.
Picking exactly one probe is a compare-and-swap on the timestamp, so whoever wins the CAS gets the probe:
fn claim_probe_slot(&self) -> bool {
let opened_at = self.opened_at_millis.load(Ordering::Acquire);
let now = now_millis();
if now.saturating_sub(opened_at) < self.cooldown_millis.load(Ordering::Relaxed) {
return false;
}
self.opened_at_millis
.compare_exchange(opened_at, now, Ordering::AcqRel, Ordering::Acquire)
.is_ok()
}
The subtle bug was late results. Suppose a slow request is admitted while the circuit is closed, the circuit then trips, and the slow request finishes successfully afterwards. A naive record_success() closes the circuit, which skips the cooldown and the single-probe rule completely. The opposite also happens: a request that times out late can reopen the circuit while the real probe is still in flight.
The fix is to give every admission a ticket:
pub enum Admission { Normal, Probe }
pub fn try_acquire(&self) -> Option<Admission>;
pub fn record_success(&self, a: Admission);
pub fn record_failure(&self, a: Admission);
-
Normalresults only count while the circuit is still closed. Otherwise they're stale and get ignored. -
Proberesults are authoritative: a success closes the circuit, and a failure opens it again. - Active health checks report as
Probe. A healthy answer therefore closes an open circuit straight away, which is what gives fast failover recovery. But a healthy/healthresponse while the circuit is closed does nothing. Otherwise an upstream whose/healthis fine but whose real endpoints fail would keep having its failure count reset.
Two smaller details mattered too:
-
Stamp
opened_atbefore publishingOpen. Another thread must never seeOpennext to a stale timestamp, or it immediately "wins" a probe. -
Keep the breaker across config reloads. Once I rebuilt breakers on every reload, in-flight requests reported their results to an orphaned breaker. Now the rebuilt table reuses the same
Arc<Breaker>for eachhost:portand only retunes its cooldown and threshold.
3. hyper-util's auto builder has a slowloris-shaped hole
I serve HTTP/1.1 and HTTP/2 on the same port with hyper_util::server::conn::auto::Builder. I set header_read_timeout, felt safe, and moved on.
A code review pointed out two problems:
-
header_read_timeoutis silently ignored unless you also give the builder a timer (.timer(TokioTimer::new())). - Even with a timer, the auto builder first sniffs the first bytes to decide between HTTP/1 and HTTP/2, and it keeps reading for as long as the bytes still look like the HTTP/2 preface
PRI * HTTP/2.0…. That sniff happens before hyper's HTTP/1 connection exists, so no header timeout applies. A client that sendsPRIand then goes quiet holds a file descriptor forever. A thousand of them andaccept()starts returningEMFILE.
The fix is a first-request deadline on every connection. A flag is set the first time the service is called:
let first_request_deadline = async {
tokio::time::sleep(HANDSHAKE_TIMEOUT).await;
if seen_request.load(Ordering::Relaxed) {
std::future::pending::<()>().await; // a request arrived; never fire
}
};
tokio::select! {
res = &mut conn => { /* normal end of connection */ }
_ = first_request_deadline => { /* nothing in flight yet: drop it */ }
}
This one timer also covers clients that finish a TLS handshake and then idle. There's a regression test that sends PRI and waits for the proxy to hang up.
Smaller things I'd tell my past self
-
A client failing is not an upstream failing. If a client aborts mid-upload, hyper reports it as a user error (
hyper::Error::is_user()), so answer 400 and leave the breaker alone. Otherwise three bad clients can open the circuit for everyone. -
HTTP/2 to HTTP/1.1 needs translating. Force the version to 1.1, set
Hostfrom:authority, and join HTTP/2's splitcookieheaders with"; "(RFC 9113 §8.2.3). -
cargo denyis worth running from day one. Default features had pulled inaws-lc-rs, which needs a C toolchain the musl build didn't have, plus a crate under a license I hadn't allowed. Turning those features off produced a 9 MBFROM scratchimage with onlyringfor crypto.
Numbers, honestly
From a CI smoke run (wrk2 at 5k rps through Docker, with an http-echo upstream on a GitHub runner):
| metric | result |
|---|---|
| errors | 0 |
| p50 | ~1.7 ms |
| p99 | ~3–4 ms |
| idle memory | 7.6 MB |
| recovery after an upstream restarts | 2.3 s (the next health check) |
The real target is 50k rps with p99 under 5 ms, and it hasn't been measured yet because that needs quiet, dedicated hardware. I'll post those numbers when I have them.
Try it and read it
- Code: https://github.com/Bunty9/ferryman (the architecture notes are in
docs/architecture.md) - Crates:
ferryman(binary + library) andferryman-core(breaker, routing, config) - Dual-licensed MIT / Apache-2.0
Next up are mTLS, JWT auth and WebSocket passthrough. Issues and "you got this wrong" comments are very welcome.
Top comments (0)