ferryman-edge is a small L7 reverse proxy in Rust. It's now on crates.io (cargo install ferryman-edge). Every request goes through four gates:
- mTLS. The client certificate must chain to a configured CA (rustls 0.23 with the aws-lc-rs provider).
-
JWT. An RS256 bearer token. Verified tokens are cached in moka, and
expis re-checked on every cache hit, so an expired token can't keep passing from the cache. -
Rate limit. A per-tenant GCRA limit (governor), keyed by the token's
sub. - Circuit breaker. One per upstream, with Closed / Open / HalfOpen states, plus an active health checker.
Certificates and routes hot-reload on SIGUSR1. Live connections keep the TLS config they handshook with; new connections get the new one.
The happy path is the easy part. This post covers the failure paths: five bugs that show up specifically because a proxy sits between two parties who can each misbehave.
1. Every HTTP/2 client got a 502
The listener advertises h2 and http/1.1 via ALPN. curl picks h2 by default. The proxy then forwarded the request to a plain http:// upstream using hyper-util's pooled client, which refused it:
UserUnsupportedVersion
hyper-util's legacy client won't send an HTTP/2-versioned request over an HTTP/1 connection. The inbound request's version had simply been carried across. The fix is one line:
// hyper-util's legacy Client rejects an HTTP/2-versioned request over
// an HTTP/1 connection (UserUnsupportedVersion) — upstreams here are
// plain http://, so always downgrade.
parts.version = http::Version::HTTP_11;
The mirror image of the bug: the upstream's response version was also copied back to the client. A Python http.server upstream answers with HTTP/1.0, so a keep-alive HTTP/1.1 client got an HTTP/1.0 status line. The response version is now reset too, and an end-to-end test sends a real h2 request.
2. A tripped breaker sent traffic to someone else's backend
Routing is longest-prefix on a path-segment boundary: /svc-a matches /svc-a/x but not /svc-abc. The original lookup was:
.find(|(prefix, up)| matches_prefix(path, prefix) && up.is_routable())
It reads fine, until /svc-a's breaker opens. Then find keeps walking, and /svc-a/x matches the / catch-all. One service's traffic quietly lands on another service's backend.
The fix is to pick the most specific route first, and only then ask whether it's routable:
pub fn lookup(&self, path: &str) -> Option<&Upstream> {
self.rules
.iter()
.find(|(prefix, _)| matches_prefix(path, prefix))
.map(|(_, up)| up)
.filter(|up| up.is_routable())
}
None becomes a 503, which is what an open breaker should look like to a client.
3. Exactly one half-open probe, without a lock
When a breaker's cooldown expires, exactly one request should go through as the probe. Everyone else should keep getting 503 until the probe reports back. The first version used the state byte as the single-flight token (a compare-and-swap from Open to HalfOpen). That leaves an ABA window: two callers can both see a stale HalfOpen, both reset it, and both probe.
The version that shipped uses the transition timestamp as the token instead:
let stamped = self.last_transition_unix.load(Ordering::Acquire);
let now = now_secs();
if now.saturating_sub(stamped) < self.cooldown_secs {
return false;
}
if self
.last_transition_unix
.compare_exchange(stamped, now, Ordering::AcqRel, Ordering::Relaxed)
.is_err()
{
return false; // someone else won the probe
}
A test spawns 8 threads behind a barrier, 200 times over, and asserts that exactly one is admitted each time. A related trap: with cooldown_secs = 0, stamped == now for every caller in the same second, so everyone probes. The config loader now rejects 0.
4. Any tenant could open a route's breaker for everyone
This is the one I'd most want another proxy author to know about.
With streaming bodies on (the boxed_body feature), the client's upload runs inside the upstream request. If the client disconnects mid-upload, or sends a chunked body over the 8 MiB cap, hyper reports it as an error on the upstream call. The proxy treated every upstream-call error as "the upstream is unhealthy" and opened the breaker.
So any authenticated tenant could take a route down for every other tenant by starting uploads and cutting them off.
The rule that came out of this: only failures the upstream caused may count against it. Client-side body errors show up in hyper's error chain as user errors, so the classification is a walk down the source() chain:
fn is_client_body_error(e: &(dyn std::error::Error + 'static)) -> bool {
std::iter::successors(Some(e), |c| c.source()).any(|c| {
c.is::<LengthLimitError>()
|| c.downcast_ref::<hyper::Error>().is_some_and(|h| h.is_user())
})
}
Two more versions of the same bug turned up in review:
- Slow uploads. A single 30-second deadline covered both reading the client's body and calling the upstream. A client could trickle its body for 29.9 s, leave the upstream 0.1 s, and trip the breaker. Now the body has its own deadline (408), and the upstream's clock starts only once the body is in hand. In streaming mode, a small body wrapper records when the upload finished, and a timeout blames the upstream only after that.
- Wasted probes. The body is now read before route lookup, because lookup may admit the request as the single half-open probe. A probe that ends in a client-side 408 never reports back, and recovery stalls for another full cooldown.
5. Connection: x-ferryman-tenant
After verifying the JWT, the proxy stamps x-ferryman-tenant: <sub> so upstreams know who's calling. Any client-supplied value is removed first. Hop-by-hop headers are stripped too, including any header named in Connection, as RFC 9110 requires.
The strip ran after the stamp. So a client sending
Connection: keep-alive, x-ferryman-tenant
had the proxy delete its own authoritative header on the way out. The fix is ordering: strip, then stamp. The regression test sends exactly that header over HTTP/1.1. I checked the test by moving the strip back after the stamp and watching it fail. (HTTP/2 isn't affected, because h2 forbids the Connection header.)
Honourable mentions
-
tokio::select!guards are evaluated once. The first attempt at "close connections that send nothing for 10 s" used_ = &mut timer, if !seen_request => …. The guard is checked whenselect!starts, not when the timer fires, so a long first request got killed too. The fix is to check the flag inside the branch. Clippy caught the surrounding code as "this loop never actually loops", which is what led me to it. -
jsonwebtoken 9 checks
iss/audonly if they're present.set_issueralone accepts a token that simply omitsiss. You also needrequired_spec_claims.insert("iss"). A test caught this before release. -
pgrep -x ferryman-edge-servernever matches. The kernel truncates a process name (comm) to 15 characters. Usepidof. -
The Docker image couldn't start (
GLIBC_2.38 not found): a trixie-based builder and a distroless bookworm runtime. The builder is now pinned to bookworm.
Numbers (measured, with caveats)
| What | Result |
|---|---|
| Hot reload under load | 3,725 / 3,725 OK across 2× SIGUSR1, 60 s, 8 curl workers, release build |
| JWT verify, cache hit vs miss | 0.68 µs vs ~150 µs (~220×), criterion |
| RSS after that run | 16 MB |
| TLS handshake p99 | 119 ms, but client and server shared one machine; not representative |
| 50k req/s target | Not measured. wrk/wrk2 can't present a client certificate, so it needs an mTLS-capable load generator first |
The reload check deliberately uses a fresh curl process per request. Every request is a new mTLS handshake, so each one exercises the config swap, not just a warm keep-alive connection.
Try it
cargo install ferryman-edge
- Guide (configuration, operations, design): https://bunty9.github.io/ferryman-edge/
- API docs: https://docs.rs/ferryman-edge-core
- Code: https://github.com/Bunty9/ferryman-edge
The reusable pieces are published as ferryman-edge-core: the reloading TLS config, the cached JWT verifier, the per-tenant limiter, and the routing table with its breaker. Issues and PRs welcome, especially from anyone with an mTLS-capable load generator.
Top comments (0)