DEV Community

Cover image for The limit that matters isn't in my code: what I learned building a P2P file transfer app
d_tomassoni
d_tomassoni

Posted on AI-assisted

The limit that matters isn't in my code: what I learned building a P2P file transfer app

A few weeks ago I tried to send a 2 GB video to a friend. Email wouldn't take it. Dropbox wanted him to sign up. WeTransfer said the link would expire in three days, and he opened it on day four.

So I built Peerino. It's a desktop app that lets you share a file by sending a link. The other person opens the link in their browser and downloads directly from your machine. No account, no server holding the file. If you shut down your PC, the link dies.

The app isn't really the interesting part here. The interesting part was a design decision I got wrong, how I discovered it, and what I'd do differently now.

The original design

When I started, Peerino used TURN as a fallback. If two peers couldn't connect directly, the file was relayed through a TURN server. Cloudflare's, in my case. Files up to 100 MB went through it. The client was supposed to limit relayed transfers to 100 MB. It worked.

Then I opened the network tab while testing, and I saw the TURN credentials sitting there in plain text. Username and password, right in the receiver's browser. I knew they were there — that's how WebRTC works — but seeing them made me realize something I'd been avoiding.

Anyone with those credentials could use the relay for their own traffic. They didn't need my app at all. And my 100 MB limit was in my client, which they weren't using.

A limit that lives only in the client isn't a limit. It's a suggestion to well-behaved users.

I'd built the relay assuming people would use my app the way I'd designed it. That assumption was doing a lot of work.

What I changed

I turned TURN into a diagnostic instead of a transport. It's still configured in the receiver's ICE servers. If the direct path isn't available, ICE can select a relay candidate pair and the connection establishes through it. But then a gate on the sender checks the selected candidate pair and refuses to send the file.

The gate reads the WebRTC stats and looks for a relay candidate. This is the core of it, simplified:

const stats = await pc.getStats();
const statsMap = new Map();
stats.forEach((report) => statsMap.set(report.id, report));

stats.forEach((report) => {
  if (report.type === 'candidate-pair' &&
      (report.selected || report.nominated)) {
    const local = statsMap.get(report.localCandidateId);
    const remote = statsMap.get(report.remoteCandidateId);
    const isRelay =
      local?.candidateType === 'relay' ||
      remote?.candidateType === 'relay';
    // if isRelay: refuse the file
  }
});
Enter fullscreen mode Exit fullscreen mode

If the pair is relay, the sender sends an error message and the transfer stops. The user sees: "Your connection requires a relay server. To share this file, connect to WiFi."

The rule I ended up with is simple: if Peerino says a transfer is direct, the file bytes must actually travel directly between the two peers. If that isn't possible, the transfer fails. Side effect: the file data never goes through third-party infrastructure.

What the redesign actually fixed

Here's the part I didn't understand until I wrote it down.

The redesign didn't stop someone with the credentials from using the relay. They still can. The credentials still go to the browser. The rate limit on the Worker would have protected the old design too.

What changed is something else: legitimate traffic through the relay is now almost zero. A few kilobytes of handshake per failed attempt. Which means anything significantly above that is a strong signal that something is wrong. The egress monitor becomes a clean signal instead of a noisy one.

The threshold is 100 GB right now. That's enormously above real usage — it's a ceiling for "something has gone wrong", not a budget to spend. A 10 GB threshold would give me a much smaller failure budget while still being far above normal usage. I'll probably move it down once I have a month of real data.

To answer the question the redesign raises: what protects the account today? Rate limiting on the Worker that issues credentials — 40 requests per hour per source IP, with 15-minute credentials. There is no hard spending cap. The egress monitor sets a shutdown flag at 100 GB, and the proxy refuses new requests when it's set, but open connections aren't killed. I'd rather say it than pretend otherwise.

The bug I found while writing this

I re-read my own code before publishing this.

The design doc said "the gate refuses the file." The code did exactly that. It sent an error message and returned. But it didn't close the connection. The WebRTC session stayed alive, the TURN allocation stayed alive, and the user saw a stuck upload at 0 bytes in the UI.

It wasn't a divergence between the doc and the code. It was an omission in the doc. I never wrote down what should happen to the connection after the refusal, so nothing happened.

Fixed it before publishing this. But the lesson is that "the code does what the doc says" isn't the same as "the code does the right thing." The doc only described half the behavior.

What Peerino doesn't protect

I had encryption, but I hadn't really solved identity.

Signaling sees peer IDs and ICE candidates. TURN sees an IP pair and a handshake when a relay path is attempted. Neither sees file content. But the signaling server isn't authenticated out-of-band, so a compromised signaling server could, in theory, MITM the connection. There's no certificate pinning.

DTLS protects the connection. It doesn't tell me that the peer on the other end is actually the person I think it is. Those are two different problems, and I'd only solved the first one. The next step is out-of-band verification of the peer identity — showing the fingerprint to both users and letting them confirm. It's on the list.

What I'd do differently

There are three things I'd keep in mind if I were building this again.

Client-side limits aren't limits. If the cost of abuse lands on your bill, the control has to live somewhere the client can't modify. Don't build a relay whose safety depends on users being nice.

DTLS is not authentication. Your connection can be encrypted and still not be verified. If your threat model includes a compromised signaling server, encryption alone doesn't help you.

A refusal isn't complete until you've closed everything. When your code rejects something, ask what's still open. The gate refused the file, but the connection was still there, holding resources I didn't want to pay for. Every time I skip this question, I find it later.

If I started over, I'd never use TURN as a transport — which is what it still does today. The diagnostic role is the one I want. I'd still want the fast failure message, because that's the whole reason TURN is there. But I wouldn't spend the first week writing relay limits that a modified client can ignore. That time would go into the credential problem instead.

Try it, and tell me where it breaks

Peerino is open source, AGPL-3.0. Windows builds are available. Linux compiles but I haven't run it on a real distro yet, so if you're on Linux and want to try, build from source and tell me what breaks. macOS is not a priority.

The TURN proxy that issues credentials isn't in the repo. It's a separate Cloudflare Worker I keep private because it contains the rate limiting logic. The desktop client doesn't depend on it. Clone the repo, build the app, and it works out of the box with a hardcoded public TURN endpoint.

Code: https://codeberg.org/Daniele-Tomassoni/peerino
Site: https://peerino.com

If you try it and something breaks, open an issue or leave a comment. If you've solved the credential abuse problem differently, I'd like to hear how.

Top comments (0)