The download half of my bufferbloat test was measuring an idle line. The code asked Cloudflare's speed endpoint for a gigabyte, got back a 403 with a one-byte body, and never looked at the status. The pings kept running, the grades kept coming out, and every result looked like data.
I build NetDiag+, an iOS network toolkit, and wrote here before about ping and traceroute without entitlements and measuring latency without root. In that last post I showed the bufferbloat classifier, including this branch:
if dnN >= upN * 2 { return .ispDownstream }
"Download latency rose much more than upload: the queue is on the provider's downstream path." It reads well. It could almost never run.
This post is about how that happened, how I found it, and the lesson I took from it: a load generator has to measure its own load. The feature that exposed the bug is a Health Score that turns one run of about 40 seconds into a number and a verdict on who is to blame.
How a bufferbloat test works
Bufferbloat is latency that appears only when the line is busy. An idle ping to 1.1.1.1 says 14 ms; start a big download on the same connection and the same ping says 180 ms, because packets are now waiting in an oversized buffer somewhere between you and the internet. That is what makes a video call stutter while someone else in the house updates a game.
The test is simple in principle:
- Baseline: ping the home gateway and an internet target for five seconds with the line idle.
- Download: saturate the downlink for ten seconds while the same two pings keep running.
- Recovery: two seconds of nothing, so buffers drain.
- Upload: saturate the uplink for ten seconds, pings still running.
The grade (A+ to F) comes from how far the internet ping rises over its baseline under load. The attribution comes from the gateway ping: if the gateway climbs too, the queue is between the phone and the router, i.e. Wi-Fi; if the gateway stays flat and the internet leg climbs, the queue is past the router. Comparing the two load phases then separates the router's upload buffer from the provider's downstream queue.
The physics behind the attribution is the old argument from every bufferbloat thread: the queue forms at the bottleneck. Toward your home, the bottleneck is usually the provider's equipment, so download bloat lives there. Toward the internet, the bottleneck is usually your own modem or router, so upload bloat is yours. That asymmetry is why comparing the two load phases can separate them at all.
It also shapes the fix: upload bloat is solved on your router with fq_codel or cake; download bloat can be tamed from home by shaping inbound traffic a few percent below line rate, or handed to the provider as evidence. (Details in the bufferbloat guide.)
Every one of those steps assumes the load phases actually load something.
The load that wasn't
Here is the download saturation as it shipped:
private static let downloadURL =
URL(string: "https://speed.cloudflare.com/__down?bytes=1073741824")! // 1 GB — we cancel by time
private static func saturateDownload(duration: TimeInterval) async {
let session = URLSession(configuration: .ephemeral)
defer { session.invalidateAndCancel() }
do {
let (bytes, _) = try await session.bytes(from: downloadURL)
let deadline = ContinuousClock.now.advanced(by: .seconds(duration))
for try await _ in bytes {
if ContinuousClock.now >= deadline { break }
if Task.isCancelled { break }
}
} catch {
// network error — measurement still yields whatever ICMP got.
}
}
Ask for a gigabyte, drain it for ten seconds, cancel. Nothing in it is wrong as Swift. Three things in it are wrong as a measurement:
-
The response is never looked at.
(bytes, _)throws away theURLResponse. Whatever status came back, the loop ran. -
Nothing is counted. The function returns
Void. There is no way for the caller to learn whether ten seconds of saturation moved ten gigabits or ten bytes. - The failure path is silent by design. The comment even says so: whatever happens, "the measurement still yields whatever ICMP got".
So I asked the endpoint directly:
$ curl -s -o /dev/null -w "%{http_code} %{size_download}\n" \
"https://speed.cloudflare.com/__down?bytes=99999999"
200 99999999
$ curl -s -o /dev/null -w "%{http_code} %{size_download}\n" \
"https://speed.cloudflare.com/__down?bytes=104857600"
403 1
Anything from 100 MB up gets a 403 with a one-byte body, with any user agent. I cannot say whether that limit was already there when I wrote the code. The code gave itself no way to notice either way. The "download phase" had been ten seconds of pings over an idle line.
What it did to the verdicts
Tracing it through the classifier is the uncomfortable part. With no download load, the download rise over baseline is roughly zero. Then:
let worstRise = [dnNetRise, upNetRise].compactMap { $0 }.max()
let grade = Grade.from(latencyRiseMs: worstRise) // effectively: upload only
...
if grade == .aPlus || grade == .a || grade == .b { return .clean }
if upN >= dnN * 2 { return .uplinkQueue } // true for any upload bloat over ~0
if dnN >= upN * 2 { return .ispDownstream } // unreachable
return .modemOrLine // nearly unreachable
- The grade was an upload-only grade. A line that was clean under upload and awful under download got an A.
- Any real bloat was blamed on the router's upload queue, because "upload rose at least twice as much as download" is true when download rose by nothing.
- "ISP downstream", the verdict the tool exists to produce for a support ticket, could not fire.
Gateway attribution still worked, since the upload phase did load the line, so Wi-Fi faults were caught. But the one number people screenshot for their provider was half a measurement.
How it surfaced
Not from a user report. It surfaced when I started reusing the load phases for something new.
The new feature, Health Score, wants a speed number, and running a separate speed test next to a bufferbloat test is wasteful: both saturate the same line. So the plan was to count the bytes the load phases already move. The first simulator run came back with:
score 89/100 · Excellent · grade A+
download 0.0000008 Mbps · upload 65.0 Mbps
An excellent connection that downloads at less than one bit per second is a contradiction you cannot miss. The upload number was plausible; the download number was the single byte the 403 carried, spread over ten seconds. The moment the load had to report its own size, the lie had nowhere to go.
That is the general lesson I took: a load generator has to measure its own load. If the part of a test that is supposed to stress the system cannot fail visibly, it will eventually stress nothing, and every number downstream of it will still look like data.
The fix
Download is now 50 MB requests back to back until the window closes, with bytes counted in the data delegate:
private static let downloadURL =
URL(string: "https://speed.cloudflare.com/__down?bytes=52428800")! // 50 MB, under the limit
private static func saturateDownload(duration: TimeInterval) async -> Double? {
let counter = ByteCounter()
let session = URLSession(configuration: .ephemeral, delegate: counter, delegateQueue: nil)
let start = ContinuousClock.now
let deadline = start.advanced(by: .seconds(duration))
let loop = Task {
while ContinuousClock.now < deadline, !Task.isCancelled {
await counter.run(session.dataTask(with: downloadURL)) // 50 MB chunk
}
}
try? await Task.sleep(nanoseconds: UInt64(duration * 1_000_000_000))
loop.cancel()
session.invalidateAndCancel()
let seconds = elapsedSeconds(since: start)
guard seconds > 1, counter.bytes > 0 else { return nil }
return Double(counter.bytes) * 8 / seconds / 1_000_000
}
Two details that matter:
-
The delegate counts whole buffers.
didReceive data:hands over chunks of tens of kilobytes. The oldfor try await _ in bytesiterated one byte at a time, which on a fast line is a CPU benchmark, not a network one: the phone gives up well before a gigabit connection does. - The function returns a rate, or nil. A nil propagates into the result as "no speed measured", and the grade logic can see it.
Upload got the same treatment for a different reason. It used to POST one 32 MB body. On anything faster than about 25 Mbps that finishes early and the uplink sits idle for the rest of the phase, exactly when the test wants the queue full. It is now 8 MB POSTs back to back until the window closes, counted in didSendBodyData.
After the fix, on the same Wi-Fi: 455 Mbps down, 62 Mbps up, grade A instead of A+, because the download phase now finally adds a few milliseconds.
Health Score: one run, one number, one culprit
With the load phases honest, Health Score is mostly orchestration of tools that already existed:
| Phase | What runs | Time |
|---|---|---|
| Idle | ICMP to 1.1.1.1 and 8.8.8.8, TCP connect to 1.1.1.1:443, gateway ping | ~8 s |
| Load | the Bufferbloat probe, now counting bytes | ~27 s |
| Recovery | ten more pings to 1.1.1.1 | ~3 s |
The phases are sequential on purpose. Throughput and latency cannot be measured cleanly at the same time: the throughput test fills the queue, and a filled queue corrupts the latency number. That corruption is bufferbloat. So the latency you care about for gaming is measured idle, and the latency you care about for "everything lags when someone downloads" is measured under load.
The score is a weighted sum of five components:
| Component | Max | What earns the points |
|---|---|---|
| Responsiveness under load | 35 | the Bufferbloat grade: A+ 35, A 30, B 22, C 12, D 5, F 0 |
| Idle latency | 20 | under 20 ms 20, under 40 ms 16, under 80 ms 10 |
| Stability | 20 | idle jitter under 5 ms 20; any loss −6, 10% or more zeroes it |
| Speed | 15 | download tiers up to 12, upload up to 3; gentler curve on cellular |
| Local link | 10 | gateway RTT and jitter; points given on cellular |
Responsiveness gets the biggest weight deliberately. Most people's complaint is not "my internet is slow" in the speed-test sense; it is that it becomes unusable at certain moments, and that is a queueing problem, not a bandwidth one.
Under the number there are three rows: Wi-Fi, Router, ISP, each ok, suspect or at fault. They come from the same rules as the Bufferbloat verdict, plus the idle numbers that load never touches: a gateway that already answers in 30 ms is Wi-Fi's fault, loss on the idle path with a healthy gateway is the provider's.
One more lesson from the first runs. Idle loss was originally the worst row: a single refused TCP handshake out of ten gave that row 10% "loss", which cost eight points and blamed the ISP for one dropped SYN. It is now the mean over the ICMP rows, twenty scored samples. Averaging where you have the samples and taking the worst where you do not is a choice worth writing down.
Nothing leaves the phone. The scoring is a pure function over the numbers in one run, and every row expands on tap to show its thresholds, so a user can see why they got 64 and not 80.
This is what the Health Score detail rows look like, from another run on the same Wi-Fi:
| Metric | Value |
|---|---|
| Bufferbloat grade | A |
| Idle latency | 14 ms |
| Jitter | 1 ms |
| Loss | 0% |
| Download | 447.5 Mbps |
| Upload | 68.4 Mbps |
| Gateway | 3 ms |
Three iOS networking gotchas from the same release
-
VPN Check. Is a tunnel up, what is your public identity, what NAT type does STUN see, and does the system resolver agree with Cloudflare DoH. On iOS, the presence of
utuninterfaces tells you nothing, since every iPhone has several for system services; a VPN shows up as autun/ipseckey under__SCOPED__inCFNetworkCopySystemProxySettings(), or as a tunnel interface in the currentNWPath. -
HTTP Headers with the full redirect chain. App Transport Security refuses plain
http://inURLSession, and the http→https hop is the one people most want to see, so that single hop is a raw HTTP/1.1 exchange overNWConnection. ATS stays on for everything else. -
Reverse DNS via
getnameinfowithNI_NAMEREQD, forward-resolving hostnames first and showing thein-addr.arpaname so the answer can be repeated withdig -x.
Limitations
- The released version still doesn't check the response code. 50 MB chunks fit under today's Cloudflare limit, but if it drops, the download would read near zero again. That is the exact failure this post is about, so in the next version a non-2xx answer stops the phase and the speed is marked "not measured" instead of a tiny number.
-
Cloudflare is the load target. If a network throttles or blocks
speed.cloudflare.com, the speeds in the report are low, and they are speeds to Cloudflare, not your plan. - One stream. Chunks go one after another over a single connection. Speed tests usually open several in parallel; on high-RTT lines one stream may not reach the plan's rate, which understates both speed and queueing.
- The phone is the bottleneck on very fast lines. Above roughly a gigabit, Wi-Fi and the device limit throughput before the line does. The score rewards 100 Mbps and up the same way for that reason.
- Health Score is a heuristic. The weights are opinions written down as code. They are public in the app's help screen, and I would rather be argued with over a table than hide it.
- Earlier Bufferbloat grades should be read as upload-only. If you saved results from 1.4.5 to 1.4.8, don't rely on the download rows: I don't know when the limit appeared, and in the latest of those versions they certainly measured an idle line.
Links
- NetDiag+ on the App Store (free; $2.99 one-time premium): https://apps.apple.com/app/apple-store/id6761954529?pt=128748487&ct=devto&mt=8
- What bufferbloat is and how to fix it: https://netdiag.online/guides/bufferbloat/
Have you had a benchmark, load test or health check that quietly stopped doing its job while still producing plausible numbers? I'm curious how you found out. And if you think the Health Score weights are wrong, the table above is there to be argued with.

Top comments (0)