DEV Community

build996
build996

Posted on

Three numbers I published about a web host that were really about something else

I was comparing four free PHP hosts by deploying the same small app to each of them and measuring it. In one afternoon I produced three numbers that were accurate, reproducible, and about the wrong thing. Two of them went live before I noticed. I corrected both within the hour, but I want to write down how each one fooled me, because the pattern is general and I suspect most people who benchmark anything have shipped one of these.

1. A "21% failure rate" that belonged to the speed test

I ran one of the hosts, FreeHostingEU, through a free multi-location speed test that probes a URL from 33 nodes. It came back 26 OK, 7 failed. Seven out of thirty-three is 21%, and I wrote it up as a reliability finding: roughly one visitor in five can't reach this host.

Then I ran the other three hosts through the same test. Every one of them came back 26 / 33. Same seven nodes failing, every time.

Those seven nodes are broken on the speed-test side. The number was a property of the instrument, and it would have come back the same for any URL I gave it. Written as a fact about the host, it was simply false.

The check that would have caught it: measure a sibling before you publish a number about one target. If the number is identical across targets that should differ, it isn't describing the targets.

2. A 1.2-second DNS delay that belonged to the domain name

Same host, same test. The median page load was 2.123 s, and when I split it into phases, 1.239 s of it was DNS resolution. I attributed that to the host: a European server, a long way from the test nodes, slow.

The next host I looked at, AwardSpace, turned out to run on the same platform. Same datacentre, and its IP address is two numbers away from FreeHostingEU's (.98 and .100 in the same range). Its DNS time from the same test: 0.027 s.

A 46× difference between two machines sitting next to each other can't come from the machines. The only thing that differed was the name. The free plan on one brand gives you a subdomain under eu5.net; the other gives you one under atwebpages.com. The first zone resolves slowly from those test locations, and the second doesn't.

That turned the finding inside out. What I had written was "this host is slow". What the data actually said was "the free subdomain is slow, and if you bring your own domain the 1.2 seconds goes away". The first version would have steered a reader away from a host for a reason they could fix in five minutes.

The check: decompose the total before you blame anything. Total load time alone can't be attributed to anyone. DNS, connect, TLS, first byte and download live in different layers owned by different parties, and the split between them usually points straight at the one that varies.

3. A fast page that wasn't the page

This one I caught before publishing, mostly because I was primed by the first two.

The same speed test gave InfinityFree a median of 0.961 s, while AwardSpace, the host I had just defended, came in at about 2.0 s. On the headline number InfinityFree was twice as fast.

The tell was a column I had been ignoring: the size of what was downloaded. AwardSpace's row said 8 KB, which matched the real page. InfinityFree's said -KB.

InfinityFree serves a JavaScript challenge page to clients it thinks are bots, and a speed-test probe is exactly what it thinks a bot is. The probe timed how fast it received a small interstitial, not the page. The measurement was real. The thing measured was a doorway.

The check: before trusting a timing, find the column that proves what was fetched. Size, status code or a string from the body will do. A response that's fast but empty should be read as a warning.

What I didn't do, and what I do now

The tempting reaction after the third one was to declare those hosts "unmeasurable" and leave the numbers out. I said roughly that to the person I was building this with, and they pointed out that a visitor who gets a slow page gets a slow page, whatever the reason. They were right. Refusing to measure would just have been a more respectable way to be wrong.

So the rule I use now isn't "don't publish numbers you can't attribute". It's:

  1. Run a sibling. Same method, same path, a different target. Anything that doesn't change across siblings can't tell them apart.
  2. Name the layer. Instrument, network path, DNS name, server or application. Then ask whether the layer I'm about to blame is the one that actually varies.
  3. Prove what was fetched before trusting how fast it arrived.
  4. Publish an unattributed number as a number. "Median 2.1 s from these nodes, 1.2 s of it DNS" is honest. The error is inventing the cause, not reporting the measurement.

None of these is sophisticated, and all three mistakes would have been caught by the first one. I only ran the sibling because I happened to be testing four hosts anyway. With one host I'd still be telling people it fails for one visitor in five.

Top comments (0)