DEV Community

pm25coder
pm25coder

Posted on

Your control arm is a probe too

A few days ago I wrote up a test that passed on Windows and failed on Linux, and one of its three lessons was: give every negative probe a control arm. A probe that asserts "nothing is listening on this port" cannot, on its own, tell "the listener is gone" from "I built the probe wrong" — so you stand up a second listener you know is running, and check the probe reports it.

That advice is right. It is also not enough, and I know because my control arm lied to me. This is the follow-up: the control is a probe too, and it is subject to every rule you wrote for the thing it is controlling.

The control that read True for the wrong reason

The probe asks a question — is nothing listening on this port? — and answers by trying to connect:

def nothing_is_listening(port):
    with socket.socket() as s:
        try:
            s.connect(("127.0.0.1", port))
            return False               # the connect succeeded: something is there
        except ConnectionRefusedError:
            return True                # the refusal: nothing is there
Enter fullscreen mode Exit fullscreen mode

The control stands up a listener and asserts the probe reports it:

srv = socket.socket()
srv.bind(("127.0.0.1", 0))
srv.listen(1)                          # the whole bug lives in this argument
port = srv.getsockname()[1]
assert nothing_is_listening(port) is False
Enter fullscreen mode Exit fullscreen mode

With a listener sitting right there, the control returned True — "nothing is listening".

listen(n) does not mean "accept up to n connections". It sets the backlog: how many fully-handshaken connections the kernel will hold in the accept queue on your behalf before it stops completing new ones. The control's listener was one listen(1) socket that nothing ever calls accept() on — a control arm is usually a bare listen() with no server behind it — and the probe was exercised against it more than once, once per spelling under test. The first connect() was handed to the kernel and parked in the queue, filling it. The second connect() found no room, was refused, and the probe dutifully reported "nothing is listening". A listener with a full accept queue is still a listener; the probe just cannot tell.

So the control was measuring the kernel's backlog, not the listener. And note which way it failed: it reported the absence of a thing that was present — the answer that makes the rest of the experiment look conclusive. A broken control fails toward "verified". If the arm that is supposed to return the other answer is itself mis-instrumented, the whole verification quietly stops being a verification.

The repair is to give every arm a listener that stays a listener:

def listener(host="127.0.0.1"):
    srv = socket.socket()
    srv.bind((host, 0))
    srv.listen(16)
    threading.Thread(target=_drain, args=(srv,), daemon=True).start()
    return srv
Enter fullscreen mode Exit fullscreen mode

One listener per probe, each draining its own queue. With that, the control read False for both spellings — a live listener seen as listening, which is what makes the blackhole reading further down the experiment trustworthy.

Narrowing the exception is not enough

A reader read the first post and made the obvious correction: except (ConnectionRefusedError, OSError) is just except OSError spelled longer, because ConnectionRefusedError is a subclass of OSError. Narrow the probe so that only a genuine refusal reads as "gone", and let a timeout or a failed socket() surface as something else. (Measured, it is exactly right: on a host that swallows the SYN, the wide spelling reads the timeout as True — "nothing is listening" — while the narrow one raises TimeoutError and refuses to answer.)

That fixes the probe. It does nothing for the caller one layer up:

# weak caller
try:
    outcome = probe(host, port)
except Exception:
    outcome = "gone"

# category-preserving caller
try:
    outcome = probe(host, port)
except ConnectionRefusedError:
    outcome = "gone"
except OSError as e:
    outcome = f"inconclusive({type(e).__name__})"
Enter fullscreen mode Exit fullscreen mode

Point both at a genuinely dead port and they agree — both say "gone". Point them at a state the probe cannot decide — a SYN that the host silently swallows so the connection times out, or a process out of file descriptors so socket() raises EMFILE — and they part company:

caller timeout EMFILE closed port
except Exception: return "gone" 'gone' 'gone' 'gone'
category-preserving inconclusive(TimeoutError) inconclusive(EMFILE) 'gone'

The first row is the original defect, one layer up: it catches every failure and turns an inconclusive measurement back into a clean shutdown — the wrapper re-collapses exactly the categories the probe had separated. And notice what EMFILE actually is: a fact about the measurer, not the measured. It says the process could not open a socket, not that the socket did not connect. Merge it into the same bucket as "the listener is gone" and the measurement has silently become a statement about the observer.

Here is the part worth slowing down for. A test that asserts "the probe raised an exception" passes for both rows. So does one that asserts "it returned a string". I wrote both and watched them pass against the weak caller. Only an assertion that names the category — assert outcome != "gone", or assert "TimeoutError" in outcome — fails on the wrong one.

An assertion that cannot fail is not evidence

Put the control arm and the caller test side by side and the same shape appears twice:

  • the control returned "nothing is listening" for a listener that was up;
  • the caller test passed whether the caller collapsed categories or preserved them.

In both cases the check held for the right outcome and the wrong one. An assertion that cannot fail is not evidence — and it is the most expensive kind of false assurance, because it is green: nobody re-reads a passing test.

The fixes rhyme with the diagnosis:

  • Run every control in both directions. The moment to trust a control is the moment you can make it fail — so make it fail once, on purpose, and watch it.
  • Assert the thing, not its proxy. "An exception was raised" is a proxy for "the inconclusive path is reachable", and the two come apart the instant someone wraps the probe in except Exception.
  • Run every new test against the code you have not fixed yet. The pre-fix red is the only evidence a test is capable of red. A test that has never failed is a test you have not met.

Checklist

  1. Every negative probe gets a control arm — and the control is itself a probe for the thing it names (listener liveness, not kernel queue depth).
  2. Each arm gets its own resource; a control that shares one is measuring the sharing.
  3. If a probe can return "I could not decide", the caller must have a distinct outcome for it — otherwise the ambiguity is deleted the moment someone writes except Exception.
  4. Assert the category, not the exception's existence: assert raises(Exception) and assert "inconclusive" not in outcome are not the same test.
  5. Before believing any of it, run it against the unfixed code.

Top comments (0)