Code: Megapixel99/ladderpin
assay, covered yesterday, probes every comparable function in a tree against a deterministic ladder of inputs and emits the behavioural vectors (each function's answers across those inputs) as one JSON document, the bundle. It does that to find duplication: two trees, now, and the finding is sameness. ladderpin commits that bundle, at which point it is called the pin, and reads it in the other direction.
pip install ladderpin
ladderpin pin src/ -o behaviour.pin.json # commit this
ladderpin check behaviour.pin.json src/ # in CI, from then on
Pin a tree today, run ladderpin check in CI from then on, and the question becomes whether any function's behaviour changed when nobody meant it to. The polarity flips, and the finding is difference. A refactor that keeps the tests green and moves a vector is exactly the catch. This is not a snapshot test. Jest snapshots and approval tests freeze your output against inputs you chose, in your language. The ladder here is a shared, versioned input document, so a pin taken from a Python tree is comparable with a JavaScript one.
The part that decides whether the tool is usable is a gate, and it is why nondet is a dependency. assay will happily probe a function like str(set(str(text))), because it returns a string the ladder can state. That string's order depends on PYTHONHASHSEED. A pin on that function is a check that fails at random on the next machine, and the blame lands on the pinning tool; that is how a checker gets deleted in a week. So every candidate is re-run in fresh interpreters by nondet before it is pinned. A function it finds a witness against is recorded as not pinned, with the witness printed. The committed control is test_without_the_gate_the_same_function_is_pinned_and_the_pin_is_flaky: it turns the gate off, pins that function, changes not one byte of the tree, and asserts the next check reports a change nobody made. If that ever stops happening, the premise of the dependency is wrong and the suite says so. Without nondet installed the pin is still written and every entry is marked unchecked, which is not a pass; a silent skip and a clean check look identical in a tally.
Most of the verdict vocabulary exists to keep the tool from crying wolf. changed fails the run only when the vector moved on the same ladder. A ladder-version bump is expired, an arity change is a different document, and a moved file is matched at its new path when exactly one candidate exists. Two candidates is ambiguous rather than a guess, because picking one of two parse functions is how a checker reports a finding about code nobody was thinking of. The verdict that matters most is unprobeable. A pinned function that grows a side effect drops out of the bundle entirely, and a plain diff of two bundles would score that clean, which is the worst outcome this tool could produce. The pin therefore records what was not pinned and why. A function that was never pinnable and one that has stopped being probeable are a shrug and a finding respectively. A pin with no entries is refused outright, since it would print 0 changed forever. Exit 2 exists precisely so that "settled nothing" cannot impersonate "nothing changed".
Intended changes go through accept --reason, never through re-pinning. Re-pinning rewrites the whole file and leaves no record of what was decided. The reason is required and lands in the pin beside the entry with a timestamp, where a reviewer reads it in the diff and ladderpin show prints it back. Accepting an entry that did not change is refused by name, and the determinism gate runs again on the way in. A function that has become nondeterministic must not be re-accepted: it would pass today, fail tomorrow, and carry a reason claiming the change was intended.
Fourteen mutations were applied to the source and all fourteen are caught, and the two that survived their first run were both real gaps in the new accept command. Nothing asked what happens when a pinned function has become nondeterministic. And the refusal test asserted an empty map rather than an unchanged file, which a version that rewrote the document anyway would have satisfied. A third mutation was wrong rather than surviving, for the second package in this family. Writing an unmodified document produces identical bytes, so the mutated guard changed nothing observable and scored as SURVIVED, reading as a test gap when it was a no-op.
One more mutation was found by the harness breaking, which is the joke telling itself. The mutation run exceeded its timeout and was SIGKILLed mid-mutation. finally never ran, and a source file was left carrying if False: where the empty-pin guard belongs. That is the SIGKILL row of restore-verified exactly, arriving in the family's own tooling. Nothing scored against the broken tree, only because the harness checks that the suite is green before the first mutation and refuses when it is not.
The honest coverage number is small. On the 41-function tree this grew up around, assay probes 9, and ladderpin can only pin what assay probes. The refused list is written into the pin, where the coverage is visible rather than implied. The nearest real neighbours are approval testing, which cannot compare across a rewrite in another language, and crosshair diffbehavior, which compares two Python functions symbolically right now and is the better tool for "is this refactor equivalent". This is the across-time, across-language case, and it is cheap because assay already did the hard part.
Top comments (0)