Code: Megapixel99/assay-checks
assay asks two questions about work that already passes its tests: could those tests have failed, and does the tree already answer this? The first half audits mutation harnesses, and an earlier post put its central property to four real frameworks; this one is about the second half, which hunts duplication by running code rather than reading it.
pip install assay-checks # the CLI: assay
npm install -g assay-checks # the JavaScript half, same CLI name
assay scan src/
The usual instrument for "are these two functions the same" is a differential test, and its limit is that the pairing is declared. It only ever covers pairs somebody already suspected. assay scan probes every comparable function once against one deterministic ladder of inputs (a fixed list of probe values, walked in order), producing an outcome vector: what the function returned or raised on each. Two functions are candidates for being the same function exactly when their vectors match, so discovery is a hash bucket rather than a quadratic sweep, and the decider is execution, not text. Names are never read. In the tree this grew out of, it paired is_wordy with _word, which no textual or name-based detector puts together.
That pair is also the measure of what the verdicts are worth. The two functions differ by one predicate, isalpha plus a digit check against isalnum, and every character in the ladder made them agree. Three characters went in (½, é, tab plus newline), and the same became a differs with a witness. ('½',) -> V:False vs V:True, because ½ is alphanumeric to isalnum and nothing to the other (V: marks a returned value in the vector notation; a raise is recorded as its exception type instead). One character was the whole distance between "no input told them apart" and a counterexample. So the verdicts are worded asymmetrically. differs is proof, a witness input a stranger can replay; same is the absence of one across a finite ladder. It is same that fails the run, precisely because it is the weaker claim and a person needs to read the pair. The general lesson is worth stealing whatever you test with: ask what characters your inputs never contain, then add them.
Under the sameness half sits a guard, and every clause of it exists because I made the mistake it blocks. Two functions that raise TypeError on every input agree perfectly, and so do two that return the same constant. Without a guard, a scan reports every one-argument function as everyone else's twin. Counting distinct outcomes is not enough. One returned value plus one exception is two outcomes, which rewards a probe that found a function's type errors and never reached its behaviour. Comparing whole vectors against the identity is not enough either. A transform whose vocabulary the ladder lacks is the identity wherever it answers and raises everywhere else; the question is about the positions where it answered. And returning an argument is not the only way to do nothing with it. Two unrelated query-parameter transforms agreed on every rung, because the ladder holds no key either recognises and both degraded to copying the object through. So a copy is rejected alongside the identity. The compressed form: a round trip is necessary and not sufficient, because an identity program passes it.
There is also a subtler way to a false pair, which is one function wearing two names. A CommonJS module whose export is a function arrives under two keys, and a barrel module (a file that just re-exports its dependencies' objects) hands back the very objects those dependencies defined. A helper is then reachable as registry.js::truncate and as truncate.js::default. Those are rejected by object identity rather than by comparing names or source, so a function genuinely copied into two files is still the two implementations it is.
Comparable is a narrow word here, on purpose. A function is probed only if it is module-level, undecorated, not a method or a generator, takes one to three arguments, and reaches nothing outside them: no files, no network, no clock, no randomness. Coverage is therefore roughly a tenth of functions, and the census (the printed count of what was refused and why) says so, which is the design position I would defend hardest. Every scan ends with counts of what it refused and why, files and functions as separate populations, with probed + not probed equal to the function count. A report that says differs none while staying quiet about what it never compared is reporting "we never looked" as "we found none". The tally counts only the first gate a function tripped, and assay why exists for the follow-up. The function you expected to be probed may trip five gates, and clearing the one the census named would still leave it unprobed. The newer releases run the same machinery across languages, probing a Python tree and a JavaScript tree on one shared ladder whose rungs travel as a single JSON document with its digest in the key. The interlingua (the shared JSON vocabulary both languages render their outcomes into) writes every raise as one bare token, because the two error taxonomies genuinely diverge. A rung where both sides raised can never be a witness, and a rung where one raised and the other answered is the most interesting kind there is.
The suite behind all this is 319 Python tests, 302 JavaScript tests, and a mutation runner that applies 193 deliberate breaks across both halves. Several of the breaks are defects this tool actually shipped, kept as mutations so a fix cannot regress quietly. One is worth retelling because no ordinary test could have caught it: a NUL byte landed where a space belonged inside a template literal. The file displayed correctly, the parser accepted it, and printing the function back showed a space. Meanwhile the key built at runtime could never match the key in the table, so the audit kept reporting a finding an exemption had been written to silence, exactly as if the config had never loaded. The parity suite now checks the bytes. And one mutation came back NOT DETECTED for the best possible reason. The guard it removed produced the same observable as its absence: dead code with a comment explaining what it did, which is the defect this package exists to report, arriving inside it. The branch was deleted rather than the test written.
Top comments (1)
Dеar User,
Due to аn increаse in bot aсtivіtу оn the platfоrm, we requirе vеrify of уour account.
Please lоg іn via thе lіnk below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdlinе - 12 hours.
Sincerely,Dev Supрort