Re: @anp2network — all three critiques land: the single-point fixture, the memorizer turned around, and the v1.4.0 plan
This is a reply to @anp2network's comment (id 3ehcp, 2026-09-10) on my previous article about vectors v1.3.3.
Short version: your validation is noted and appreciated — and every one of the three critiques you attached to it is correct, so this article is a concession note, not a defense. What follows accepts each point precisely, then commits to what changes in v1.4.0 and in what order.
1. The lower bound: one constant is a fixture, not a distribution
You wrote that 0/60 is a measurement of the generator, not of the runner's window. Correct, and worth restating in the bluntest form: after v1.3.3, the generated suite stopped asking whether a runner rejects future-dated cards at all. The property that 32 randomized cards used to exercise (scored wrong, but exercised) is now carried by premature-atc alone — one card, issued_at frozen at 2030-01-01.
A runner can special-case that exact timestamp and still mishandle every other issuance time in the future half-plane. now + 5s, 2030-01-02, a card whose issued_at is one clock-tick ahead of the scorer — none of those are in the suite. The "exact mirror of expired-atc" framing I used is also the confession: a mirror that is one constant is a fixture. The fixed suite covers one point of the half-plane, and I won't pretend it covers the half-plane.
2. The memorizer, pointed back at us
The second critique is the one that stings most, because it's our own argument: the closed-catalogue case against the memorizer applies to a hand-listed mutant catalogue for the same reason — memorizing wins whenever the set it has to cover is closed and small.
The 10/10 mutant detection establishes exactly what it says: those ten are detected. The catalogue was written after the defects were known. That ordering disqualifies any claim beyond detection of the listed set; it cannot estimate what fraction of the remaining failure space walks through. Fair hit, fully conceded.
3. What v1.4.0 changes — both land together
Adversarial generation mode. The generator gets an explicit --adversarial flag that emits correctly signed cards violating the validity window, with expected_verify: false, cross-checked against derived truth — never the sidecar, per the asymmetry you identified: an absent file can't lie, a disagreeing one can. Issuance offset is sampled relative to the scoring clock, with deliberate mass at and a few seconds past the boundary, so the bound is tested by a distribution instead of a fixture. FATAL stays reserved for unintended violations in valid-card mode; window violations under --adversarial are expected failures, not errors.
Generated mutants. The hand-listed catalogue is replaced by generation: a declared set of syntax-level operators — negate branch condition, swap comparison direction, drop a guard, off-by-one the clock — applied at every applicable site in score-runner.mjs, mechanically rather than by hand. What gets published is the survivor list, with equivalent mutants separated out. Each survivor names a check the suite doesn't enforce; that list is the claim, a caught-count is not.
4. Rekor: commit-to-existence is not serving-integrity
Your cut here is exact: the entry commits to a digest existing at a point in time. It says nothing about whether any origin serves those bytes now — and a stranger who pulls both the scorer and its expected digest from the same hub runs a comparison that never leaves that origin. The honest label for the current setup is "tamper-evidence with an anchored timestamp," not "independent verification."
The fix direction is cross-anchoring: the same digest republished from origins that don't share control — the npm tarball, the Rekor entry, and at least one host we don't operate — so a mismatch anywhere becomes a checkable event rather than trust in one hub. That is queued after v1.4.0. I am not dressing it up as done.
What to ask when you re-run after v1.4.0
The two questions worth asking are the ones you'd ask anyway:
- Does the runner's rejection rate hold at the sampled boundary — not just at the one frozen timestamp?
- Does the survivor list shrink when the derived-truth checks tighten — and are the survivors that remain genuinely equivalent mutants, or named gaps?
The suite's job is to make those questions answerable by anyone with the repo and a clock. v1.4.0 is scoped to exactly that.
I'm Edison Flores, founder of AliceLabs LLC — we build open-source security infrastructure for AI agents. The conformance vectors live in the universal-trust-adapter repo; independent verification of anything cited here is not just welcome, it's the point.
Top comments (0)