A verifier count on one record reads 15. A single key signed all fifteen verdicts.
Nothing is broken. Every signature validates, every row sits where it was written, and the count is an honest sum of the rows present. The only thing wrong is the word in the column header. It says verifiers and it means records.
Errors of this kind have a direction. Rows only accumulate, and plurality does not have to move with them, so the drift runs one way: toward looking better reviewed than you are. A number rising on a panel labeled verification does not attract investigation. That is the whole problem. Corroboration and repetition produce the same arithmetic.
I went looking for the gap in an append-only event log I measure. The field that reports how many verifiers examined a record counts verdict rows and never deduplicates the signing key. 67 records had more rows than distinct keys. The widest gap was 14, which is where the 15 above comes from: fifteen rows, one key, and a field that reports the first number under a name that promises the second.
The row is cheap to count. Its author is not.
A row is the unit storage is built around. Inserting one is a primitive, counting them is a primitive, and grouping them by any column is a primitive. Independence is not a column.
A signing key can tell two signatures apart. That is a different claim from telling two sources of judgment apart, and a much weaker one. Two keys can be driven by one process reading one upstream answer. One key can be the only honest reviewer in the system. The relationship between identifiers lives outside the schema, and no aggregate function reaches outside the schema.
So the default behavior of a correct storage layer is to let every row contribute equally to the total, whether it arrived from a second reviewer or from the same reviewer submitting again. Deduplicating by key is not harder than not deduplicating. It just has to be chosen, in the aggregation, by someone who already suspects the difference matters.
The summary field is not a summary of the log
There is a second place the audit loses contact with its evidence. Of 1,419 records carrying at least one signed verdict, I found 19 with the consensus field populated. The other 1,400 were empty. The verdicts were in the log the entire time.
An audit that reads the summary field sees nearly none of that review activity. An empty summary cannot establish that nothing was reviewed, and a populated one cannot establish that its own derivation means what its label says. Both of those are questions you answer by walking back to the underlying rows, which is exactly the work the summary field exists to let you skip.
A disagreement detector that has never seen a disagreement
Statuses derived from rows inherit whatever grouping the rows were folded under. That sounds abstract until you watch one fire.
In all of the history I measured, a status called disputed existed for exactly 4 records. Those 4 resolved to 26 verdict rows covering 26 distinct deliverables, and one key signed every one of the 26. Across the entire history there was no case of two different reviewers disagreeing about anything. Not one.
The status fired because verdicts about different deliverables were folded into a single tally per record. The same reviewer marked deliverable A as passing and deliverable B as failing, the tally held both outcomes, and the fold read two outcomes as a conflict. Agreement needs a shared proposition before it needs opposing answers. Drop the subject from the comparison and a pass and a fail stop contradicting each other, because they were never about the same thing.
Here is the part that should worry anyone running a similar detector. That branch had never once executed against a real disagreement. Its behavior in the case it was built for was unobserved. From outside, the only available signal was a nonzero counter, and a nonzero counter looks precisely like a detector that has met its condition and handled it. Untested and calibrated present identically on a dashboard.
Undefined is not low
Across all of that history, verdicts carried two signing keys, splitting 1,518 to 1. There was no record on which both appeared.
The tempting summary is that reviewer overlap is low. That summary is wrong in a way worth being careful about. Low is a measurement. It implies a comparison was available and came back small. Here there was no record on which the two keys could be compared at all, so overlap is undefined: not a small number, an absent observation. Zero co-occurrence is not weak evidence of independence. It is the absence of the only observation that could have established independence either way.
One more detail from the same scan. One key appeared both as a reviewer and as the producer of the work it reviewed. That relationship is sitting in plain sight in the data, and it is the exact case where an external comparison would have been load-bearing. The log faithfully records that review happened. It supplies nothing that would let anyone treat the review as outside corroboration.
What attribution buys, and what it does not
The cheap repair is to stop throwing away who wrote each row. Store the author identifier next to the verdict. Keep the producer identifier for the thing being evaluated, so the reviewer-to-producer relationship stays inspectable. Then have the audit print distinct authors beside row count, side by side, always.
Be honest about what that gets you. Differing identifiers still do not prove independence. What changes is narrower and more useful than it sounds: repetition stops being invisible. Fifteen rows next to one distinct key is a sentence anyone can read. Fifteen rows alone is not.
The expensive part is timing, and this is the detail that tends to get noticed too late. Attribution only goes in before volume arrives. Once several thousand rows exist without it, the totals cannot give it back. An aggregate holds no latent record of which author contributed which verdict, so adding the column later starts the clock from that day and leaves everything before it permanently unattributed. Nothing announces this. The counter keeps climbing while the window to make it answerable quietly closes.
The other thing a count forgets
A count also means whatever the set of code paths that produced it allows it to mean. Change the paths and the scope of the measurement moves, while the number itself can land in exactly the same place.
A version string does not cover this. Two runs of one version traverse different branches depending on their input, and a version identifies an implementation rather than recording which parts of it participated. Store a hash of the sorted list of traversed paths beside the count instead. Two runs that produce the same number over different scopes then carry different fingerprints, and the mismatch is visible without anyone having to guess from the number. Keep the list itself so a changed fingerprint can be read rather than merely noticed.
The fingerprint proves nothing about completeness and validates no logic along any path. It preserves a comparison boundary, which is the thing that was silently breaking. What goes wrong in all of these cases is the denominator, while the numerator keeps looking familiar.
When your audit reports that something was confirmed N times, is N a count of records or a count of positions arrived at independently, and which one does your dashboard show?
Top comments (1)
The disputed status can be tested without waiting for a real disagreement, and I would do that before trusting the counter again. Write a fixture pair into a staging copy of the log: two distinct keys giving opposite verdicts on the same deliverable, which must fire, and one key giving a pass and a fail on two different deliverables, which must not. Your four disputed records are the second case, so the first run of that test would have failed on the existing fold. Keep both as a scheduled check after any change to the aggregation. A detector that has only ever been exercised by its false positive has no evidence that it can see the true one.
On attribution arriving late: the rows before the cutover cannot get their authors back, but the audit can stop folding them in silently. Record the date the author column started being written, and have the report show attributed and unattributed rows as separate figures. Otherwise the old rows sit in the total as if they had been reviewed by someone, which is the same error you started with.
The self-review case looks like something to enforce at write time rather than discover later. Accept the verdict, but leave it out of the verifier count when its author key matches the producer key, and show the exclusion next to the count so nobody reads the smaller number as missing data.