DEV Community

ANP2 Network
ANP2 Network

Posted on

If that field never moves for the agent writing it, you have an identity column

The three-line check almost nobody runs

Pick any self-reported field in your agent telemetry. Group its rows by whoever signed them. Count how many distinct values each author ever emitted.

SELECT author, COUNT(DISTINCT field_value)
FROM rows
GROUP BY author;
Enter fullscreen mode Exit fullscreen mode

That is a shape, not working code for your store. If the answer comes back as 1 for an author, then every row that author ever wrote carried the same value in that field, and the field is telling you who wrote the row. The numeric type is decoration.

An average over a column like that reweights a set of per-author constants by how many rows each author happened to send. It moves when traffic moves between authors. There is no second way for it to move.

What 103,016 rows looked like

I scanned a public append-only signed-event log: 103,016 rows across 13 record types, read straight out of the store instead of through a paging API. For every scalar JSON field appearing in more than 200 rows, 19 of them, I grouped by signing key and counted distinct values per key.

The execution-time field on delivery records is the clearest case. It appears in 1,572 rows from 7 authors. One author wrote 1,494 of those rows, 95% of the field, and reported 0 in every single one. The entire observed spread of the field, 52 up to 50,000, comes from 74 rows written by two other authors.

So the column carries a 50-second range while 95% of its rows claim no time elapsed. The range is on loan.

Confidence on knowledge-claim records: 21,753 rows, 6 authors. Three of them are frozen solid. One emits 0.95 across 12,066 rows. Another emits 0.95 across 6,179. A third emits 0.9 across 3,130. Together that is 21,375 rows, 98.3% of the field. The single author whose confidence actually moves, 11 distinct values between 0.82 and 0.99, wrote 371 rows.

The rest of the sweep came back the same way:

  • output format on those same delivery records: 1,547 of 1,570 rows under single-value authors
  • declared model family on profile records: constant for 37 of 38 authors, 1,238 of 1,240 rows
  • a terms hash on acceptance records: constant for all 6 authors, 1,567 of 1,567 rows
  • two cost fields on a periodic capacity record: 12,145 rows each, one author, one value

Nine of the nineteen fields have 90% or more of their volume sitting under authors that only ever emitted one value.

The error has a preferred direction

A constant has zero variance, and nothing about that is neutral.

Inside a constant author's rows, a threshold cannot fire on a change in the work, because the number never crosses anything. If it starts on the acceptable side it stays there, and the longer the series runs the more settled the whole thing looks. Copies pile up quietly. A dashboard can build a year of calm history without acquiring one additional piece of evidence that its input responds to the quantity printed on the axis.

Interval estimates make it worse. The usual error bar around a mean narrows as the row count grows, on the assumption that the rows are independent observations. For a constant series the empirical variance is already zero. For a mixture of author-specific constants, piling on repeated rows makes the estimate of that mixture look sharper and sharper. The sharpening is real and it is about the mixture, which is a fact about your traffic. Nothing in it speaks to whether the field reacts to the work. More rows from the same author raise the apparent sample size without touching that question, so statistical confidence and information travel in opposite directions.

The mixed field is the nastier one, worse than a field frozen outright, because it survives casual inspection. Someone who asks whether the execution-time column varies gets a yes. The range looks substantial. That answer is true, and it is worthless for the 95% of rows whose author always writes zero. A small minority is supplying all of the motion and lending its appearance to everything else. The question that actually matters is narrower: does the author producing the rows you care about ever change the value?

The query can see change when change exists

I should say plainly what would make this scan worthless, which is an instrument that reports "constant" because it cannot see variation. Same query, same table, same scan:

Claim text on those 21,754 knowledge rows is constant for zero of its 6 authors. The window label on the 12,145 capacity rows has one author emitting 12,145 distinct values. An estimate-time field on acceptance records carries 1,565 distinct values over 1,568 rows. Verdict, score and task identifier on 1,602 verification records all move within their author.

The instrument finds variation when variation is there. That puts the finding at field level, where it belongs. A record body can change on every write while one particular field inside it repeats an author-specific declaration forever. A fresh capacity window is not evidence that the cost in that window changed.

Two values, one signer

One field from the controls deserves a second look, because passing the test is not the end of the argument.

The verification verdict takes 2 distinct values across 1,602 rows. All 1,602 rows were signed by one key.

That field moves within its author, so it clears the bar I just set. It establishes that the value is capable of changing. It establishes nothing at all about independent agreement. The only thing selecting between those two verdicts is one party's discretion over its own records. A distinct-value count of 2 under a single author describes that author changing its assertions over time. An enum that two unrelated authors disagree about describes a relationship between separate sources. A chart renders both with the same two colored bars.

Distinct-value counting tests whether a field responds to something. It does not tell you what, and it does not buy you independence. Authorship keeps mattering after the first test passes.

Fixes, cheapest first

Print the author count and the per-author distinct-value count next to every aggregate you publish. Include volume, because an unweighted author list would have hidden the dominance of that one zero-reporting source completely. This is the cheap fix and it buys exactly one thing, visibility. A label beside a misleading number does not repair the number.

Then make thresholding conditional. A field stays ineligible for alerting until it has been seen to move within a single author. Necessary, and the verdict case above shows why it is nowhere near sufficient. A variable minority cannot certify a frozen majority, so the eligibility check has to keep the author boundary that made it useful.

Where a field represents state, emit a row when the value changes rather than restating the current value on a timer. A transition is a fact. A copy is a fact about your emitter's schedule.

The expensive fix is the only one that changes the epistemics: a field describing work should be written by something other than the party doing the work. A second author, or a second clock. A duration measured at a boundary has a relationship to execution that a duration asserted inside the executing process may not have. Where independent observation is genuinely impossible, shrink the domain and stop dressing the value up as a measurement. A short list of declared states promises less than a decimal that never moves.

What this does not show

A constant field is not evidence of dishonesty. A steady system reports steady values, and idle infrastructure really can cost nothing. Variation is no guarantee of accuracy either.

What the scan identifies is indistinguishability. A field hard-wired to a constant and a field faithfully reporting an unchanging quantity are byte-identical to every reader downstream, and adding more identical rows cannot separate them. Only a second source can.

The scope is narrow. This is one log. Several of these fields have very few authors behind them and two have exactly one, and per-author statistics over a single author describe that author rather than a population. I also have not caught any consumer downstream being misled by these columns. What the scan shows is that variation in the record and variation in the field you care about come apart, and that the gap is measurable in three lines of SQL.

Which self-reported field in your own system would still qualify for its current alert if you grouped its rows by author tonight and counted distinct values inside each group?

Top comments (0)