DEV Community

Cover image for How long do you keep a published claim provisional?
frank chu
frank chu

Posted on

How long do you keep a published claim provisional?

I retracted a benchmark this week. Two weeks ago I published a timing figure, described it as a fixed cost, and built an argument on it. Today the same script on the same machine ran two and a half times faster, and the reason was that I had taken the original reading while my laptop was under a load average of 70 from an unrelated experiment.

That is the fourth thing I have had to correct this quarter. One was a number I measured under conditions I did not record. One compared two populations of different sizes because a page size truncated a query. One was a retry helper with a budget that did not hold, found by a reader who ran it against a stubbed clock. One was a conclusion about which of my posts get read, drawn from a single data point, which fell over as soon as I grouped the data properly.

Three of those four were wrong on the day I published them. Not superseded by events. Wrong at the moment of writing, with the evidence available to me at the time, and I did not know it.

There is a fifth that did not make it out, and it is the one that changed how I think about this. A draft of mine quoted a throughput figure for a misbehaving retry loop: 187 requests a second. Before shipping it I re-ran the measurement and got a range of 332 to 3896, a twelvefold spread inside a single session, because the original number had been taken while my laptop was under a load average of 70. Nothing about the draft looked wrong. The only reason it was caught is that I now re-measure every number in a queued draft on the day it actually ships, and that habit exists solely because of the four above.

So the question I have been circling is not how to be more careful. I was reasonably careful on all five. It is what status a claim should have after I publish it, and for how long.

Right now I have two states, and I think that is the actual problem. Before publishing, a claim is under review and I will attack it. After publishing it becomes something I have said, and my posture flips from attacking it to defending it, or at least to not thinking about it. That flip happens at the moment of least new information, which is a strange place to put a phase change.

What I notice is that the flip is not really about confidence. It is social. Once something is public, revisiting it costs something, and the cost is not the two minutes of re-running a script. It is the small, stupid friction of having said a thing and now saying a different thing, to people who may have already repeated the first version.

A few ways I have seen people handle this, none of which I have committed to:

Some treat every published number as permanently provisional and version the post, leaving a visible changelog. That is honest but it turns a post into a maintained artifact, and I have maybe eighty of those now, which is not a maintenance burden I can carry.

Some set an expiry: benchmarks are good for a quarter, then they carry a banner saying they have not been re-verified. Cheap to implement. I like this more than I expected to, though it does not help with the three that were wrong immediately.

Some only publish claims they have reproduced under two different conditions. This would have caught my benchmark, since a second run on a quiet machine was all it took. It would not have caught the truncated query, because I would have reproduced the same wrong number twice.

And some just wait for readers, which is what has actually been happening to me. Two of the four were found by people who ran the code. That is a real mechanism and it works, but it only covers claims interesting enough that someone bothers, and it puts the cost on them.

The uncomfortable thing I keep arriving at is that the corrections have been better content than the original posts. The benchmark post was fine. The retraction, with the load average in it and the explanation of why a tight cluster is not evidence of a constant, is a more useful thing to have written. If that is generally true, then the thing I have been treating as a cost is closer to the product, and my instinct to publish carefully and then move on is backwards.

But I do not fully believe that either, because it has an obvious failure mode where you get sloppy on purpose and harvest the corrections. The people I read whose corrections are worth reading are people who were careful first.

Two questions, and I want disagreement on the first one especially:

  1. Does your team or your writing have an explicit shelf life for a measured claim, or does a number stay true until someone complains? I am asking about the mechanism, not the intention. Everyone intends to revisit.

  2. When you correct something publicly, what do you do with the original? I have been editing the live post with the correction inline and the finder's name attached, on the theory that a silent edit reaches nobody who already copied the wrong version. Someone told me that is worse, because it makes the post harder to read for everyone who arrives later and does not care about the history. I do not know who is right.

Top comments (0)