DEV Community

Rain S
Rain S

Posted on

When a Record Has No Word for "Unknown", an Empty Value Will Say "Everything Is Fine"

Something happened recently that I want to start with.

We fixed a defect: four different kinds of "nothing" were collapsed into a single switch, so the readings couldn't tell "by design" from "broken". Then we moved on to the next thing. The day after that new field landed, a reader pointed at it in a comment, named the same defect again, and wrote down what to do about it.

This isn't the same bug coming back. It's the same class of bug growing back in the place where we thought we had learned it.

This piece is about one thing: when a record has no slot for "unknown", "unknown" reads as "known" — and specifically as the strongest statement that field can make.

  1. Four kinds of "nothing", one switch Background first.

When the AI writes a record, we keep a snapshot of what it looked like before and after. Revocation needs to know where to go back to.

The trouble was in "we couldn't capture the after-snapshot". There are at least four reasons that happens:

no captor was wired at all — an optional piece by design
the entity didn't resolve, and the fallback data was empty too
the entity resolved but the row is gone — an anomaly
the capture threw — an error
The first two are what the design chose. The last two are something went wrong.

We had one boolean field for all of it. Four reasons, one switch.

So when someone asked "why are these snapshots empty", there was no answer: you couldn't tell by design from broken. The worse part is the direction. That switch defaults to false, which reads as "the snapshot is fine".

"Nothing" is not one value, it is four different values. Collapsing them into one boolean lets the strongest reading answer for all of them.

We split it into four causes, each stored and readable on its own. While splitting it we also found that two of the four names we had first proposed didn't belong in that class at all. One lives on a different path (the table consulted before the write), and the other is a mechanism rather than a cause. So the four that landed are not the four we started with.

  1. The next day, in the field we had just written With that fixed, we moved to the next thing: before a revoke, record which members this attempt intends to compare.

The promise matters here. Before the revoke it's a promise; after it, an actual. Only the difference between them lets you say "a member I promised to compare turned out not to be comparable". The implementation put a column on the group's root row to hold that difference. If this attempt found no difference, the column is written empty.

That looked fine. Then someone wrote this in a comment:

A per-reason column sits on the member row and gets rewritten by whichever revoke attempt ran last, so a second attempt erases the first attempt's broken promise. Append-only keeps them apart.

He's right, and it's slightly worse than he put it.

When the attempt finds no difference, that column is written empty. And that empty means two things at once: "the last comparison found no difference", and "this row was never revoked at all". So:

the first revoke found "promised, but couldn't compare" — recorded
the second revoke went fine — and wrote over the first attempt's record with an empty value
reading that row later: empty. You can't tell "there was once a broken promise", and you can't tell "it was cleared"
A retry that succeeds erases the earlier attempt's failure record.

To be fair to ourselves: that clearing is deliberate. The reasoning is in the design record from when it landed — a stale difference shouldn't outlive the comparison that produced it. The reasoning isn't wrong. What's wrong is that the empty value had no second meaning available.

Then we changed it to what he described.

One row per attempt that reached the comparison, instead of one shared cell. The two readings are separated where they are written: this row is either "a difference was found" or "checked, and there was none". So:

a later attempt no longer erases an earlier one — it writes its own row
an empty history (no attempt ever reached the comparison) and "checked, clean" are no longer the same reading
the old rule — a stale difference shouldn't outlive the comparison that produced it — was withdrawn. It was wrong about using one cell for two things, not about keeping old records
The old column wasn't dropped. It's still read, as a legacy entry marked "written at a time we can't know". Deleting it would erase exactly the evidence this whole argument was about.

  1. One place in the same system got it right Both of the above are "no slot for unknown". But one place in the same system does have a slot.

When a revoke goes out to an external system, all we can do is send the request; the terminal state is on their side. Their 2xx only proves the request arrived, not that they undid anything. So that row doesn't write "revoked". It writes "compensation requested, outcome unknown" — a dedicated reading.

The effect is immediate: nobody misreads it as done. The interface says plainly that the result is with the target system. No gloss needed.

So this isn't impossible. It's a question of whether you noticed that "unknown" needed a slot. Give it one and it stays put. Don't, and it falls into the strongest reading available.

  1. Why it's hard It's hard because the default is silent.

Every field decides for you, and it always picks the cheapest reading:

a null value reads as "no such item"
an empty collection reads as "no difference"
false reads as "not marked", and from there as "fine"
Those readings are correct almost every time. The problem is the rest: when "nothing" and "unknown" share a value, you have a field that lies — and it stays quiet while doing it, until somebody acts on it and gets "everything is fine" on your behalf.

Worse: fixing one instance doesn't remove the class — including while fixing it.

Here's something that happened inside that fix. Once "no difference" had its own reading, there was still a function whose empty return meant two things: everything really was compared, and there was nothing to promise in the first place (with no captor wired, "compare" isn't defined). Writing "checked, clean" is right for the first. Writing it for the second would vouch for a check that never ran.

We nearly did. What we changed it to: nothing was promised, so no row is written.

That's why the entry points for this class are every "nothing" there is. Fix one, and it waits for you in the line of code where you write the fix.

  1. A check you can run yourself Don't start with "is my record complete". Start with something more basic:

List every value in your system that means "none / unknown / empty", and ask of each: what does it read as?

Anything that carries two meanings at once is a potential misreport. Three common shapes:

Shape It also means
a null value "no such item" / "looked, found none" / "never looked"
an empty collection "genuinely none" / "couldn't read it out"
false / a default "not marked" / "marked as no"
The test is simple: when this value is empty, can a reader still say why it's empty? If not, the field is guessing on their behalf.

Closing
This kind of problem is hard to find because it doesn't throw. The system runs, the interface renders, and one cell quietly says something untrue.

And the place it shows up most is the place you just fixed — because that's when you're busy already knowing how to do it.

Where we stand: we know the pattern now, and we've fixed two instances — the second one's shape came from the reader. We have not done a systematic pass: how many of those three shapes are in our record layer, and what each one reads as, is not a list we have.

So this one doesn't close on a conclusion. Run the check against your own system and you'll likely find a few. On our side, we only know we haven't finished looking.

Drafted with an AI assistant. The system, the positions and the mistakes are mine.

Top comments (2)

Collapse
 
chrissellers profile image
Chris Sellers •

The line that stuck with me is that a retry which succeeds erases the earlier attempt's failure record. Outbound tool calls have the same shape: the first request times out after it was sent, the retry succeeds, and a shared status cell ends up saying ok with no trace that the first attempt may also have landed. One row per attempt is the right fix there too.

Your "compensation requested, outcome unknown" reading is the one I'd copy everywhere. I'm building Agent Middleware (still pre-revenue), so I'm biased, but we landed on the same idea: when the response is lost after the send, the call is recorded as delivery uncertain rather than failed, and it is never re-sent automatically. A retry with the same key returns the original receipt instead of making a second call.

The part I haven't solved cleanly is who gets to turn an unknown into a known later. Does your reconciliation step overwrite that row, or does the resolution become its own appended row?

Collapse
 
rain6fish profile image
Rain S •

On the question: both, in different places.The status column is current state and gets overwritten. What doesn't get overwritten is the audit. Every attempt writes its own entry, and the code says outright that a failed attempt has to be recorded too, rather than leaving a blank. So the history lives in the chain and the current state lives in the column.

That split held right up until it didn't. The thing this piece is about is exactly the case where the audit entry wasn't enough: the comparison a revoke promised to run left evidence that belonged on the row rather than in a log, and one shared cell couldn't hold both "checked, clean" and "never checked". So the rule we ended up with isn't that the audit covers it. It's that evidence which has to outlive its attempt needs its own append-only home, and the status column stays current state.

On who turns an unknown into a known: nobody, automatically. We don't poll and we don't reconcile. When the target system owns the terminal state, our row says requested-and-unknown and stays that way, and the interface says the same thing. That's a deliberate refusal rather than an unfinished feature, but it has a cost we should name: a person has to close it, and until then the row is unresolved.

Your outbound case landed closest, and it's narrower than you'd expect. Revocation separates "we don't think it arrived" from "it arrived and nobody answered", and so does one of our two outbound paths: the external tool path says outright that a request may have reached the target and that nothing will be retried, and it holds the idempotency key on failure. The proxy path doesn't. There, everything that isn't a timeout reads as "unreachable", which asserts the request never arrived, and a connection reset after the send is exactly the case that assertion can't support. So the shape you describe isn't hypothetical here. It's one word, in one path.