What a refractometer taught me about observability
In July, I wrote about a farmer walking a vineyard row at dawn, crushing a grape onto a refractometer prism, and typing "13.5 Brix" into a chat window. The agent on the other end didn't care how the number arrived. To the Digital Scribe, I wrote, a number is just a number.
In the vineyard, that's a defensible position. Fresh grape juice is exactly what a refractometer is built to read.
This fall, I've been pointing the same kind of instrument at juice that's fermenting. It turns out the agent should have cared.
Three batches and a prism
Over the last several weeks I've started three small batches: a blackberry wine that's now bulk aging, a pear wine that began fermenting on September 24, and a fireweed honey mead that got its yeast two days later. They're one-gallon batches, which matters for this story, because at that size every sample you pull is wine you don't get to drink and oxygen you've let in.
So I track them with a refractometer. A few drops of liquid on a glass prism, close the cover, hold it up to the light, read a number off the scale. The number is Brix, roughly the percentage of dissolved sugar. It's fast, cheap, and costs a few drops per reading. During the blackberry's primary fermentation I took readings twice a day, every time I punched down the cap of fruit.
And once fermentation starts, the refractometer is wrong.
The instrument isn't lying, exactly
A refractometer measures how much light bends as it passes through a liquid. Dissolved sugar bends light, so in fresh juice the amount of bending is a good proxy for the amount of sugar. The scale on the instrument is calibrated on that assumption.
Fermentation breaks the assumption. Yeast converts sugar into alcohol, and alcohol bends light too. A few days into primary, the refractometer is reporting the combined effect of the sugar that remains and the alcohol that's been produced, and presenting all of it as sugar. The reading comes out higher than the real sugar content. Taken at face value, it says the fermentation is further behind than it actually is.
To get something usable, you run each reading through a correction formula. Most home winemakers use calculators built on formulas originally developed for brewing, and every one of them needs the same extra input: the original reading, taken before the yeast went in.
Here's what that looked like on the blackberry:
| Day | Raw refractometer (Brix) | Corrected value |
|---|---|---|
| 08/29/2026 | 20.7 | 20.7 |
| 08/30/2026 | 18.4 | 16.5 |
| 08/30/2026 | 17.0 | 14.3 |
| 08/31/2026 | 13.2 | 8.1 |
| 08/31/2026 | 10.8 | 4.1 |
| 09/01/2026 | 8.7 | 0.7 |
| 09/01/2026 | 8.3 | 0.1 |
Nothing is malfunctioning. The instrument is doing exactly what it was built to do. What changed is the system it's pointed at, and a number that was a reliable proxy in one state became a misleading one in the next.
Neither instrument measures sugar
The traditional alternative is a hydrometer, a weighted glass float. The deeper it sinks, the less dense the liquid. Sugar makes liquid denser, so falling specific gravity means sugar is being consumed.
Except the hydrometer doesn't measure sugar either. It measures density, and alcohol is less dense than water, so as fermentation proceeds the alcohol pulls the reading down on its own. A finished dry wine routinely reads below 1.000, lower than plain water, which would be impossible if the hydrometer were really a sugar meter.
So the refractometer measures how light bends and the hydrometer measures how heavy the liquid is. Neither measures what I actually care about: how much sugar is left, whether the yeast is still working, whether this wine is done. I infer those. The instruments give me evidence.
Even a steady reading is ambiguous. The traditional sign that fermentation has finished is the same reading several days running. But a fermentation that has stalled also produces the same reading several days running. Mead is notorious for this: honey has very little natural buffering, the pH drifts down as fermentation proceeds, and the yeast can slow or stop well short of dry. A flat line can mean done or stuck, and the instrument can't tell you which. You tell them apart with other evidence: where the line flattened, what the pH is doing, what you expected to see.
That distinction between evidence and state sounds pedantic until you notice it's one we get wrong in software constantly.
Your dashboard has the same problem
CPU utilization isn't health. A service at 15% CPU can be deadlocked, and one at 90% can be doing exactly what you want. An HTTP 200 isn't correctness; it tells you the server returned a response, not that the response was right. A green health check tells you the health check endpoint answered. Model confidence isn't truth; it's a number the model produced about its own output.
Each of these started as a reasonable proxy under a particular set of conditions. Then the system changed, a new failure mode appeared, or someone started optimizing the number itself, and the proxy drifted away from the state it was meant to represent. The dashboard kept rendering it with exactly the same authority.
That's the refractometer problem. The metric was calibrated for one state of the system and is now being read in another, and nothing on the display tells you so. And like the flat fermentation line, a steady metric can mean two opposite things: a quiet error rate might mean a healthy service, or one that stopped receiving traffic an hour ago.
The failure isn't collecting the metric. It's treating the metric as the state instead of as evidence about the state.
A reading needs its history
Back to that correction formula. To interpret today's reading, you need the reading from before fermentation began. Without it, the current number is close to uninterpretable. The same Brix value could describe a fermentation that has barely started or one that's nearly finished, depending entirely on where it began.
The meaning of a measurement depends on its lineage.
This is where I think observability practice most often falls short. We store the value and the timestamp and call it a record. What we usually drop is everything needed to interpret it later: which instrument produced it, under what conditions, what calibration assumptions it carried, what baseline it should be read against, and what state the system was believed to be in at the time.
My fermentation log keeps the raw reading and the corrected value side by side. The corrected value is what I act on. The raw value is what I re-derive from if I later learn the correction was off, find a better formula, or discover I misread the original. A corrected value on its own is a conclusion with its evidence thrown away.
That's a pattern worth stealing. Store the observation separately from the interpretation. Keep the provenance that makes the observation meaningful. Let state be something you derive, with its evidence attached, rather than something you overwrite. It's the same idea behind the forensic receipt work I've been doing: a claim should carry what it was based on.
Watching isn't free
One more thing the one-gallon batches have taught me. I use a refractometer because a hydrometer needs a much larger sample, and in a small carboy that sample is a real fraction of the batch, and pulling it lets oxygen in. The act of measuring changes the system being measured.
It also means deciding how often to look. The blackberry had to be punched down twice a day anyway, so twice-daily readings were free. The pear and the mead don't need that kind of handling, so I've settled on a reading every couple of days: often enough to catch a stall, rare enough to leave the batch alone. That's a measurement policy, and it means my log has deliberate gaps in it. Those gaps are their own story, and I'll come back to them.
Software has the same trade-off: tracing overhead, probes that add latency, sampling that shifts timing. Usually the effect is small. Sometimes it isn't, and how often to look is a design decision with costs on both sides.
Evidence, then state
Here's what I've taken from a few weeks of squinting at a prism.
Separate what you observed from what you concluded. "The refractometer read 9 Brix" and "fermentation is about two-thirds done" are different kinds of statement. They belong in different places, and they should be allowed to disagree.
Keep the raw reading. Corrections, normalizations, and aggregations are interpretations. If you keep only the interpreted value, you can never revisit the interpretation.
Record the conditions that make a reading meaningful. Instrument, baseline, calibration assumptions, the believed state of the system. A metric without its context is a number waiting to be misread.
Treat state as a derived claim. "Fermentation is complete" isn't a reading. It's a conclusion drawn from several readings over time, plus whatever evidence rules out "stuck," and it should be traceable back to all of it.
Notice when your proxy stops being a proxy. Every metric was calibrated for some range of system behavior. When the system leaves that range, the metric keeps reporting with exactly the same confidence.
The agent in July was right about one thing: it shouldn't matter whether a number comes from a clipboard or a five-thousand-dollar probe. Where it was wrong was in believing the number could travel without its history.
The pear wine is still in primary as I write this. The raw refractometer reading says it has a long way to go. The corrected number says otherwise, and the corrected number is only trustworthy because I wrote down what the juice read before the yeast went in.
The number on the instrument isn't the system. It never was.
Top comments (2)
hold up. green settlement tiles are not a signed hop tip.
1 cut: when chargeback week opens, can anyone GET the queryable tip after the vendor UI flips, or only another dashboard seal?
receipts > seals. marker0929h2128-dt
I think I follow the distinction you're making: a green settlement status is an interpreted state, while a durable, queryable receipt preserves the evidence behind that state. If so, that maps closely to what I'm arguing here: keeping the observation separate from the conclusion.
I'm not familiar with “signed hop tip” or “dashboard seal,” though. What do those mean in the system you're describing?