DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Our OCR endpoint signs a receipt, because the confidence the client sends us proves nothing

When a barcode is not in Munchable's catalogue, the app points you at the ingredients panel instead. The photo goes to an OCR endpoint, the text comes back, the row is saved, and the next person who scans that barcode gets a verdict instead of a shrug. Contributors are paid for it in subscription discount.

So there is a request that creates catalogue data and spends money, and the only things the server receives are strings. A capture POST carries the ingredient text, a source field and a read confidence. All three are typed by the client, and no amount of validating their shape makes them evidence.

Without something more, this is the whole attack: skip the camera, POST plausible text twice from two accounts, reach consensus with no photo behind it, collect the discount. And since other people's verdicts come out of that row, the attack is not only against our billing.

The receipt

The OCR route issues an HMAC over the account, a hash of the text it extracted, the confidence of the read, and an expiry:

v2.<exp>.<conf>.<sig>
Enter fullscreen mode Exit fullscreen mode

The capture endpoint recomputes it from the submitted text. Verified, the row persists and can be credited. Not verified, it is a 422 and nothing is written at all.

export function verifyOcrReceipt(
  receipt: unknown,
  args: ReceiptText & { userId: string },
): ReceiptCheck
Enter fullscreen mode Exit fullscreen mode

Nothing about the receipt is stored. The saved revision keeps a boolean, so there is no table of signatures to leak and nothing to replay out of our own database later.

The useful property is that the receipt binds three things that were previously independent: the account that did the read, the exact text that came out of the model, and the model's own confidence in it. The confidence travels inside the signed blob specifically so the verifier cannot be handed a better number than the model gave.

Version one covered the right bytes and the wrong meaning

The read is not one string. It is three: the ingredients list, the "Contains" statement, and the precautionary "May contain" line. The first hash concatenated them with a separator, which is the obvious implementation and is wrong in a way that took a while to see.

A hash of a + sep + b + sep + c proves that those three strings were read. It does not prove which field each one came from. So a receipt issued for a precautionary "May contain nuts" also verified when the same words were submitted back as the "Contains" line. In an app where people are managing allergies, that is the difference between a warning and a hard stop, and the swap could be made in whichever direction suited the sender.

The fix is to hash named, length-prefixed fields:

function textHash(text: ReceiptText): string {
  const h = createHash('sha256');
  const field = (name: string, value: string | null | undefined) => {
    const v = typeof value === 'string' ? value.trim() : '';
    h.update(`${name}:${Buffer.byteLength(v, 'utf8')}:`, 'utf8');
    h.update(v, 'utf8');
  };
  field('ingredientsText', text.ingredientsText);
  field('allergenStatement', text.allergenStatement);
  field('precautionaryStatement', text.precautionaryStatement);
  return h.digest('hex');
}
Enter fullscreen mode Exit fullscreen mode

The name closes the hole from one side: the digest now commits to a field, not to a position in a list. The byte length closes it from the other: no separator can be smuggled inside a value to fake a field boundary.

The general shape of this bug is worth keeping in your head, because it is not about food labels at all. An HMAC authenticates bytes, not meaning. If your serialisation lets two different structures produce the same byte string, your signature covers both of them, and an attacker gets to pick which one you believe. Canonicalise with field names and lengths, or sign a format that is already unambiguous, and do it before the first production signature rather than after.

Absent fields contribute a zero-length entry, which keeps the useful property that a receipt over the ingredients alone still hashes exactly as it did before the allergen lines existed.

The confidence floor was in the wrong function

"Is this read good enough" is two different questions that happen to be about the same number, and for a while one function answered both.

  • Is it good enough to save? Almost always yes. A low-confidence read is still the only information anybody has about that barcode, the server derives its own structure from the text, and other signals catch a bad row later.
  • Is it good enough to pay for? That is a money decision and deserves a floor.

The floor originally lived inside the verifier, which meant a perfectly genuine receipt for a slightly blurry photo came back as unverified. The app then told a user who had done everything right that their read could not be trusted, which is both false and the most discouraging possible message for the one feature that depends on volunteer effort.

Now the verifier returns the confidence and judges nothing:

const rewardCandidate =
  gate === 'miss' &&
  receipt.ok &&
  receipt.confidence >= OCR_RECEIPT_MIN_CONFIDENCE &&
  !user.isAnonymous &&
  isValidGtin(barcode);
Enter fullscreen mode Exit fullscreen mode

A genuine low-confidence read saves the row and simply does not earn. Every term in that expression is a fact the client cannot forge: the gate that opened the barcode for contribution, the signed receipt, the account kind, and the barcode's own check digit.

Two parsing details that are easy to get wrong

The receipt contains a decimal, so you cannot split on dots. v2.<exp>.<conf>.<sig> looks like four dot-separated parts until you notice conf is 0.910. split('.') over-splits it and then the error you get is a signature mismatch, which sends you looking at the wrong half of the system entirely. One regex for the whole shape instead:

const m = /^v2\.(\d+)\.([01]\.\d{3})\.([A-Za-z0-9_-]+)$/.exec(receipt);
Enter fullscreen mode Exit fullscreen mode

Check cheap things first, then bound the expiry in both directions. A length cap before any parsing, an integer check on the expiry, and then a rejection of expiries that are too far in the future as well as ones in the past, so a receipt cannot claim a lifetime nobody issued. The signature comparison is constant time, and the text hash is always recomputed server-side and never read off the request.

Fail closed means refusing our own feature

If the signing secret is not configured, no receipt can exist. The tempting behaviour is to carry on and simply never credit anybody, since the money is the thing being protected.

That is wrong, because the receipt protects the catalogue too. With no receipt the endpoint is a catalogue any JSON client can write to, and two hand-written POSTs with similar text can reach consensus with no photograph behind either of them. So production refuses every capture with a 501 and persists nothing, and development keeps working with the gap said once in the logs.

A missing secret means "nothing is eligible". It never means "the check is off".

The part I did not expect

Munchable's capture flow has no human confirm-and-edit step. You photograph the panel, the machine reads it, and what the machine read is what gets submitted. The only thing a contributor types is the product name and brand, which are display-only and not covered by the receipt.

That decision was made for data quality: agreement between independent captures of the same product is measured as text similarity, and hand edits blur exactly that signal. Requiring a verified receipt to persist a row turns it into something stronger, because the two properties collapse into one. Every shared row is an unedited machine read of a physical label. Trust in the data and eligibility for the reward stop being separate systems with separate rules, and become the same fact checked once.

What this is not

It is not protection of the photograph. Our privacy policy says label images are read once and discarded, with no image store and no backup of them, and that is why the receipt is over the text rather than over the image: there is no image left to refer back to. The licences page is the plain-language version of the same deal.

It is also not identity. The receipt is bound to the account that performed the read, so one person's receipt cannot be used to credit somebody else's contribution, and it expires.

See the thing it protects

The capture flow is in the app, but everything it feeds is public:

If you take one thing from this: before you key anything valuable on a client-sent field, write down the smallest request that produces it. If that request does not involve the hardware you assumed was involved, you do not have a measurement, you have a parameter.

Top comments (0)