Most products have a licenses page that nobody reads: a wall of MIT notices, a data attribution line, a credit for an icon set. Ours is four short sections and it credits nobody, which took a while to be comfortable with. The useful part is what that page's shortness did to the architecture behind it.
You can read the whole thing at munchable.app/licenses. The core of it is this:
Munchable builds and maintains its own product database. We curate, correct, and extend the product and ingredient data ourselves so that scans stay fast, accurate, and consistent.
When you scan a product we do not have yet, you can photograph its label to add it. Those entries are built from the physical label and become part of our database, which we maintain and keep up to date. Photos are read and then discarded, never stored.
Three sentences in the middle of a legal page, and the second half of the last one is a deployment diagram.
"Read and then discarded" is a list of things that do not exist
Munchable is a barcode scanner for digestive conditions. Scan a pack, get a verdict worked out on your own device. When the catalog has no data for a barcode, the app offers the camera instead: photograph the ingredients panel, the server reads the text off it, and the row joins the catalog so the next person who scans that pack gets an answer immediately.
The obvious way to build that is a bucket. Upload the image, store the key on the contribution row, run the read against the stored object, keep the original so you can re-read it later with a better model, so you can moderate it, so you can debug a bad parse.
We decided the image lives for exactly one request. The photographs arrive base64 in a single POST, the text and product name are read out, and the response goes back. Nothing writes the image anywhere.
Here is what that one decision removed from the stack. Not deferred, removed:
- No object storage. Not a bucket, not a blob container. Grep the repository for a storage client and there is nothing to find.
- No signed URL issuing, no expiry policy, no CDN in front of user content.
- No lifecycle rules, no image retention job, no nightly sweep for orphaned objects whose rows were deleted.
- No image deletion path in the account deletion flow, and no "delete my photos" endpoint, because there is no photo to delete.
- No images in the database backups, which keeps the nightly dump small enough to be boring.
- No moderation queue of user-submitted photographs, which is a content problem I am extremely happy not to have.
- No question about where in the world the image is stored, which is the kind of question that turns into a paragraph in a transfer impact assessment.
The cost side is real and worth stating. We cannot re-read an old capture with a better model, we cannot eyeball the original when a parse looks wrong, and a bad read has to be fixed by somebody photographing the pack again rather than by us reprocessing. The flow accepts that: if a result looks wrong, the app has a button that flags the row and sends you straight back to the camera, and a new capture supersedes the old read.
For a product whose entire promise is that it does not accumulate information about you, trading a debugging convenience for a whole category of data we can never leak is not a close call.
What survives the request, and what is deliberately missing from it
What persists from a capture is a short list: the barcode, the ingredient text and product name as read, the confidence of the read, the timestamp, and the contributor's account identifier, which is what gives quality checks, contribution rewards and abuse prevention something to work with.
What is not on that list is the thing people assume is on it. Your conditions are not attached to a contribution, and the reason is not a policy, it is that they were never sent. The lookup and the capture routes are both condition-blind by construction; the verdict is computed afterwards on the device. There is no field to forget to clear.
And when an account is deleted, the contribution link is replaced with an anonymous marker while the product data stays in the catalog, so one person leaving does not take answers away from everyone else. The row stops being a fact about a person and goes on being a fact about a food.
Owning the data means provenance is about trust, not credit
The thing an attribution page normally does is answer "who gets credited for this row". We do not have that question, so the provenance fields on a catalog row answer a different one: how much should this row be believed.
A row carries its source and its status, it accumulates revisions rather than being overwritten, captures of the same pack by different people are compared for agreement, reports demote a row, and a re-capture supersedes a read. A row that somebody photographed is not allowed to claim the engine's top confidence badge, and that is a database constraint rather than a code path, because a code path is one refactor away from being bypassed.
That is the inversion I did not expect going in. Dropping third-party data did not reduce the amount of provenance machinery, it changed what the machinery is for. Nobody needs crediting, and every row still has to prove itself.
Legal pages are versioned artifacts, not prose
One more thing the licenses page taught us, which is really about its two neighbours.
/terms and /privacy each have a version number in a TypeScript constant:
export const TERMS_VERSION = '1.3';
export const PRIVACY_VERSION = '1.3';
Account creation records which versions you accepted. When a constant here is ahead of what an account recorded, the re-consent notice appears on that account's next visit, and a script can email everybody still on the old version. Bumping the number is the entire mechanism.
Which means a legal document change is a deploy with a migration-shaped consequence, and it is reviewed like one. When the publisher behind the apps changed name, that was not a find and replace: it changed who the data controller is, so it bumped the version, which notified every existing account.
Look at it yourself
- munchable.app/licenses is the page this post is about. It is short on purpose; the point is that there is nothing missing from it.
- munchable.app/privacy is the long version, with the capture paragraph spelled out: what is sent, what is kept, what is discarded, and the lawful basis for each. It names every processor we use, including the one that reads the label text.
- munchable.app/answers is what the curated catalog produces at the other end: a page per ingredient and condition, each one answered by the fixed rule set rather than by a language model. Is inulin low FODMAP is a representative one.
- munchable.app/terms carries the version constant above in its header, so you can see which number a new account is currently accepting.
If you are building something that takes user photographs, the question worth asking early is not how to store them well. It is whether the request can be the only place the image ever exists. When the answer is yes, a surprising amount of infrastructure turns out to have been optional.
Top comments (0)