Menu digitization fails quietly when an image looks fine to a human but carries the wrong orientation, color profile, or dimensions for downstream search. Short answer: inspect metadata at upload, reject or normalize unsafe files immediately, and defer expensive cleanup until a dish is actually indexed or viewed. That split keeps the upload SLO predictable without making every menu photo wait for a full processing pipeline.
I model the incident as a bounded lunch-rush upload: a restaurant sends 240 phone photos in ten minutes, including a 12 MB JPEG rotated through EXIF and a PNG with an alpha channel. My first job is not to make either image beautiful. It is to decide, with a small and explainable gate, whether the bytes can enter the searchable dish record flow. A bad decision here becomes a support ticket later: a sideways thumbnail, a blank transparent background, or OCR that reads a dish name from the wrong crop.
The invariant is simple: metadata is a control signal, not decoration. Store the original object, record the inspected facts, and make every transformation reproducible from that record. If inspection cannot establish a safe type and bounded size, return a client error such as 422 and keep the object out of indexing; do not silently guess. In practice, that means the intake service writes an immutable original key and a small inspection row in the same workflow, then publishes a cleanup job only after both are durable. The row contains the source hash, dimensions, detected format, policy version, and a state such as accepted, rejected, or ready. A retry sees the same hash and policy and can decide that work is already complete instead of decoding the photo again. When a phone sends a JPEG whose EXIF says “rotate 90,” the row records that fact and the canonicalizer applies it once; the search index receives the normalized derivative, while an auditor can still retrieve the untouched upload. When the decoder cannot read the bytes, the response is a clear 422 with a stable reason code, and the object is quarantined under a retention policy rather than being fed to OCR. This is deliberately boring state management, but it gives on-call engineers a bounded question to answer: did the gate reject the file, did the queue delay the derivative, or did indexing lag after the derivative was ready?
What should metadata inspection decide before image cleanup?
At upload, inspect the declared media type and the bytes, then capture width, height, orientation, color model, and an integrity hash. The declared Content-Type is useful for routing but is not proof of the payload. A decoder should confirm the format, and a pixel-dimension limit should be enforced before allocating a large raster.
Here is a compact Go gate. It reads only enough to identify the format and dimensions, while the caller can stream the original to durable storage separately.
package media
import (
"bytes"
"crypto/sha256"
"encoding/hex"
"fmt"
"image"
_ "image/gif"
_ "image/jpeg"
_ "image/png"
"io"
)
type Facts struct {
Format string
Width, Height int
SHA256 string
}
func Inspect(r io.Reader, maxBytes, maxPixels int64) (Facts, error) {
limited := io.LimitReader(r, maxBytes+1)
b, err := io.ReadAll(limited)
if err != nil {
return Facts{}, err
}
if int64(len(b)) > maxBytes {
return Facts{}, fmt.Errorf("payload exceeds byte limit")
}
cfg, format, err := image.DecodeConfig(bytes.NewReader(b))
if err != nil {
return Facts{}, fmt.Errorf("unsupported or corrupt image: %w", err)
}
if int64(cfg.Width)*int64(cfg.Height) > maxPixels {
return Facts{}, fmt.Errorf("pixel count exceeds limit")
}
sum := sha256.Sum256(b)
return Facts{Format: format, Width: cfg.Width, Height: cfg.Height,
SHA256: hex.EncodeToString(sum[:])}, nil
}
The decision record should include the policy version and the reason for rejection. That gives operators a useful metric: metadata_gate_rejected_total{reason="pixels"} tells us whether the limit is wrong or the intake contract is drifting. It also prevents a later cleanup worker from treating a missing field as permission to improvise.
Choosing upload-time or on-demand processing under an SLO
Upload-time normalization is appropriate when every consumer needs the same orientation and color conversion, and when indexing cannot proceed without a canonical raster. It consumes CPU at the busiest moment, though, so capacity must be based on bursts rather than daily averages. I reserve headroom for a two-minute arrival spike and keep the synchronous path to metadata plus a bounded normalization step.
On-demand cleanup is a better fit when menus are rarely searched, when originals must remain untouched for audit, or when several derivative sizes are likely. The first search or view pays the processing latency, and a cache key must include the transform policy; otherwise a policy change serves stale pixels under a familiar URL.
The practical compromise is a two-lane pipeline: gate and persist synchronously, enqueue canonicalization asynchronously, and let the index consume only a derivative whose status is ready. A small placeholder record can carry the dish text while image readiness is pending. That keeps the write path short without hiding that the thumbnail is not yet available.
| Decision | Upload-time normalization | On-demand derivative |
|---|---|---|
| Latency seen by uploader | Higher, bounded by worker budget | Low after metadata gate |
| Burst capacity | Needs CPU headroom at intake | Queue absorbs bursts |
| First search/view | Predictable | May include decode and resize |
| Audit posture | Canonical copy is ready early | Original remains primary |
| Not suitable when | Upload clients have strict timeout budgets | Search SLO requires an image on first response |
How do you prevent orientation, OCR, and cache failures?
The most expensive bugs are valid files with invalid assumptions. EXIF orientation is a classic example: the pixel matrix can be 3024 by 4032 while the displayed dish is landscape. Normalize orientation before OCR and before generating dimensions used by a cropper. Strip metadata that is not needed by search, but retain the inspected facts and hash in your record so the operation remains auditable.
Color management deserves the same treatment. Convert to a documented working color space, flatten alpha against an explicit background, and test a plate with pale text on a transparent layer. “Looks right in my browser” is not an acceptance test. Use fixture images with known orientation, profile, and alpha combinations, then compare OCR fields and thumbnail dimensions in CI.
For retries, make the transform idempotent: (source_hash, policy_version, width) should identify one derivative. A worker can safely retry after a timeout, and a duplicate queue message will not create a second object. I log duration, bytes read, output bytes, and queue age; I do not log the raw menu image in application logs.
Three words: measure the queue.
An SLO should name both freshness and correctness, for example, “99% of accepted uploads have a ready 320-pixel thumbnail within five minutes, and 99.9% of indexed records have orientation verified.” The exact targets belong to your traffic and staffing model; I'm not sure a single threshold transfers between a ten-store chain and a national delivery marketplace, so I would validate them against an arrival histogram and a replay of representative fixtures.
Where this design is a poor fit
This split is not suitable when a legal or clinical workflow requires a fully rendered image before the upload request is acknowledged; use a synchronous, capacity-reserved path there. It is also a poor fit for tiny, fixed-format menus where an existing batch job can process everything cheaply and the operational complexity of a queue buys nothing. Stick with a simpler batch conversion when the search index is rebuilt nightly and freshness is not part of the contract.
The catch is ownership: deferred work needs a dead-letter policy, retention rules for originals, and an on-call alert tied to queue age rather than worker CPU alone. If the team cannot staff those controls, upload-time processing with a conservative concurrency limit may be the more honest choice, even if it raises median upload latency.
Top comments (0)