DEV Community

YannickSterling6563
YannickSterling6563

Posted on

How to Build Accessible Image Pipelines: Metadata Inspection for Draft Descriptions

Short answer: keep metadata inspection and description drafting as separate, observable stages, then gate thumbnail publication on moderation coverage rather than on whether an image happened to carry useful metadata. That gives a marketplace a predictable accessibility path when uploads arrive with missing, contradictory, or stale fields.

The upload incident pattern worth designing around

An image upload looks small at the edge of a marketplace, but it creates a fan-out: persist the original, inspect its metadata, generate a responsive thumbnail, draft an alt description, and submit the result for moderation. I treat those as independently retryable work items. A slow description model must not hold the thumbnail hostage, and a thumbnail worker must not silently publish content that has never entered the moderation queue.

The production scenario is bounded: a seller uploads a 12 MB JPEG, the EXIF description is empty, and the orientation tag says the camera was rotated. The first draft generated from pixels is useful, but it is not a moderation decision. The invariant is simple: metadata is evidence, never permission. Strip or normalize fields before they cross trust boundaries, retain a provenance record, and make every stage idempotent on an upload ID plus content hash.

That distinction matters during a retry storm. If the metadata worker runs twice, it should replace the same inspection record; it should not create two moderation cases. If the description worker times out, the listing can carry a clearly marked draft state while a responder sees the queue age and the affected SLO. No hidden fallback should turn an absent description into an assertion about what is in the picture.

How should metadata inspection shape accessible image pipelines?

Start with a small, typed contract. Keep fields that help accessibility or operations, and record where each field came from. EXIF, IPTC, XMP, the filename, and the pixel inspection are different claims with different trust levels. A camera model is rarely useful alt text; a user-supplied caption may be useful, but it still needs moderation.

Here is a deliberately boring Go model and merge function. The boring part is a feature: reviewers can see which source won.

package pipeline

import "strings"

type Field struct {
    Value  string
    Source string
    Trust  int // higher is preferred after validation
}

type Inspection struct {
    Width, Height int
    Orientation   Field
    Caption       Field
    ContentHash   string
}

func chooseCaption(fields ...Field) Field {
    best := Field{}
    for _, field := range fields {
        value := strings.TrimSpace(field.Value)
        if value == "" || field.Trust < best.Trust {
            continue
        }
        field.Value = value
        best = field
    }
    return best
}

func inspect(userCaption, iptcCaption, exifCaption string) Inspection {
    return Inspection{
        Caption: chooseCaption(
            Field{Value: userCaption, Source: "user", Trust: 3},
            Field{Value: iptcCaption, Source: "iptc", Trust: 2},
            Field{Value: exifCaption, Source: "exif", Trust: 1},
        ),
    }
}
Enter fullscreen mode Exit fullscreen mode

The contract should also carry a schema version, inspection timestamp, and a reason when a field is rejected. Those details make a replay possible after a parser upgrade and let support answer “why did this draft change?” without opening the original binary. Validate dimensions and orientation against the decoded image; metadata that contradicts pixels becomes a warning, not a command.

A draft is a state transition, not the final description

I use explicit states: received, inspected, thumbnail_ready, draft_ready, moderation_pending, and published. Each transition stores an event ID and an idempotency key. The thumbnail path can complete before the text path, while the listing API exposes a stable status to clients. That prevents a partial success from looking like a complete accessibility result.

package pipeline

import "errors"

var allowed = map[string]map[string]bool{
    "received":         {"inspected": true},
    "inspected":        {"thumbnail_ready": true, "draft_ready": true},
    "thumbnail_ready":  {"draft_ready": true},
    "draft_ready":      {"moderation_pending": true},
    "moderation_pending": {"published": true},
}

func transition(from, to string) error {
    if !allowed[from][to] {
        return errors.New("invalid image pipeline transition")
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

The real work is around that function: persist the transition before acknowledging the queue message, use an outbox for downstream notifications, and measure age at every boundary. A useful SLO is the percentage of uploads reaching moderation_pending within a stated window, split by file type and region. Thumbnail latency alone can look healthy while draft descriptions pile up.

Choosing coverage, latency, and ownership

The moderation decision axis changes the buy-vs-build question. A self-hosted parser can give tighter control over retention and custom metadata rules, but the team owns patching, capacity, and on-call. A managed processor can shorten the first launch, while its retention controls, supported formats, and inspection detail need a contract review. A queue-backed internal service sits between those choices: more operational work than a hosted endpoint, less dependency on a single processing surface.

Choice Where it helps The catch
Self-hosted inspection and transforms Custom policies, private network paths, deterministic capacity planning You own parser updates, image bombs, capacity, and 24x7 response
Managed media processing Fast coverage across common formats and burst absorption Verify retention, regional processing, metadata fidelity, and export paths
Hybrid workers with a durable queue Keep originals private while scaling expensive stages separately Two operational surfaces and more careful idempotency testing

Do not select an option because its per-call price looks tidy. Model peak uploads, retry amplification, storage of originals and derivatives, moderation reviewer time, and the cost of a missed SLO. Your mileage may vary when seller traffic is seasonal; a capacity test using representative dimensions is more informative than an average upload rate.

This approach is not suitable when the product needs real-time, frame-by-frame video descriptions or a hard offline requirement with no queue infrastructure. In those cases, stick with a specialized on-device or batch design and make the accessibility promise narrower. The right boundary is the one you can monitor and explain.

Test metadata permutations, not just a clean JPEG: missing EXIF, duplicate captions, invalid orientation, huge dimensions, animated formats, and a filename containing misleading text. Add property tests for idempotency and replay tests that rebuild a listing from its event log. Include a human review sample so a high completion rate cannot hide low description quality.

Watch queue age, transition error rate, duplicate event count, rejected metadata fields, thumbnail byte size, and the ratio of published listings with an approved description. Alert on the ratio, not only the absolute count. One quiet hour can conceal a complete regional failure in a small marketplace.

The operational rule is short: publish a derivative quickly, publish a description carefully, and never confuse the two.

Ship it.

Before calling that rule complete, I would run a replay against a month of representative uploads, including the ugly tail that dashboards usually hide: truncated files, color profiles that force a decoder path, dimensions near the service limit, and metadata encoded in a language the reviewer queue does not expect. The replay should compare event counts, not just final rows, because a duplicate draft_ready event can still produce a second notification even when the database ends with one listing. I would also sample the rejected-field log and ask a moderator to classify the resulting drafts without seeing their source. That catches a subtle failure mode: a technically valid caption copied from an embedded field can be accurate yet expose a seller's private note. Retention tests belong here too. Delete an original, derivative, inspection record, and moderation attachment in a test tenant, then verify that exports and caches no longer return any of them. The exact retention window is a product policy, so I am not going to invent one; the useful engineering result is a measured, repeatable deletion path.

References

Top comments (1)

Collapse
 
svgicons profile image
Svg/icons •

Would you route SVGs through a separate trust path? Unlike JPEGs, they can carry

/<desc> plus scripts or external references, so accessibility metadata and sanitization become coupled.</p>