DEV Community

SullivanReed1247
SullivanReed1247

Posted on

Python Automated Caption Moderation Coverage Against Manual Catalog Photo Audits

An ecommerce upload pipeline has one awkward constraint: the text that matters may exist only inside the photo, while the decision to publish is expected immediately. Processing every image at upload maximizes early coverage, but it also puts OCR latency and ambiguous text in the checkout or listing path. Processing only on demand keeps ingestion light, yet leaves unrequested images unexamined.

TL;DR: Run cheap structural checks at upload, queue OCR for every publishable catalog image, and reserve human inspection for policy-sensitive or uncertain cases. On-demand OCR still belongs in reprocessing and investigation workflows, not as the primary moderation trigger. Machine captions broaden screening coverage; people supply judgment. Neither is a substitute for the other.

How should automated caption moderation and human image review share coverage?

A raw percentage is a poor definition. A pipeline can report that it processed 100% of uploads while missing text in a secondary gallery image, accepting a stale OCR result after replacement, or treating a generated caption as a policy decision. The useful unit is a specific image version that reached a specific policy check before publication.

I would track four separate states: received, decodable, text-extracted, and policy-resolved. That split matters because each state has a different owner and retry rule. A corrupt file is an ingestion outcome. Empty OCR can mean no visible text or a failed extraction. A policy result can be automatic, manually resolved, or still pending. Collapsing those outcomes into moderated=true creates an attractive dashboard and a weak audit trail.

The catalog boundary matters too. A seller may upload a hero photo plus seven gallery photos, then replace only the third. Coverage should follow the immutable asset version, not the listing ID. Otherwise an approval attached to the old bytes can appear to cover the replacement.

This is where a deliverability mindset helps: accepted is not delivered, and processed is not resolved. Name the stages.

Counts can lie.

Derive the pipeline from the publication deadline

Start with the business rule: can a listing become visible before image text has been resolved? If the answer is no for a regulated category, OCR and policy resolution sit on the publication path even if the work happens asynchronously. The API can accept the upload, but publication remains pending. For a lower-risk category, the listing might publish after structural validation while later findings can restrict it. That is a policy choice, not an OCR setting.

At upload, validate the declared media type against bytes, enforce size and dimension limits, assign an immutable asset version, and store the original. Image formats have different browser support and characteristics, so normalization should be an explicit operation rather than an assumption based on a filename. The MDN image-format guide is a useful reference for the format boundary.

Then enqueue extraction with an idempotency key built from the asset version and extractor revision. Keep the OCR text, regions, language hints, confidence evidence, and revision together. A later policy evaluator should consume that record and produce its own versioned decision. Separating extraction from evaluation lets a policy change reclassify known text without decoding every image again.

Do not let an OCR timeout silently become approval.

A compact Python model makes those boundaries visible:

from dataclasses import dataclass
from enum import Enum


class Resolution(str, Enum):
    ALLOW = "allow"
    REVIEW = "review"
    BLOCK = "block"


@dataclass(frozen=True)
class Extraction:
    asset_version: str
    extractor_revision: str
    text: str
    regions_found: int
    complete: bool


def route(extraction: Extraction, policy_hits: list[str]) -> Resolution:
    if not extraction.complete:
        return Resolution.REVIEW
    if policy_hits:
        return Resolution.REVIEW
    return Resolution.ALLOW
Enter fullscreen mode Exit fullscreen mode

The example deliberately refuses to turn a policy hit into an automatic block. Some rules can support deterministic blocking, but that requires a separately reviewed rule and evidence model. Human review is the conservative default for ambiguous language, context-dependent claims, and incomplete extraction.

Machine captions and people cover different gaps

A generated caption is useful for triage because it turns visual material into searchable evidence. It can flag likely text-bearing images, describe scene context, and help prioritize a queue. OCR should still retain detected text and regions separately; a fluent caption can omit a small disclaimer or paraphrase a prohibited claim. Store evidence, not merely prose.

Human inspection covers context that a text-first path cannot settle: whether words belong to the product or a background sign, whether a before-and-after collage makes an implied claim, and whether tiny text changes the meaning of a prominent headline. People also make mistakes, especially in repetitive queues. Give reviewers the original image, highlighted regions, extracted text, applicable rule version, and a small set of reason codes. Random ordering, duplicate suppression, and escalation paths are operational controls, not interface polish.

Coverage layer Best role Known gap Publication treatment
Structural validation Reject unreadable or disallowed input Cannot interpret content Block ingestion or request replacement
OCR Locate and preserve visible text evidence Weak context and uncertain extraction Route incomplete or relevant cases
Machine caption Add scene-level triage signals May summarize away decisive detail Never treat the caption alone as final evidence
Human inspection Resolve contextual policy questions Queue delay and inconsistent judgment Use reason codes and escalation
Audit sampling Measure misses after decisions Finds problems after the first decision Feed policy and test-set changes

The comparison is therefore asymmetric. Automation can touch every queued asset at stable machine speed, while human capacity should concentrate on the cases where judgment changes the outcome. Full machine contact is not full policy coverage. Full manual inspection is not automatically complete either if reviewers see resized previews that hide small text or if replacement images bypass the queue.

There is a real trade-off. Upload-time OCR is not suitable as a blocking dependency when the business permits immediate publication and the review queue cannot meet that deadline; in that case, publish after structural validation, process asynchronously, and define the exact later action for a policy finding. Conversely, on-demand extraction is a poor primary control for categories that require resolution before publication because untouched images remain outside the moderation path. Human review has its own limitation: inspecting every ordinary image consumes finite attention that is better reserved for uncertain evidence, while sampling must still test automated allows for misses. The right split follows the publication rule and the consequence of a miss, not a blanket preference for machines or people.

No layer gets a free pass.

Make uncertainty an explicit queue input

One score is tempting and usually underspecified. Route on named conditions instead: extraction incomplete, relevant term detected, text region too small for the retained rendition, language unsupported by the current review pool, asset replaced after decision, or policy rules changed after extraction. Each condition needs an observable counter and a terminal state.

Avoid inventing a universal confidence threshold. Calibrate routing against a labeled sample from the actual catalog categories and image transformations in use. Keep false-negative review separate from reviewer agreement: one measures what automation missed, while the other measures whether the policy can be applied consistently. If disagreement clusters around one rule, adding more OCR capacity will not repair the rule.

The edge cases are mundane and expensive: animated inputs whose first frame differs from later frames, orientation metadata that changes the rendered view, transparent text over a checkerboard, repeated uploads of the same bytes, and a crop generated after the original was approved. The ingestion contract should choose which rendered artifact is moderated and retain enough lineage to reproduce it.

Observability should follow that contract. Count assets by version through each stage; measure queue age by risk class; record retries separately from unique assets; and alert on unresolved publication holds. For sampling, draw from allowed, reviewed, and blocked outcomes. Sampling only the review queue cannot estimate the misses that never entered it.

A compact rollout from on-demand to upload-time OCR

Begin in shadow mode: extract from new uploads without changing publication, then label a stratified sample of automated allows and review candidates. Use those results to define routing conditions and reviewer guidance. Next, hold publication only for the narrow categories where unresolved image text creates the clearest policy exposure. Keep a kill switch that changes the publication rule, but never converts extraction failure into approval without an explicit policy decision.

Backfill by immutable asset version, with lower priority than new uploads. Re-evaluate stored extraction when policy rules change; re-extract only when the source image, rendering procedure, or extractor revision requires it. This distinction controls load and preserves an explainable history.

The durable design is straightforward: upload-time processing establishes broad, versioned evidence coverage; manual inspection resolves selected uncertainty; and on-demand execution supports appeals, investigations, and controlled reprocessing. The gate is the policy resolution, not the existence of a plausible caption.

References

Top comments (0)