Storage fan-out changes the moderation answer: a customer-support photo may be retained as an original, normalized into another representation, sent through OCR, and cached again before anyone decides whether it is acceptable. Short answer: text checks can evaluate extracted words, filenames, and captions; they cannot establish that the pixels are safe, that OCR preserved the meaning, or that an unsupported visual condition was harmless. Treat text and image analysis as separate signals, then make one policy decision with an explicit review outcome.
For capacity planning, I would model this as an intake pipeline rather than a single moderation call. Consider a deliberately bounded load model, not a benchmark: 10 million photos per month, a 6 MB admission limit, one normalized derivative, and one retry budget. Those inputs force the team to account for retained bytes, temporary bytes, cache keys, and duplicate work. They do not predict a real workload; measured arrival rate and object sizes must replace them before procurement.
The invariant is blunt. No successful text verdict can clear the image channel.
What can upload moderation cover across text and image?
Picture an incident exercise in which an OCR worker is slow while object ingestion remains healthy. Captions continue to pass text rules, queued photos accumulate, and clients retry requests whose status they cannot see. If the system equates "caption accepted" with "photo accepted," it has silently converted an unavailable image signal into approval. If it stores every retry under a new key, the same control failure also becomes a storage event.
The useful lesson is not that OCR is unreliable in some vague sense. OCR answers a narrower question: what text can be extracted from this raster under the decoder and recognition conditions presented? A phrase denylist, a language-aware classifier, or a case-management rule can then evaluate that text. None of those operations checks an uncaptioned visual symbol, the relationship between words and imagery, or material that the decoder never exposed to OCR.
Unknown blocks promotion.
Format handling makes the boundary concrete. Browsers and services encounter JPEG, PNG, GIF, WebP, AVIF, SVG, and other image types with different capabilities; MDN's image format guide documents those differences and the need to consider support. Admission should therefore identify a permitted type, decode it under resource limits, and normalize the representation used by downstream analysis. A filename suffix alone is not that decision.
This is where an SLO needs an honest denominator. Report the percentage of admitted photos that reached each required signal, the percentage sent to review, and the age of the oldest unresolved item. A single "moderation success rate" hides partial coverage. It also encourages a dangerous fallback: counting a text-only result as complete when image analysis timed out.
Separate evidence before combining policy
I use three states for each required signal: pass, fail, and unknown. Unknown covers timeout, decode rejection, exhausted retry budget, or any other condition in which the signal did not produce a policy verdict. It is not a softer pass.
The policy combiner can then stay boring. A known failure rejects or quarantines the item; all required passes allow it; any unknown routes it to a bounded review queue or keeps it pending. The exact action depends on the harm model, but the state transition should never depend on a worker returning an empty string.
| Evidence channel | What it can support | What remains outside that signal | Cost driver |
|---|---|---|---|
| User text | Rules over captions, filenames, and form fields | Content visible only in pixels | Text request volume and retained audit data |
| OCR text | Rules over text successfully extracted from a decoded image | Missed, distorted, occluded, or purely visual content | Decode work, OCR work, retries, and cached results |
| Image analysis | Rules over the normalized visual representation it actually processed | Business context and ambiguous intent | Derivatives, model work, and review volume |
| Human review | Contextual resolution within the reviewer's policy and view | Items never routed or evidence not shown | Queue capacity, tooling, and retention |
Keep the evidence records distinct even if the final policy is compact. That preserves the reason for a decision, permits a targeted re-evaluation when one detector changes, and avoids rerunning OCR merely because a caption rule was updated.
A preventative intake path
The following Go sketch shows the control point, not a vendor SDK. The content hash gives retries a stable identity, while each analyzer must return an explicit status. Production code still needs authenticated callers, streaming size enforcement, decoder resource limits, durable state transitions, and access controls around sensitive support material.
package moderation
import (
"context"
"crypto/sha256"
"errors"
)
type Status uint8
const (
Unknown Status = iota
Pass
Fail
)
type Evidence struct {
Text Status
Image Status
}
type Analyzer interface {
Check(ctx context.Context, normalized []byte) (Status, error)
}
type Store interface {
PutIfAbsent(ctx context.Context, key [32]byte, normalized []byte) error
}
func Admit(ctx context.Context, normalized []byte, text, image Analyzer, store Store) (Status, error) {
if len(normalized) == 0 {
return Unknown, errors.New("empty normalized image")
}
key := sha256.Sum256(normalized)
if err := store.PutIfAbsent(ctx, key, normalized); err != nil {
return Unknown, err
}
textStatus, textErr := text.Check(ctx, normalized)
imageStatus, imageErr := image.Check(ctx, normalized)
if textErr != nil || imageErr != nil {
return Unknown, nil
}
evidence := Evidence{Text: textStatus, Image: imageStatus}
if evidence.Text == Fail || evidence.Image == Fail {
return Fail, nil
}
if evidence.Text == Pass && evidence.Image == Pass {
return Pass, nil
}
return Unknown, nil
}
One detail deserves skepticism: hashing the normalized bytes deduplicates that representation, not every semantically equivalent photo. Re-encoding, rotation, cropping, or metadata changes can produce different bytes. Perceptual matching is a separate mechanism with false-match consequences, so it should not be smuggled into an idempotency claim.
Cache keys also need the analyzer version and policy version alongside the content identity. Otherwise an old pass may survive a policy change. Conversely, blindly including a request identifier defeats reuse and causes retry amplification. The useful cache hit is one whose input representation, analyzer configuration, and policy interpretation are all the same.
Buy, build, or combine?
The operational choice is larger than model quality. It changes who owns decoder patching, queue behavior, data retention, capacity, and the evidence needed during an appeal. Each option has real limitations: managed analysis gives up some operational control, self-hosting adds on-call work, and a combined path increases state-space and test cost.
| Approach | Platform team owns | Main advantage | Main exposure |
|---|---|---|---|
| Managed analysis | Admission, policy combination, review, and vendor failure handling | Less detector infrastructure on call | Data boundary, external dependency, and switching cost |
| Self-hosted analysis | The full runtime, models, scaling, patching, and evaluation | Direct control over deployment and retention | Larger on-call and capacity-planning burden |
| Combined path | Routing plus both operational contracts | Different treatment for different risk classes | More states, more tests, and reconciliation work |
I would not choose from that table until a replay set represents the support queue: screenshots, camera photos, rotated receipts, low-contrast text, multiple languages, and files at the admission boundary. Labels need a documented policy and disagreement process. Compare false approvals, false rejections, unknown rates, review minutes, peak queue age, stored bytes per accepted item, and cache hit rate. Aggregate accuracy alone cannot reveal whether the expensive errors sit in the highest-risk category.
Measure the queue.
Set the SLO around the user-visible decision. For example, define separate objectives for automated decisions and reviewed decisions, then budget capacity for the arrival distribution rather than its monthly average. The storage forecast should include originals retained by policy, normalized objects, evidence, and temporary retry overlap. Delete temporary material on a defined lifecycle; do not rely on a successful happy path to clean it up.
Where this advice does not apply
Some uploads do not need OCR. A workflow that accepts only machine-generated text documents should parse the document structure when the format and trust boundary permit it. At the other extreme, a channel whose risk policy requires every image to be reviewed should not pretend that additional OCR converts review into automation.
The three-state combiner may also be too coarse for policies with category-specific thresholds or legal escalation. Preserve the same principle: missing required evidence remains visible, and each final action records which policy version consumed which signals.
Cost is a constraint, not the verdict. The defensible design stores each expensive artifact once when its identity is stable, reuses it only across equivalent analyzer inputs, and measures the queue created by uncertainty. It never asks text moderation to vouch for pixels it did not examine.
Sources
References:
- MDN Web Docs, "Image file type and format guide": https://developer.mozilla.org/en-US/docs/Web/Media/Formats/Image_types
Top comments (0)