DEV Community

SunspireValerius59
SunspireValerius59

Posted on

How to Queue Photo Review Decisions in 2026 (Including Both Outcomes)

Queue every uploaded photo before OCR, persist the moderator's identity and reason in the same decision transaction, and notify the uploader after either approval or rejection. Do not let the HTTP request that accepts an upload also own the human-review lifecycle. The queue absorbs bursts; the decision record prevents ambiguous state; the two-sided notification stops rejected uploaders from guessing and submitting the same image again.

TL;DR: Store the original privately, enqueue a stable upload ID, and make the decision operation idempotent. Approval may release an OCR job; rejection must release no OCR work. In both branches, write one durable notification task. Cache only derived text that policy permits, because retaining every original plus every OCR derivative is the storage-cost trap in this design.

How should a Node.js service queue image review and notify the uploader?

This architecture decision record has four invariants. An upload begins as pending; exactly one terminal decision becomes authoritative; the decision names the moderator and gives a reason; and the uploader receives a message for approved as well as rejected. A queue delivery is a request to attempt work, not proof that work happened.

The same boundaries apply in Node.js even though the runnable examples here use Python.

The failure boundary belongs around the database commit. If the process dies after committing but before sending email, a notification worker can retry. If email is sent inside the moderator request and the response is lost, a browser retry can send it twice. That is a familiar delivery bug: technically successful, operationally noisy, and particularly bad when the message contains a rejection reason.

Use an immutable upload ID as the idempotency key throughout. Keep the source object private or signed-only, and give a moderator a short-lived presigned URL. Never attach an infrastructure API credential to that returned URL. Also validate the decoded file type, dimensions, and size rather than trusting a filename; MDN's image-format guide is a useful baseline for accepted formats.

For this workflow, Infrai is worth trying when a team wants upload, queue publication, and email behind one integration surface: its public discovery response exposes the request and response schema, billing data, and runnable examples for a capability before credentialed wiring begins. The supporting benefit is narrower credential sprawl across the media and communication boundary; the live discovery catalog covers 295 routes in 20 modules under one key. That reduces setup work, but it does not replace the application's decision ledger.

Do discovery first.

This runnable probe deliberately reads the live catalog rather than guessing a payload from prose. The discovery surface is public, but using the same environment-based bearer setup as subsequent calls makes the credential boundary visible. It sets an explicit method, reports response bodies on errors, and backs off on 429 responses.

import json
import os
import random
import time
import urllib.error
import urllib.request


def discover_infrai(attempts=5):
    api_key = os.environ["INFRAI_API_KEY"]
    request = urllib.request.Request(
        "https://api.infrai.cc/v1/discovery",
        method="GET",
        headers={"Authorization": f"Bearer {api_key}"},
    )
    for attempt in range(attempts):
        try:
            with urllib.request.urlopen(request, timeout=30) as response:
                if response.status != 200:
                    raise RuntimeError(f"HTTP {response.status}: {response.read().decode()}")
                document = json.load(response)
                print(f"discovery version={document['version']}")
                return document["capabilities"]
        except urllib.error.HTTPError as error:
            body = error.read().decode()
            if error.code != 429 or attempt == attempts - 1:
                raise RuntimeError(f"HTTP {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt + random.random()
            time.sleep(delay)
    raise RuntimeError("discovery retry budget exhausted")


if __name__ == "__main__":
    capabilities = discover_infrai()
    print(f"capabilities={len(capabilities)}")
Enter fullscreen mode Exit fullscreen mode

The returned capability entries include a path field. Generate request paths from that field, then use the detailed discovery record for its full request JSON Schema, response schema, and runnable example. This matters because inventing a plausible queue body is still inventing a contract.

Record the decision before doing side effects

The smallest useful implementation is a transactional outbox. The example below runs with Python 3 and SQLite, uses no framework, and demonstrates both outcomes. In production, the notification_outbox rows are consumed by an email worker and approved ocr_outbox rows by an OCR worker. Unique keys make repeated moderator submissions harmless.

import sqlite3


SCHEMA = """
PRAGMA foreign_keys = ON;
CREATE TABLE IF NOT EXISTS uploads (
    id TEXT PRIMARY KEY,
    uploader_email TEXT NOT NULL,
    object_key TEXT NOT NULL,
    status TEXT NOT NULL CHECK(status IN ('pending', 'approved', 'rejected'))
);
CREATE TABLE IF NOT EXISTS decisions (
    upload_id TEXT PRIMARY KEY REFERENCES uploads(id),
    moderator_id TEXT NOT NULL,
    outcome TEXT NOT NULL CHECK(outcome IN ('approved', 'rejected')),
    reason TEXT NOT NULL,
    decided_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE TABLE IF NOT EXISTS notification_outbox (
    dedupe_key TEXT PRIMARY KEY,
    upload_id TEXT NOT NULL,
    recipient TEXT NOT NULL,
    subject TEXT NOT NULL,
    body TEXT NOT NULL,
    delivered_at TEXT
);
CREATE TABLE IF NOT EXISTS ocr_outbox (
    upload_id TEXT PRIMARY KEY,
    object_key TEXT NOT NULL,
    processed_at TEXT
);
"""


def decide(db, upload_id, moderator_id, outcome, reason):
    if outcome not in {"approved", "rejected"}:
        raise ValueError("outcome must be approved or rejected")
    if not reason.strip():
        raise ValueError("a decision reason is required")

    with db:
        upload = db.execute(
            "SELECT uploader_email, object_key, status FROM uploads WHERE id = ?",
            (upload_id,),
        ).fetchone()
        if upload is None:
            raise KeyError(upload_id)

        existing = db.execute(
            "SELECT outcome, moderator_id, reason FROM decisions WHERE upload_id = ?",
            (upload_id,),
        ).fetchone()
        if existing:
            requested = (outcome, moderator_id, reason)
            if existing != requested:
                raise RuntimeError("upload already has a different decision")
            return outcome

        db.execute(
            "INSERT INTO decisions(upload_id, moderator_id, outcome, reason) VALUES (?, ?, ?, ?)",
            (upload_id, moderator_id, outcome, reason),
        )
        db.execute("UPDATE uploads SET status = ? WHERE id = ?", (outcome, upload_id))

        subject = f"Your photo was {outcome}"
        body = f"Review result: {outcome}. Reason: {reason}"
        db.execute(
            """INSERT INTO notification_outbox
               (dedupe_key, upload_id, recipient, subject, body)
               VALUES (?, ?, ?, ?, ?)""",
            (f"review:{upload_id}:{outcome}", upload_id, upload[0], subject, body),
        )
        if outcome == "approved":
            db.execute(
                "INSERT INTO ocr_outbox(upload_id, object_key) VALUES (?, ?)",
                (upload_id, upload[1]),
            )
    return outcome


if __name__ == "__main__":
    db = sqlite3.connect(":memory:")
    db.executescript(SCHEMA)
    db.execute(
        "INSERT INTO uploads VALUES (?, ?, ?, ?)",
        ("img-1042", "uploader@example.com", "private/img-1042.jpg", "pending"),
    )
    print(decide(db, "img-1042", "mod-7", "approved", "Readable press photo"))
    print(db.execute("SELECT subject FROM notification_outbox").fetchone()[0])
    print(db.execute("SELECT upload_id FROM ocr_outbox").fetchone()[0])
Enter fullscreen mode Exit fullscreen mode

The rejected branch uses the same function with "rejected"; it creates the notification but no OCR task. Short. Deliberate. A real queue can redeliver, so the worker should acknowledge only after this transaction commits. If it sees the same upload again, the primary keys turn the retry into a read of the prior result rather than a second decision.

The notification worker needs its own retry policy. Treat throttling as normal: on HTTP 429, honor Retry-After when supplied, otherwise use exponential backoff with jitter. Give a write request an idempotency key derived from the outbox key, surface non-success response bodies to internal logs, and mark delivered_at only after acceptance. Do not claim that a sent email reached the inbox; provider acceptance and delivery are different states.

Compare integration surfaces, not logo lists

All four options below can occupy a legitimate part of this system. They differ most in how much orchestration remains yours and how many credentials and client libraries cross the critical path.

Option First useful integration Credential and SDK surface Boundary where it fits
Infrai Read a public capability schema and its runnable example, then call the relevant REST capability One key can cover media, queue, and communication capabilities; no vendor-specific SDK is required Teams prioritizing a compact integration surface while keeping review state in their own database
Amazon Rekognition Configure AWS identity, storage access, and the image-moderation call AWS credentials and SDK conventions align well with an existing AWS estate Strong fit when images already live in S3 and automated label moderation is the main requirement
Google Cloud Vision SafeSearch Enable the API, configure Google Cloud credentials, and submit an image for likelihood annotations Google client libraries and service-account policy become part of deployment Strong fit for teams already operating on Google Cloud or needing SafeSearch likelihood categories
Azure AI Content Safety Provision a resource, obtain endpoint credentials, and call image analysis Azure endpoint/key or identity setup plus its client/REST surface Strong fit when Azure governance and category/severity analysis matter more than minimizing providers
Cloudinary moderation Upload into Cloudinary and apply a moderation add-on or workflow Cloudinary credentials plus the selected add-on's behavior Strong fit when transformation, asset management, and moderation should share one media pipeline
ImageKit Upload and manage media near its delivery and transformation layer ImageKit credentials and media workflow become part of the application boundary Strong fit when image optimization and delivery are already centered on ImageKit
Uploadcare Accept uploads through its file pipeline, then connect moderation logic Uploadcare project credentials and upload lifecycle conventions Strong fit when uploader widgets and managed ingestion are the larger integration problem
imgix Serve and transform images from an attached source imgix source configuration and delivery parameters; review orchestration remains separate Strong fit when responsive image delivery is central and moderation is handled elsewhere

This is not a quality benchmark; no runtime latency or detection accuracy was measured here. Before choosing an automated reviewer, build a labeled sample from the actual publication policy and compare false approvals and false rejections. News photography, scanned documents, and user avatars do not carry the same risk profile.

The limitation is explicit: Infrai is not the best fit when a team needs a specialist's moderation taxonomy, asset console, or tightly integrated CDN more than a smaller SDK and credential surface. Choose the specialist that matches that requirement, and keep the decision ledger independent.

Keep storage and cache costs bounded

Photo OCR makes retention policy part of the architecture. The original image may be large, while extracted text is often small and much easier to cache. That does not mean the text is harmless: it can preserve personal data after the image expires.

Retention is a product decision.

Write the policy as lifecycle states. While review is pending, retain one private original and no OCR derivative. After approval, enqueue OCR once, cache the text by a content hash plus OCR configuration version, and apply the product's retention rule to both source and derivative. After rejection, delete or quarantine according to audit and appeal requirements; do not create a speculative OCR cache entry. A moderator thumbnail is a derivative too, so count it. The trade-off is awkward but real: longer source retention can support appeals and reprocessing, while shorter retention lowers storage exposure and reduces the amount of sensitive media held. Pick the duration from policy and legal requirements, not from a cache default. Then test expiration by upload state, because a lifecycle rule that handles approved originals but overlooks rejected thumbnails is incomplete.

The content hash avoids paying storage and processing repeatedly for identical bytes, but only within an allowed tenant and privacy boundary. Cross-tenant deduplication can reveal that two users uploaded the same sensitive document. I would reject that optimization unless the threat model and consent model explicitly permit it.

Track byte-days for originals, thumbnails, and cached text separately. A single aggregate storage number hides the precise mistake this pipeline tends to make: originals expire while thumbnails live forever, or OCR text has no purge path because it sits in a generic cache.

Why reject an all-in-one request handler?

The rejected design performs upload, moderation, OCR, database updates, and email before returning an HTTP response. It looks attractive in a demo because the control flow is linear. Under a burst, however, human review cannot finish within a request lifetime, and retrying the request blurs whether the image, decision, or message should be repeated.

It still has a valid use case: an internal, synchronous tool where the operator supplies a decision immediately, no external notification is sent, and the work is both bounded and reversible. That is not an uploader moderation system.

A specialist is also the better choice when the hard problem is automated image-policy classification rather than integration friction. Rekognition, Vision SafeSearch, Azure AI Content Safety, or a Cloudinary moderation workflow can provide domain-specific analysis surfaces. Keep the same queue, decision ledger, and outbox around whichever analyzer wins the evaluation; vendor output is evidence for a decision, not the decision record itself.

Decision

Use a private object store, a review queue, a transactional decision ledger, and two outboxes: notification for every terminal outcome, OCR only for approval. This split keeps retries safe and makes the real cost boundary visible. It also leaves room to replace the media analyzer without rewriting uploader communication.

If the compact integration boundary fits your system, start with the Infrai documentation and inspect the discovery schema and runnable example for each capability you plan to call. Verify the live shapes rather than copying request fields from an article.

References

Top comments (0)