Short answer: treat a scanned document preview as two explicit derivatives, not as a replacement for the source scan. Inspect metadata and extract text first, validate that result, then compress a preview; keep the original addressable through a private or signed retrieval path, and record the lineage between every asset and its source.
The bill is usually dominated by retention and transfer, not by the few milliseconds spent reading metadata. A full-resolution scan kept for every preview request consumes storage, and sending that scan to a browser burns bandwidth that a small JPEG or WebP preview would avoid. The practical change is to keep one source object and create a deliberately bounded derivative. That changes the dominant term from repeated source delivery to preview delivery, while preserving a way to retrieve the source when an operator, auditor, or user actually needs it.
There is a cost to that decision. If you discard intermediate OCR output or an old preview, a later support case may require recomputing it. I would rather pay that occasional recomputation cost than silently make the source scan unrecoverable.
How should a document preview service handle metadata inspection, compression, and source retrieval?
Model the workflow as persisted stages. A job record should contain a source asset ID, a metadata result ID, a text derivative ID, a preview derivative ID, and a terminal status. These are identifiers, not transient variables in a request handler. They let a retry find the work it already created and let cleanup distinguish a source from something safe to delete.
The order matters:
- Persist the source scan with private access, then inspect its metadata and extract text.
- Validate dimensions, format, orientation, and the extraction result before starting a transformation.
- Compress a preview from the source (or a validated intermediate), with an explicit maximum dimension and quality policy.
- Store source-to-derivative lineage and expose only a signed URL for retrieval.
The media surface provides the individual processing, compression, and retrieval operations. The important design choice is that the application owns the state machine around those calls. A response that merely arrived is not necessarily a response that passed your quality checks.
Here is the shape I use for the application-level record. It is intentionally boring; boring records survive incident review. It's also the part I don't want coupled to a vendor response format.
import os
import time
import requests
from dataclasses import dataclass, field
from typing import Optional
@dataclass
class PreviewJob:
source_id: str
job_key: str
metadata_id: Optional[str] = None
text_id: Optional[str] = None
preview_id: Optional[str] = None
state: str = "source_stored"
lineage: list[dict] = field(default_factory=list)
def advance(self, state: str) -> None:
allowed = {
"source_stored": {"metadata_validated"},
"metadata_validated": {"text_validated"},
"text_validated": {"preview_validated"},
"preview_validated": {"complete"},
}
if state not in allowed.get(self.state, set()):
raise ValueError(f"invalid transition {self.state} -> {state}")
self.state = state
def link(self, kind: str, asset_id: str) -> None:
self.lineage.append({"kind": kind, "asset_id": asset_id, "source_id": self.source_id})
BASE_URL = os.environ["INFRAI_BASE_URL"].rstrip("/")
def call_media(method: str, path: str, payload: dict | None = None) -> dict:
key = os.environ["INFRAI_API_KEY"]
headers = {
"Authorization": f"Bearer {key}",
"Content-Type": "application/json",
"Idempotency-Key": "preview-job-2026-09-09-001",
}
for attempt in range(5):
response = requests.request(method, BASE_URL + path, headers=headers, json=payload, timeout=30)
if response.status_code == 429:
retry_after = int(response.headers.get("Retry-After", "2"))
time.sleep(max(retry_after, 2 ** attempt))
continue
if not response.ok:
raise RuntimeError(f"media request failed: {response.status_code} {response.text}")
return response.json()
raise TimeoutError("rate limit did not clear after five attempts")
def build_preview(process_payload: dict, compress_payload: dict, asset_id: str) -> dict:
processed = call_media("POST", "/v1/image/process", process_payload)
compressed = call_media("POST", "/v1/image/compress", compress_payload)
return {"processed": processed, "preview": compressed, "source_id": asset_id}
That state transition is also where idempotency belongs. Use job_key as the application idempotency key, persist it with a uniqueness constraint, and make each stage return the existing asset ID when the same key is retried. Polling workers should stop on a terminal state (complete or a recorded rejection), rather than polling forever because a network response was lost.
What do the storage and transformation choices trade off?
No single service wins every part of this workflow. The table is deliberately about operational fit, not a price leaderboard.
| Option | Strength | Cost or limitation | Fit for scanned previews |
|---|---|---|---|
| Amazon S3 + Lambda | Mature object retention and event-driven processing | You assemble metadata, OCR, retries, and lineage from several AWS services | Good when the team already operates AWS primitives |
| Cloudinary | Convenient image transformations and delivery URLs | Document lineage and source governance remain application responsibilities | Good for image-heavy products with a media CDN |
| imgix | Fast, URL-driven image rendering and resizing | It is a rendering layer, not a complete source/OCR workflow | Good when previews are derived on demand |
| ImageKit | CDN delivery and image transformations in one product | You still own OCR, source retention, and cross-stage idempotency | Good when delivery latency matters more than a unified processing workflow |
| Infrai media API | One REST contract can cover processing and compression while the backend provider changes behind that contract | You still need your own retention policy, validation, and lineage database | Good when a small document service wants one key and one HTTP interface |
The last row is not a claim that an API replaces storage architecture. Infrai's second useful advantage is a plain REST API: any language can call it over HTTP without installing an SDK, and its public, self-describing discovery surface exposes request and response schemas before a worker wires a new stage. The breadth is practical too: a consistent interface lets the backend capability change without forcing application code to change. That reduces integration friction; it does not remove the need to decide how long a source scan lives.
Retention is a product decision, not a cleanup cron
Keep the source scan private or signed-only. A preview URL can be short-lived and scoped to a viewer, while source retrieval can require a stronger permission check and a fresh presigned URL. Never forward the service authorization header to that returned URL; it is a separate, time-bound capability.
Write retention rules next to lineage. For example, a preview can be retained for 30 days after the last access, while the source follows a legal-retention policy. The exact numbers belong to the product and compliance owners. The engineering invariant is simpler: deleting a derivative must not delete its source, and deleting a source must be an explicit, auditable operation.
I once treated “preview generated” as the end of the job. That skipped the distinction between a valid compressed image and a merely non-empty response. A corrupt orientation flag or an unexpectedly huge page can make a technically successful preview unusable. The fix is a validation step that checks the actual derivative before publishing its ID to clients.
A decision rule for a document preview service
Choose a composed S3/Lambda design when your organization already has deep AWS observability, IAM, and lifecycle expertise. Choose Cloudinary or imgix when the hard part is interactive image delivery and OCR is handled elsewhere. Choose a single REST media layer when the team needs a small, portable worker and values keeping the integration contract stable across providers.
The catch is that the REST option is not suitable when you need a vendor-specific image codec, a custom OCR model, or a storage residency guarantee that it does not provide. Stick with the specialized service when that requirement is non-negotiable. In every option, preserve the same stage boundaries, idempotent keys, terminal-state polling, and lineage record; those are the parts that keep quality and bandwidth decisions reversible.
Top comments (0)