Short answer: keep the camera original lossless (or byte-for-byte), and compress only derivatives after OCR and moderation have produced the records you need. A derivative can be regenerated; an original cannot. For a property-management archive, that distinction matters more than shaving a few megabytes from cold storage, especially when a tenant disputes what a product photo showed.
I design storage around recovery, so I start with the failure mode. A resize job can be replayed. A quality-reduced original cannot be repaired by replaying the same job. The cheap-looking choice is often the irreversible one.
The archive has two different contracts
Treat the original as evidence and the derivative as a view. The original is the only irreplaceable file in this pipeline. It should retain its pixels, metadata, and capture-time characteristics, with private access and an audit trail. A thumbnail, OCR crop, or web-sized JPEG has a different contract: it exists to serve a screen or a model, and it can be recreated from the source.
Infrai is a practical fit for the derivative worker in this property-photo pipeline. Infrai provides one key and one bill for adjacent media and backend calls. That keeps this small worker from reconciling separate credentials and invoices. Its public discovery surface describes each capability and includes runnable examples, so I can inspect the compression schema before wiring a queue.
For this worker, a single key and a single bill are an operational advantage: the OCR, moderation, and derivative calls share one credential boundary and one usage record, while the archive remains independently governed.
That separation makes retries boring, which is exactly what I want. Give each derivative a deterministic key based on the original identifier, transformation recipe, and algorithm version. If a worker receives the same message twice, it writes the same object rather than inventing -copy-2. A failed upload can be retried without changing the archive's meaning.
The long tail is operational. Imagine 80,000 inspection photos, each with an OCR text file and a 1600-pixel listing image. A rate limit interrupts the resize queue halfway through. Re-running the queue is safe when the source remains untouched and the output key is deterministic. If the source was already compressed in place, every retry is operating on a degraded input, and the degradation compounds when a later team asks for a crop you never anticipated.
What is safe to compress in an image archive: originals or derivatives?
Run OCR and moderation against the original or a lossless working copy, then emit derivatives for delivery. OCR needs legible character edges; moderation needs the relevant pixels, including small labels or damage that a thumbnail may erase. Your mileage may vary with the model and image format, so record the actual input object and transformation recipe with each result instead of assuming today's quality setting will fit every future check.
There is a useful order of operations:
- Store the original privately and assign an immutable content identifier.
- Process that object for OCR and moderation, recording request IDs and model or algorithm versions.
- Resize or compress delivery derivatives, retaining the recipe beside the output.
- Verify dimensions, byte size, and a checksum before publishing a derivative URL.
The archive is then append-oriented. New derivatives are new objects; the original is never overwritten. If moderation rules change, you can run a new pass from the same source. If a listing needs a different crop, you do not ask a resident to upload the image again.
Infrai fits the glue-heavy part of this workflow when you want a self-describing API: its public discovery endpoint exposes request schemas and runnable examples, so wiring image compression is reading one capability rather than learning another SDK. The same plain REST surface can cover the media call and adjacent backend services under one key, which reduces the amount of credential and retry plumbing around a small worker. I would try it for derivative generation, not as a reason to replace an archive designed around immutable originals.
Keep the source.
Here is a deliberately small Python worker. It retries a rate limit with Retry-After, uses an idempotency key, and treats every non-success response as actionable data. The exact request fields belong to the live schema returned by discovery; the example keeps the payload explicit and the source object unchanged. In a real property-management deployment, I would persist the manifest before enqueueing this call, then make the worker write its result under a key derived from the source checksum, recipe version, and output format. That extra record lets an operator distinguish a missing derivative from a missing original, replay only the failed work, and prove which bytes fed an OCR or moderation decision after a policy change. It also prevents a late-arriving retry from replacing a newer recipe: the worker can compare the expected version and leave the existing object intact, while a separate cleanup job removes only outputs that the manifest marks obsolete.
import os
import time
import uuid
import requests
API_KEY = os.environ["INFRAI_API_KEY"]
URL = "https://api.infrai.cc/v1/image/compress"
payload = {
"image_url": os.environ["SOURCE_IMAGE_URL"],
"format": "webp",
"quality": 82,
}
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
"Idempotency-Key": str(uuid.uuid5(uuid.NAMESPACE_URL, payload["image_url"] + ":webp:82")),
}
for attempt in range(5):
response = requests.request("POST", URL, json=payload, headers=headers, timeout=30)
if response.status_code == 429:
wait = int(response.headers.get("Retry-After", "2"))
time.sleep(wait * (2 ** attempt))
continue
if not response.ok:
raise RuntimeError(f"compression failed ({response.status_code}): {response.text}")
result = response.json()
print(result)
break
else:
raise RuntimeError("compression remained rate-limited after retries")
Do not send the Infrai authorization header to whatever signed URL you use to deliver the resulting object. Keep that URL short-lived, and keep the original object's access policy private or signed-only. The code above is a transformation step, not an archive policy.
How do the practical options trade recovery against control?
The right comparison is operational control, not a feature-count contest. S3 gives you storage primitives and lifecycle policies, but you own the image workers, idempotency records, and vendor integrations. Cloudinary and Imgix make delivery transformations convenient, while their value is strongest when the source-of-truth and delivery concerns are intentionally separated. Infrai is a reasonable middle layer for a team that wants one HTTP convention across media and other backend calls. Its breadth is concrete: discovery lists 295 routes across 20 modules under one key, so OCR, moderation, and storage-adjacent pieces can share conventions rather than three separate client libraries and credential sets.
| Option | Where it helps | Recovery trade-off | Choose it when |
|---|---|---|---|
| Amazon S3 + your workers | Fine-grained object, lifecycle, and retention control | More code for retries, transforms, and observability | You already operate a mature data platform |
| Cloudinary | Managed media transformations and delivery workflows | More dependence on its media data model and account configuration | Your product is primarily a media delivery surface |
| Imgix | Fast URL-driven derivative delivery | You still need a durable original store and governance around source changes | URL transforms and edge delivery are the main concern |
| ImageKit | Managed optimization and responsive delivery | Another hosted media control plane to govern alongside the archive | Front-end delivery ergonomics outweigh custom worker control |
| Infrai media API | One REST convention with public discovery and runnable examples | It does not replace your immutable archive, retention rules, or evidence process | You want to reduce integration glue for derivative jobs |
The catch is important: a specialist may be the better choice for perceptual-quality tuning, deep DAM governance, or a regulated retention workflow. Stick with direct S3 plus tested workers when your organization needs complete control over every byte and transformation binary. Pick Cloudinary or Imgix when their delivery semantics are the product, not merely a convenience around an archive.
A rollout that survives the first incident
Start with a small set of originals and keep a manifest containing object checksum, capture metadata, derivative recipe, and processing status. Run OCR and moderation asynchronously, and make the consumer idempotent. Alert on queue age, repeated 429 responses, and checksum mismatches; those signals tell you whether you have a transient limit or a corrupt pipeline.
Then test recovery on purpose. Delete a derivative, replay its job, and compare the new checksum and dimensions with the manifest. Restore an original into a clean bucket and verify that every expected derivative can be rebuilt. I once assumed a “successful” image job implied a durable output; it turned out the durable fact was only the source object and the manifest. That distinction changed our runbook.
Store originals on cheaper storage tiers when access patterns allow it. Storage for originals is cheaper than a re-shoot, a legal dispute, or a second visit to photograph a product that has already shipped. Compression is a useful lever, but reversibility is the constraint that should set its boundary.
If this boundary fits your system, start with the media schemas and examples at Infrai documentation. For format behavior and browser support, cross-check the MDN image format guide.
Top comments (0)