Short answer: validate dimensions, color intent, alpha behavior, file integrity, and policy metadata before conversion, then keep the original and the decision record long enough to investigate a rejected product photo. The converter should receive only assets that passed those gates.
In healthtech catalogs, a product photo can be both a sales asset and evidence of what a patient will receive. A transparent background that turns white, a 96 DPI export that is actually 640 pixels wide, or an embedded profile stripped during conversion can make a printed package misleading. The expensive part is rarely the conversion call. It is the rework: another upload, another proof, another moderation review, and a support ticket that arrives after the print slot closes.
I use a gate model because it gives the operations team a precise answer to "why was this file rejected?" It also keeps the conversion engine replaceable. The record says what was checked and which input hash was checked, rather than trusting a filename or a human memory.\n\nKeep it boring.\n\n## How do print-on-demand artwork metadata checks work before conversion?
Start with facts that can be measured without decoding every pixel. Record the byte size, media type detected from the signature, width, height, frame count, color model, profile name, alpha presence, and any orientation flag. For print-on-demand artwork, also record the requested physical size and the minimum effective pixels-per-inch (PPI). PPI is a relationship, not a magic field: 2400 pixels placed across 8 inches gives 300 PPI, while an EXIF value of 300 on a 640-pixel image does not create detail.
The gate should be deterministic. A 4000 x 5000 image with an sRGB profile can pass one rule set; a CMYK file may need a different conversion path; an animated file should be rejected when the product template expects one still image. Do not silently rotate or flatten during validation. Those are transformations, and transformations belong after the decision record is written.
Metadata is not trusted just because a library parsed it. Compare the declared MIME type with the file signature, cap decompression work, and reject impossible dimensions before allocating a large raster buffer. This is the same edge-case discipline I apply to OTP payloads: a small input can trigger a large downstream cost if the boundary is vague.
How do print-on-demand artwork metadata checks work before conversion?
- \1 detect the format from bytes, not the extension.
- \1 verify the upload hash and make sure the file is complete.
- \1 enforce width, height, aspect ratio, and frame count.
- \1 calculate effective PPI from pixels and the requested print box.
- \1 capture color model and profile; define an explicit conversion policy.
- \1 state whether alpha is allowed and what background is used when it is not.
- \1 attach product SKU, locale, consent/provenance fields, and a moderation decision ID.
That last gate is easy to skip because it is not a property of the image bytes. In a healthtech workflow, it is still part of the artwork contract. A photo can be technically perfect and belong to the wrong product.
Here is a compact Python boundary object. It does not convert anything; it produces a reviewable decision. The probe function is intentionally an adapter around your chosen decoder, so the business rules do not depend on one library.
from dataclasses import dataclass
from hashlib import sha256
from pathlib import Path
@dataclass
class ArtworkDecision:
accepted: bool
reasons: list[str]
metadata: dict
def validate_artwork(path: str, box_inches: tuple[float, float], probe) -> ArtworkDecision:
raw = Path(path).read_bytes()
digest = sha256(raw).hexdigest()
info = probe(raw)
reasons = []
if info["mime"] not in {"image/png", "image/jpeg", "image/webp"}:
reasons.append("unsupported media type")
if info["frames"] != 1:
reasons.append("artwork must contain exactly one frame")
width, height = info["width"], info["height"]
req_w, req_h = box_inches
effective_ppi = min(width / req_w, height / req_h)
if effective_ppi < 300:
reasons.append(f"effective PPI {effective_ppi:.1f} is below 300")
if info["alpha"] and not info["alpha_allowed"]:
reasons.append("alpha channel is not allowed for this template")
if info["profile"] not in {"sRGB", "Display P3"}:
reasons.append("color profile requires explicit review")
metadata = {
"sha256": digest,
"mime": info["mime"],
"width": width,
"height": height,
"effective_ppi": round(effective_ppi, 2),
"profile": info["profile"],
"alpha": info["alpha"],
}
return ArtworkDecision(not reasons, reasons, metadata)
The exact 300 PPI threshold is a template policy, not a universal law. Some printers publish different requirements. Store the policy version beside the decision so a later recheck does not guess which rule was active.
How should conversion, moderation, and retention work together?
Conversion is a state transition: received -> validated -> moderated -> converted -> proofed. Each transition should be idempotent on the input hash and policy version. If a worker retries after a timeout, it should read the existing decision instead of producing a second, slightly different derivative. A stable correlation ID should follow the asset through object storage, the queue, the converter, and the print proof.
Moderation coverage is the primary decision axis here. A background remover may make a clean cutout while leaving a prohibited claim, a dosage label, or a patient identifier untouched. Run moderation on the original and on the converted preview when the conversion changes visible pixels. Keep the moderation result separate from the technical metadata result; a pass on one is not a pass on the other.
The catch is retention. Keeping every original and every intermediate derivative makes incident review much easier, but it increases storage exposure for health-related imagery. When a print proof is challenged, an investigator usually needs a narrow chain: the original hash, the metadata snapshot, the policy version, the moderation decision, and the final proof hash. They do not need six abandoned resized copies sitting in the same bucket. A practical policy is to retain the original, the final proof, the hashes, and the decision record, while expiring disposable intermediates on a short, documented schedule. Expiry must be observable: emit an event, keep the object ID in the audit record, and make a restore request an explicit, authorized action. That is not suitable when a regulation or contract requires full reconstruction; in that case, stick with a longer, access-logged retention class and encrypt it.
Failure modes that deserve explicit tests
Test metadata, not just happy-path pixels. Include truncated files, a valid JPEG renamed as PNG, EXIF orientation set to rotate the long edge, a huge declared canvas with little compressed data, an ICC profile that the converter cannot preserve, and a PNG whose transparent pixels contain sensitive RGB values. Also test duplicate uploads with different names; the hash should make their identity obvious.
I like table-driven tests with an expected reason code such as META_DIMENSIONS_LOW or META_PROFILE_REVIEW. A 422 response is useful to an API client; a generic 400 is not. For queue workers, emit one structured event per decision and alert on a rise in review reasons, not only on 5xx counts. Your mileage may vary on thresholds, especially across printers, so make them configuration with an owner and an expiry date.
Choosing an implementation boundary
A self-hosted decoder gives control over data residency and exact library versions, at the cost of patching and capacity planning. A managed media service reduces that operational work, but you must verify its accepted formats, metadata preservation, region behavior, and deletion semantics. A command-line tool behind a sandbox can be a good middle ground when you need reproducible builds and a narrow attack surface.
Make the adapter contract small: probe(bytes) -> metadata, convert(bytes, policy) -> derivative, and delete(object_id). Keep vendor-specific fields out of the decision schema. This lets the team change an implementation without rewriting moderation, audit, or print-proof code.
Do not choose on unit price alone. The dominant cost is often review and reprint churn caused by ambiguous metadata, plus the retention and access controls required for sensitive photos. Measure rejection reasons, median time from upload to proof, duplicate conversion rate, and the percentage of assets needing manual color review. Those metrics tell you which gate to improve next.
A good gate is boring: it rejects with a reason, preserves evidence, and lets a safe file move on.
Top comments (0)