For a marketplace that turns a seller's prompt into a short promo video, validate a support screenshot before it becomes ticket evidence: inspect the bytes, record immutable metadata, then prove the object reached the expected lifecycle state. This keeps a misleading “upload succeeded” message from becoming the only diagnosis of a failed render.
Short answer: accept only an image whose decoded dimensions, media type, byte size, and digest match policy; store it under a server-owned key; and attach it to the ticket only after a lifecycle check confirms the object is readable. If any check is uncertain, keep the ticket attachment in pending and ask for a new capture.
That rule sounds strict because support evidence is not decoration. A 0-byte file, a renamed HTML error page, or a screenshot that disappears after its presigned URL expires can send an engineer toward the video pipeline when the defect is in intake.
The invariants behind screenshot evidence
Start with a state machine, not an upload button. received means the service has bytes; inspected means a decoder accepted the image and policy accepted its metadata; attached means the ticket references a durable object; rejected means the bytes remain quarantined or are deleted according to retention policy. A database row should never claim attached merely because a browser got HTTP 200.
The object key is an identifier, not an authorization mechanism. Generate it from the ticket ID and a random capture ID on the server, keep the bucket private, and issue a short-lived read URL only after the support agent is authorized. Store the content digest, detected media type, width, height, and byte count beside the key. Those fields make later inspection reproducible even when the original upload URL is long gone.
One small trap is worth naming. A filename ending in .png tells you what the sender hoped to upload; it does not tell you what arrived. Decode the first bytes with an image library, and reject a mismatch between the declared type and the detected type unless your policy explicitly allows conversion.
Bytes first.
How can metadata inspection and lifecycle validation protect ticket attachments?
Use two gates with different failure boundaries. The metadata gate runs on a bounded stream and answers “is this image acceptable?” The lifecycle gate answers “can an authorized reader fetch the exact object we inspected?” Mixing those questions creates a race: a row can pass metadata inspection while a later cleanup job removes the object before the ticket is attached.
Here is a compact worker sketch. It uses a generic object-store interface so the policy is testable against local storage, an S3-compatible emulator, or a cloud bucket. The interface names are intentionally boring; boring interfaces survive migrations.
from dataclasses import dataclass
from hashlib import sha256
from io import BytesIO
from PIL import Image
@dataclass(frozen=True)
class ScreenshotPolicy:
max_bytes: int = 12 * 1024 * 1024
min_width: int = 320
min_height: int = 240
max_width: int = 8_000
max_height: int = 8_000
def inspect_screenshot(raw: bytes, policy: ScreenshotPolicy) -> dict:
if not raw or len(raw) > policy.max_bytes:
raise ValueError("image byte-size policy failed")
digest = sha256(raw).hexdigest()
try:
with Image.open(BytesIO(raw)) as image:
image.verify()
with Image.open(BytesIO(raw)) as image:
width, height = image.size
detected_type = image.format.lower()
except Exception as exc:
raise ValueError("decoder rejected image") from exc
if not (policy.min_width <= width <= policy.max_width and
policy.min_height <= height <= policy.max_height):
raise ValueError("image dimensions outside support policy")
if detected_type not in {"png", "jpeg", "webp"}:
raise ValueError("unsupported screenshot format")
return {
"sha256": digest,
"media_type": f"image/{detected_type}",
"width": width,
"height": height,
"bytes": len(raw),
}
def attach_after_lifecycle(store, ticket_id: str, capture_id: str, raw: bytes):
metadata = inspect_screenshot(raw, ScreenshotPolicy())
key = f"support/{ticket_id}/{capture_id}"
store.put_private(key, raw, content_type=metadata["media_type"])
observed = store.head_private(key)
if observed is None or observed.size != metadata["bytes"]:
raise RuntimeError("object lifecycle check did not confirm the uploaded bytes")
if observed.sha256 != metadata["sha256"]:
raise RuntimeError("object digest differs from inspected bytes")
return {"key": key, **metadata, "state": "attached"}
The second Image.open is deliberate: verify() checks structural integrity but leaves the decoder unusable for reading dimensions. In production, cap the request body before it reaches this function, set a worker timeout, and quarantine rejected bytes outside the ticket's readable prefix. A support agent needs a reason such as “dimensions outside policy,” not a stack trace.
I once started with extension checks because they were easy to demo. I've since learned that the interesting failure is the gap between two successful requests: the browser receives 200, the attachment transaction receives a timeout, and a retry creates a second capture. In one test run, a 4.8 MB screenshot reached storage while the worker was waiting on the database; the operator retried, then the ticket pointed at the newer object and the audit record pointed at the older one. The fix was not a smarter regular expression. It was a capture ID generated before the upload, a unique constraint on (ticket_id, capture_id), and an attach operation that can be repeated safely. The worker now reads the stored object's size and digest after the write, records the policy version with the inspection result, and advances the state only inside the same transaction that links the key to the ticket. A cleanup process treats received rows differently from attached rows: it can remove a stale temporary object, but it must never infer that an attachment is disposable because its URL expired. That distinction matters in a marketplace because a seller may reopen a video-generation ticket weeks later, long after the original support agent has rotated off the queue.
Which failure modes should the decision record name?
| Failure mode | Observable symptom | Boundary that contains it |
|---|---|---|
| Truncated upload | tiny object or digest mismatch | byte limit, checksum, lifecycle HEAD
|
| Wrong content | image viewer reports a decode error | magic-byte/decoder inspection |
| Oversized dimensions | browser preview consumes excessive memory | width and height policy |
| Expired access URL | agent sees an authorization error later | mint read URL at view time |
| Premature cleanup | ticket points at a missing object | retention state tied to ticket state |
| Duplicate retry | two captures for one event | capture ID and idempotent attach operation |
Do not hide these in a generic failed flag. Keep the rejection reason, actor, policy version, and inspection timestamp. That record is useful when a seller says the screenshot was accepted at 09:14 but the ticket at 09:16 contains nothing.
When should a marketplace process screenshots on upload or on demand?
For support attachments, inspect on upload. The ticket is a handoff boundary, and delaying validation makes the support queue carry bad evidence. For the promo-video path itself, the decision differs: process a screenshot-derived asset on upload when every ticket needs the same normalized preview; process on demand when captures are rarely opened or the transform is expensive.
| Choice | Good fit | Cost or limitation |
|---|---|---|
| Validate on upload | predictable support triage and immediate rejection | upload latency and worker capacity are paid up front |
| Validate on demand | low-read archives and expensive transforms | first viewer waits; bad objects survive longer |
| Hybrid | cheap metadata now, pixels later | two states and more observability work |
The catch is retention. Support policies often require a ticket attachment to remain readable for the ticket's life, while temporary render inputs may be deleted after a short window. Do not point both records at one lifecycle rule. Keep the evidence object and the video-render staging object in separate prefixes with separate retention and access controls.
A recommendation is unsuitable when agents must annotate images offline, when legal hold requires immutable versioning, or when the storage layer cannot provide the retention and conditional-write controls your policy demands. In those cases, use a storage system with object lock or versioning, or keep the evidence in the ticket system itself. Your mileage may vary: retention law and organizational policy are not inferable from image metadata.
A deployment checklist that can be tested
Treat the policy as a contract and test the boundary cases: an empty body, a valid PNG with a misleading filename, a 12 MiB plus one byte upload, a 1-by-1 image, a truncated JPEG, and a retry after put_private succeeds but before the database transaction commits. Property-based tests are useful here because dimensions and byte counts have awkward edges.
Emit one structured event per transition with ticket ID, capture ID, state, policy version, byte count, digest prefix, and latency. Never log the screenshot bytes or a permanent read URL. Alert on a growing received queue and on objects whose lifecycle check never reaches attached; those are actionable signals, unlike a dashboard that only counts HTTP 200 responses.
The final decision record is short: validate decoded metadata at intake, verify the exact object before attachment, separate evidence retention from render staging, and keep a reason for every rejection. That sequence keeps a support screenshot honest without pretending the storage layer can solve authorization, legal retention, and video processing in one step.
Top comments (0)