DEV Community

BrennanCross2167
BrennanCross2167

Posted on

Evidence Intake Pipelines That Keep Originals While Generating Review Derivatives

Customer-support archives have an awkward constraint: an agent needs a fast crop for a ticket, while the legal record must retain the exact image that arrived. The operational choice is to make the original an append-only object and treat every crop, resize, or format conversion as a separately identified review derivative.

Short answer: hash and store the incoming bytes first, then create multi-ratio review copies that point back to that immutable record; never make the cropped file the source of a later crop.

That boundary matters when a phone photo contains a serial number at the edge of the frame. A 1:1 crop can be perfect for a support queue and still be the wrong artifact for an investigator. Keep both, with a relationship an auditor can follow.

Keep it immutable.

How should evidence image intake preserve originals while producing review copies?

Start with an intake transaction that has one job: establish what was received. Capture the detected media type, byte count, acquisition timestamp, case identifier, and a SHA-256 digest. The digest is an identity check, not a claim that two visually similar images are the same. Store the object under an opaque ID and make that ID the parent of all later work.

The derivative record should describe an operation, not just a URL. Record the parent ID, requested aspect ratio, crop anchor, output dimensions, output format, and the processor version. A support agent might request 4:5 for a ticket panel, 16:9 for a handoff, and a square thumbnail for a search result. Those are three children of one source. If one needs regeneration, create a new child and keep the old record available for audit.

Here is a deliberately small Python model and queue handoff. It leaves storage and image libraries behind an interface, which keeps the evidence rules independent of a particular vendor.

from dataclasses import dataclass
from hashlib import sha256
from typing import Optional


@dataclass(frozen=True)
class Original:
    object_id: str
    media_type: str
    byte_count: int
    digest: str


@dataclass(frozen=True)
class ReviewCopy:
    object_id: str
    parent_id: str
    ratio: str
    width: int
    height: int
    operation: str


def intake_original(data: bytes, media_type: str, object_id: str) -> Original:
    digest = sha256(data).hexdigest()
    return Original(object_id, media_type, len(data), digest)


def enqueue_crop(original: Original, ratio: str, width: int) -> dict:
    return {
        "parent_id": original.object_id,
        "ratio": ratio,
        "width": width,
        "preserve_source": True,
    }
Enter fullscreen mode Exit fullscreen mode

The important field is parent_id. preserve_source is an invariant your worker must enforce, not a checkbox the browser gets to reinterpret. Commit the original record before publishing this job. If a queue retry runs twice, idempotency should deduplicate the same requested derivative without replacing the parent.

Which crop rules keep a support review useful without changing the record?

Cropping is a product decision disguised as a media operation. Define an anchor policy for each aspect ratio: center, detected subject, or a manually supplied focal point. Preserve the full source dimensions in metadata so an agent can open the original when a crop removes context. For screenshots and documents, “subject detection” can be actively misleading; text near an edge is often the evidence.

Build a small acceptance corpus from the support channel. Include rotated phone photos, dark scenes, transparent PNGs, screenshots with tiny labels, and the largest permitted upload. Add images where the relevant detail sits at each corner, because a center-biased algorithm can pass a happy-path demo while deleting the only useful clue. Include a contact sheet with several objects, a receipt whose text runs along the border, and a photo with an intentional blank margin; those cases reveal whether the crop policy is predictable or merely attractive. For every class, write an unacceptable result before testing: clipped text, missing transparency, changed orientation, or a crop that hides the object under discussion. MDN’s media-format guide is a useful reference for browser compatibility, but your corpus decides whether a chosen output is readable. You don't need a perfect benchmark; you need fixtures that make a regression obvious in review.

I once approved a center crop because it looked fine in a dashboard mock-up. On the first real ticket, the damage was in the far-right corner and vanished from the 1:1 preview. That was a design error, not a bad JPEG. The fix was to preserve the source and add a “full frame” action beside every derivative.

Keep quality and transfer size as separate measurements. A 1600-pixel review image may be sensible for reading a label; a 320-pixel thumbnail may be right for search. Do not let a bandwidth optimization silently become an evidentiary decision. A derivative can be lossy. The parent cannot.

What lifecycle checks expose missing or misleading image copies?

Validate at each boundary. At upload, compare declared and detected media types and calculate the digest before the object is visible to reviewers. During processing, create a pending derivative row, then mark it ready only after the output can be opened and its dimensions match the request. On retrieval, authorize by case and artifact ID, not by a filename that a user can guess.

Make failure states boring and visible. A worker timeout leaves a pending row for reconciliation; it does not trigger an overwrite retry. If a crop cannot be generated, the original remains available and the ticket says which review size is missing. Retention jobs need the same split: deleting a thumbnail is a different event from deleting the source, with different approval requirements.

Useful metrics include accepted originals, pending derivatives, completed derivatives by ratio, and requests that fall back to the full frame. Alert on a queue that grows for ten minutes or on a sudden drop in one ratio. Logs should include opaque IDs, operation names, and status codes. They should not include image bytes, signed URLs, or personal content.

A compact event shape helps incident review:

event = {
    "case_id": "case-2048",
    "artifact_id": "review-4x5-7f2c",
    "parent_id": "orig-a91e",
    "state": "ready",
    "operation": "smart_crop",
    "ratio": "4:5",
}
Enter fullscreen mode Exit fullscreen mode

The event is evidence of a transformation, not a replacement for the source digest. Keep the two concepts distinct in dashboards and exports.

When is an immutable-original pipeline the wrong fit?

The catch is operational weight. Originals plus several derivatives require a retention schedule, access controls, restore tests, and a clear policy for legal holds. This approach is not suitable when the product has no need to retain source media or when every upload must disappear immediately after a transient preview. A short-lived processing object may be enough there.

It is also a poor fit for workflows that promise pixel-identical display across every browser. Media codecs and color management vary; your mileage may vary, and the acceptance corpus should include the browsers your support team actually uses. If a jurisdiction requires a signed evidence package, add that requirement explicitly rather than assuming a hash alone satisfies it.

The practical decision rule is narrow: preserve first, derive second, and make every derivative traceable to one parent. Roll out with one ratio, compare agent task time and crop misses, then add more ratios without changing the original schema. Keep the full-frame path one click away.

References

Top comments (0)