DEV Community

ColeMitchell4991
ColeMitchell4991

Posted on

How to Choose Image Conversion or Compression in Python — Property Photos

Short answer: to choose between image conversion and compression, convert when a target device or browser needs a different format or shape; compress when the format already works and the bytes are the problem. For a property-management system serving the same listing photo to a phone card, a desktop gallery, and a printed inspection packet, I usually make a small set of predictable derivatives at upload, then keep an on-demand path for an unplanned ratio. That keeps the hot request boring and leaves room to revise the crop policy later.

The distinction sounds tidy until a listing contains a phone portrait, a wide exterior shot, and a transparent floor-plan overlay. “Make it smaller” can mean changing pixels, changing encoding, or both. Those operations have different failure modes, so I want the pipeline to say which one it is doing.

Keep it explicit.

How should Python choose image conversion or compression for property photos?

Start with the output contract, not with a library call. A feed that promises 640x480 JPEG thumbnails needs a geometry decision (crop or letterbox), a format decision (conversion), and a quality decision (compression). A detail view that already accepts the source format may need only a byte budget. Treating all three as “optimization” hides the thing that will break.

Here is a deliberately small implementation. It uses Pillow because the example needs an executable local transform, while the policy remains independent of any particular storage or CDN. The input is a listing image and a requested ratio; the returned bytes are suitable for handing to an object-store adapter.

from io import BytesIO
from pathlib import Path
from PIL import Image, ImageOps


TARGETS = {
    "card": (4, 3, 640),
    "hero": (16, 9, 1600),
    "square": (1, 1, 800),
}


def smart_crop(source: Image.Image, ratio: tuple[int, int]) -> Image.Image:
    """Center-crop to a ratio; a saliency model can replace this policy later."""
    width, height = source.size
    wanted = ratio[0] / ratio[1]
    actual = width / height
    if actual > wanted:
        new_width = round(height * wanted)
        left = (width - new_width) // 2
        box = (left, 0, left + new_width, height)
    else:
        new_height = round(width / wanted)
        top = (height - new_height) // 2
        box = (0, top, width, top + new_height)
    return source.crop(box)


def render_derivative(path: str, target_name: str) -> tuple[bytes, str]:
    ratio_w, ratio_h, max_width = TARGETS[target_name]
    with Image.open(path) as opened:
        # EXIF orientation is metadata, but users judge the displayed pixels.
        image = ImageOps.exif_transpose(opened).convert("RGB")
        cropped = smart_crop(image, (ratio_w, ratio_h))
        scale = min(1.0, max_width / cropped.width)
        resized = cropped.resize(
            (round(cropped.width * scale), round(cropped.height * scale)),
            Image.Resampling.LANCZOS,
        )
        output = BytesIO()
        resized.save(output, format="JPEG", quality=82, optimize=True, progressive=True)
    return output.getvalue(), "image/jpeg"


payload, content_type = render_derivative("listing-1842.jpg", "hero")
Path("listing-1842-hero.jpg").write_bytes(payload)
print(content_type, len(payload))
Enter fullscreen mode Exit fullscreen mode

The judgment is explicit: crop only when the requested ratio differs, resize only when the width ceiling requires it, and encode to JPEG only because this derivative's contract says JPEG. For a source with an alpha channel, silently converting to RGB would discard meaningful pixels; that target should use a format and background policy that preserves transparency instead.

A first pass like this is useful in a notebook. Production needs a stable derivative key, such as (asset_id, policy_version, target_name), and an idempotent write. Store the original separately. Never overwrite it with a “better” crop; a new policy needs a new key so old listing pages do not change underneath a reviewer.

Upload-time derivatives versus on-demand rendering

Upload-time processing wins when the target set is small and known. Property cards, search results, and the default gallery tend to use the same ratios repeatedly, so generating those variants once avoids repeating CPU work during traffic spikes. It also lets an upload worker reject an unreadable file before the listing becomes visible. Picture the failure that prompted this rule: a leasing agent uploads a tall phone photo, the listing is published immediately, and three clients then ask for three different widths at once. If each request decodes the original, crops it, resizes it, and writes a cache entry, the page may look fine in a quiet test while a busy release burns the same CPU repeatedly. A precomputed card derivative makes that common path a read, while an explicit queue handles the unusual export. The important part is the key: source hash plus policy version plus target name. Without the policy version, a crop-policy change can serve yesterday's pixels under today's URL, which is much harder to diagnose than a slow job.

On-demand rendering wins when the request space is open-ended: an export tool may ask for a custom width, an agent may request a square contact sheet, or a new device layout may appear next month. Cache the result using the source content hash, policy version, ratio, and output format. A cache miss should enqueue work or return a clear pending state; it should not make the listing API wait on a multi-megapixel decode.

The hybrid rule is practical: precompute the three variants that product analytics can defend, and render everything else on demand with a bounded queue. Your mileage may vary when storage is unusually expensive or when uploads are rare but exports are massive. Measure derivative hit rate and worker time before moving the boundary.

There is a catch. Upload-time crops spend work on photos that may never be viewed, while on-demand work can create a thundering herd for a newly published listing. A small lock keyed by the derivative key, plus a short negative-cache entry for invalid requests, prevents duplicate jobs without hiding permanent input errors.

Conversion, compression, and the quality budget

Conversion answers “which representation can the consumer decode?” Compression answers “how much information and how many bytes can this representation use?” A browser-compatible format is useless if a 12 MB original blocks a mobile listing page; a tiny file is useless if a floor-plan label becomes unreadable.

Use the source type and the visual job to set a budget. Photographs with continuous tones generally tolerate lossy compression. Screenshots, maps, and text-heavy plans need a lossless or visually conservative path. Animated content and transparency add separate constraints. The Media Formats Guide from MDN is a useful compatibility reference, but it does not choose your business threshold for you.

I keep quality decisions in a policy object rather than scattering numbers through handlers:

from dataclasses import dataclass


@dataclass(frozen=True)
class OutputPolicy:
    format: str
    quality: int | None
    max_bytes: int


POLICIES = {
    "listing_photo": OutputPolicy("JPEG", 82, 450_000),
    "floor_plan": OutputPolicy("PNG", None, 900_000),
}
Enter fullscreen mode Exit fullscreen mode

The numbers are starting points, not universal truths. I would run a small eval set of exterior shots, dark interiors, room numbers, and plan annotations; compare visual defects at the actual card size; then record the chosen policy with its version. A file-size histogram alone cannot tell you that a “successful” crop removed the front door.

Testing the crop before users find the mistake

Image tests should include geometry and meaning. Geometry checks can assert that a 4:3 derivative is exactly 640x480, that EXIF rotation is applied once, and that an oversized source is reduced. Meaning checks need fixtures with a known focal subject near each edge. A center crop is deterministic, but it is not smart for every listing.

I use a tiny contract test before connecting a queue or object store:

from io import BytesIO
from PIL import Image


def test_card_contract(tmp_path):
    source = Image.new("RGB", (2400, 1600), "white")
    source_path = tmp_path / "source.jpg"
    source.save(source_path, format="JPEG")

    data, media_type = render_derivative(str(source_path), "card")
    result = Image.open(BytesIO(data))

    assert media_type == "image/jpeg"
    assert result.size == (640, 480)
Enter fullscreen mode Exit fullscreen mode

The test does not prove that a person is still visible. That belongs in a reviewed fixture set or a crop-scoring eval, where a failed score blocks a policy rollout rather than silently changing every listing. I started by assuming one center crop would be enough; edge-heavy exterior photos changed my mind. Keep that kind of correction in the test corpus.

Operational checklist for a safe rollout

Give every derivative an owner and a deletion rule. Record source hash, policy version, dimensions, media type, byte size, processing duration, and a reason for rejection. Alert on queue age and decode failures, but sample a few rendered images in an internal review page; metrics will not show a missing balcony.

Limit decoder work by pixels, not only by file bytes. A highly compressed image can still expand to an unsafe canvas. Validate the requested ratio and format against an allowlist, strip untrusted metadata when policy permits, and keep the original outside the public path. Retry transient storage writes, not malformed image input.

The recommendation is unsuitable when every request is a one-off creative crop or when the source must remain pixel-perfect for legal records. In those cases, stick with on-demand, lossless outputs and an explicit approval step. For ordinary listing photos, the hybrid boundary gives a predictable upload cost, fast common reads, and a place to experiment with better crop scoring without rewriting the serving API.

References

Top comments (0)