The page says compressed product images look blurry enough to breach the marketplace thumbnail SLO, and the on-call sees a healthy queue, successful encoders, and a bandwidth graph that looks almost suspiciously good. The least complex fix is to stop treating every product asset as a photograph: classify graphics separately, preserve transparency and hard edges for them, and tune lossy photo compression against a measured visual-quality floor rather than one global quality number.
Short answer: record the source type, dimensions, alpha channel, chosen output format, encoded bytes, and a quality score for every transformation. Use those fields to isolate the bad class before changing an encoder setting. A low byte count is not success when search results have unreadable seller marks.
This is a quality-versus-bandwidth capacity problem. It only looks like an encoder problem after the page fires.
Why do compressed product images look blurry after a quality change?
The useful early signal is not a process error rate. A codec can produce a valid file while destroying the detail that matters, so the service-level indicator needs to join technical validity with perceptual usefulness. For a marketplace media library, I would define separate indicators for photo-like assets and graphic assets, because their failure surfaces differ: photographs tolerate gradual loss in texture, while a logo can become visibly wrong when a narrow stroke or transparent boundary is disturbed.
Do not invent one universal quality threshold. Build a small, reviewed reference set from your own catalog, including dark products, fine fabric, text-bearing packaging, transparent seller marks, and flat-color graphics. Keep the originals. For each candidate policy, compare transformed images at the actual display dimensions and record both encoded bytes and the selected objective metric. Human review remains the release gate because an objective score is a proxy, not the business definition of acceptable.
The page should represent a sustained breach of the class-specific visual-quality SLO, not a single difficult asset. Ticket or dashboard the leading indicators: a shift in bytes per output pixel, growth in upscaling, unexpected alpha removal, or a sudden change in the distribution of chosen formats. Those signals explain the eventual quality breach and give the on-call somewhere concrete to look.
One detail matters operationally: preserve the asset class as a low-cardinality label, but do not put filenames, seller IDs, or content hashes into metric labels. Put per-image identifiers in structured logs or traces instead. Prometheus documents why every unique label set creates another time series; a catalog-sized identifier in metrics turns a useful alert into a capacity incident of its own.
Trace the alert back to one transformation decision
Start with one affected image and reconstruct the decision, rather than nudging a global slider and waiting for the next complaint. The transformation record should answer six questions: what pixels arrived, whether alpha was present, how the asset was classified, whether it was resized or enlarged, which output family was selected, and how many bytes left the encoder.
A compact event is enough:
{
"asset_class": "graphic",
"source_format": "png",
"source_width": 800,
"source_height": 320,
"has_alpha": true,
"output_format": "png",
"output_width": 400,
"output_height": 160,
"encoded_bytes": 18420,
"policy_version": "catalog-v3"
}
Dimensions such as 800 by 320 above are example data, not a promised threshold. The important field is policy_version: without it, a deployment and a catalog shift are hard to distinguish during an incident.
Then inspect the source at 100% scale and the delivered derivative at its CSS display size. If the derivative was enlarged, fix the requested dimensions or srcset selection before touching compression. If alpha vanished, follow the format-selection branch. If only graphics are affected, inspect classification. If only photos are affected, compare the encoded result with the reference set at adjacent quality levels. This branching order prevents a format or sizing error from being disguised by a higher bitrate.
Instrument the Go image path
The following program is deliberately small and standard-library-only. It decodes an input, identifies transparency, accepts an explicit class, encodes photos as JPEG and graphics as PNG, then emits the fields needed for the first diagnostic pass. It does not resize; keeping resampling out makes the compression decision observable on its own. This approach is not suitable when the delivery contract requires modern formats or dynamic resizing: use an encoder and resampler that implement those requirements, while retaining the same event fields and class-specific tests. The trade-off is extra dependency and upgrade ownership in exchange for more format choices and tighter bandwidth control.
package main
import (
"encoding/json"
"errors"
"flag"
"fmt"
"image"
"image/jpeg"
"image/png"
"io"
"os"
)
type Event struct {
AssetClass string `json:"asset_class"`
SourceFormat string `json:"source_format"`
Width int `json:"width"`
Height int `json:"height"`
HasAlpha bool `json:"has_alpha"`
OutputFormat string `json:"output_format"`
EncodedBytes int64 `json:"encoded_bytes"`
PolicyVersion string `json:"policy_version"`
}
type countingWriter struct {
w io.Writer
n int64
}
func (c *countingWriter) Write(p []byte) (int, error) {
n, err := c.w.Write(p)
c.n += int64(n)
return n, err
}
func alphaPresent(img image.Image) bool {
b := img.Bounds()
for y := b.Min.Y; y < b.Max.Y; y++ {
for x := b.Min.X; x < b.Max.X; x++ {
_, _, _, a := img.At(x, y).RGBA()
if a != 0xffff {
return true
}
}
}
return false
}
func encode(out io.Writer, img image.Image, class string) (string, error) {
switch class {
case "photo":
return "jpeg", jpeg.Encode(out, img, &jpeg.Options{Quality: 82})
case "graphic":
enc := png.Encoder{CompressionLevel: png.DefaultCompression}
return "png", enc.Encode(out, img)
default:
return "", errors.New("class must be photo or graphic")
}
}
func main() {
class := flag.String("class", "", "photo or graphic")
input := flag.String("in", "", "source image")
output := flag.String("out", "", "encoded image")
flag.Parse()
src, err := os.Open(*input)
if err != nil {
panic(err)
}
defer src.Close()
img, sourceFormat, err := image.Decode(src)
if err != nil {
panic(err)
}
dst, err := os.Create(*output)
if err != nil {
panic(err)
}
defer dst.Close()
counted := &countingWriter{w: dst}
outputFormat, err := encode(counted, img, *class)
if err != nil {
panic(err)
}
b := img.Bounds()
event := Event{
AssetClass: *class, SourceFormat: sourceFormat,
Width: b.Dx(), Height: b.Dy(), HasAlpha: alphaPresent(img),
OutputFormat: outputFormat, EncodedBytes: counted.n,
PolicyVersion: "catalog-v3",
}
if err := json.NewEncoder(os.Stdout).Encode(event); err != nil {
panic(fmt.Errorf("encode event: %w", err))
}
}
Run it against a representative source after saving it as main.go:
go run main.go -class photo -in product.jpg -out derivative.jpg
go run main.go -class graphic -in seller-mark.png -out derivative.png
The value 82 is a starting point for an experiment, not a generally correct JPEG quality. The Go documentation defines the accepted JPEG quality range as 1 through 100 and uses 75 when the options are nil, but the number is encoder-specific and says nothing by itself about visible acceptability. Calibrate it against the reference set. Also note that alphaPresent scans every pixel; on a high-throughput path, emit alpha information from an existing decode or analysis stage instead of paying for another full traversal. Capacity planning includes CPU and memory bandwidth, not only egress.
Set policies by content, then test the boundary
A policy should express intent, not merely codec knobs. photo can permit controlled lossy encoding after a quality check. graphic should default to preserving alpha and crisp boundaries, with a lossless path until testing proves another path acceptable. Misclassification must be visible because a perfect implementation of the wrong policy is still wrong.
| Decision | Photo-like product image | Logo or flat graphic |
|---|---|---|
| Detail at risk | texture, gradients, small packaging text | hard edges, thin strokes, transparency |
| First diagnostic | delivered dimensions and lossy quality | alpha, palette changes, classification |
| Release evidence | reference-set metric plus visual review | pixel-level or lossless check plus visual review |
| Bandwidth response | search within the accepted quality floor | resize correctly; avoid assuming lossy output is safe |
Test the boundary cases, not just median catalog images. A one-pixel transparent edge, a narrow dark stroke on a light field, repeated textile texture, and readable label text each exercise a different weakness. Decode the produced file in tests and assert its dimensions and expected alpha behavior. Keep golden outputs only when the encoder version and generation process are controlled; otherwise a semantic assertion is less brittle than a byte-for-byte comparison.
For rollout, shadow-encode a representative sample, calculate the added CPU and storage demand, and compare output distributions before serving the new derivatives. Then canary by a stable partition and watch both the visual-quality indicator and bytes per output pixel. Roll back by policy version, not by hand-editing quality values on live instances. This is where the platform decision becomes a buy-versus-build question rather than a codec preference.
| Option | On-call burden | Control | Lock-in surface | Best fit |
|---|---|---|---|---|
| Build the pipeline | team owns decoding, resampling, encoding, rollout, and evidence | highest | codecs, object schema, and internal policy | image behavior is core and staffing covers the pager |
| Operate self-hosted components | team owns capacity and upgrades; library behavior is inspectable | high | component APIs and stored derivatives | control matters more than minimizing operations |
| Use a managed transformation service | provider owns much of fleet operation; team still owns quality policy and monitoring | constrained by exposed controls | URLs, policy syntax, caches, and migration of derivatives | reducing on-call load outweighs custom control |
No row wins universally. Estimate peak transforms per second, source-pixel volume, cache-miss behavior, memory per worker, and reprocessing time for the full catalog. A service that meets steady-state traffic but needs days to regenerate derivatives after a policy correction has hidden recovery risk.
Close the loop without creating a noisy pager
The final trap is making the quality alert so sensitive that ordinary image diversity wakes someone up. Objective metrics vary by content, and marketplace uploads are adversarial in the mundane sense: screenshots, scans, tiny images, mislabeled formats, and assets that already contain compression damage all arrive through the same door. Segment the dashboard before setting the page threshold, exclude transformations that cannot meet the target because the source is undersized, and route isolated asset failures to a queue rather than the primary pager.
Page only when the error-budget consumption indicates a user-impacting, sustained policy failure. Use a ticket for a slow drift in bytes per pixel or classification mix. Keep individual decode failures in logs with bounded reason codes and sample identifiers. This division preserves the page as a demand for immediate action.
A threshold set too low hides real blur; one set too high converts harmless catalog variance into interruptions, eventually teaching the on-call to distrust the signal. That false-positive cost belongs in the design review beside bandwidth and visual quality. The correct policy is the one that preserves the marketplace's reviewed quality floor, stays within its capacity envelope, and produces an alert with an executable first step.
Top comments (1)
woоow