Short answer: to keep an image archive safe, don't compress originals with lossy settings; compress only replaceable derivatives. That costs more durable storage, but it prevents an irreversible encoding decision from becoming the ceiling for every future crop, zoom, and short promo video; control the bill with retention classes, deduplication, bounded renditions, and cache policy instead of repeatedly rewriting the only master.
The production scenario is deliberately narrow: a B2B SaaS system accepts product photos, generates short promotional videos from a prompt, and serves preview images along the way. If I am reviewing an incident in that system, I first ask whether the damaged object can be regenerated. A soft preview can be rebuilt from a master. A master overwritten by lossy output cannot. The invariant is therefore simple: irreversible work belongs on replaceable bytes.
Should an image archive compress originals or only derivatives?
The tempting sequence is upload, decode, resize, encode, and overwrite. It saves visible capacity immediately. It also merges ingestion with delivery policy, even though those two lifecycles change at different speeds. The archive now contains the output of yesterday's codec settings rather than the user's input.
That distinction matters because lossy formats discard information to reduce size, while lossless formats preserve the reconstructed image data. MDN's image format guide makes the boundary explicit and also shows that format choice depends on the image and use case. A product-photo master may later feed a square catalog tile, a detail crop, a color-sensitive approval screen, or frames in a generated promo. None of those downstream requests can recover information already discarded upstream.
One bad write is enough.
I initially find the storage graph persuasive when it shows one large master beside several small variants. Then I redraw the graph around failure domains: one master is the recovery source; every derivative is cacheable, expirable, and reproducible. The apparent duplicate bytes are not all carrying the same operational value.
Model capacity before choosing an encoding policy
Do not start with a codec. Start with object counts and service objectives. Let M be retained master bytes, D the average bytes across retained derivatives, C the cache footprint, and N the number of photos. The planning baseline is N * (M + D) + C, plus whatever headroom the storage system requires. This is a capacity model, not a price forecast.
For a concrete but hypothetical review, consider 1,000,000 photos, three retained delivery variants per photo, and a temporary set of video-render intermediates. I would run the model with measured p50 and p95 object sizes from our own corpus, not a sample image from a codec demo. The worksheet needs separate rows for uploaded masters, current derivatives, superseded derivatives awaiting expiration, render intermediates, and cache copies; otherwise a tidy average conceals the multiplication caused by variant count. Next I would replay two conditions against the same model: ordinary catalog browsing, then a campaign that makes old photos hot while new video jobs are arriving. That second condition is the one people skip. Cache misses trigger derivative reads or regeneration, regeneration competes for decode and encode capacity, and video assembly needs some of the same source objects. The model should state which work is shed first, how old a preview may become, and how long the regeneration queue may grow before the SLO is missed. None of those values should come from a generic recommendation. They have to come from measured object sizes, observed access distributions, and the product's own tolerance for stale previews. Only after those inputs are visible would I compare the durable master footprint with the compute and cache capacity required to rebuild derivatives. The review is skeptical by design: deleting recoverable output is a capacity action, while degrading the sole recovery input is a data-policy change.
The SLO needs separate indicators: master durability and readability, derivative generation success, preview latency, and regeneration backlog age. A single availability percentage hides the important failure mode. Serving a stale preview while regeneration catches up can be acceptable; silently returning a degraded master as the new source is not.
| Policy | Durable capacity | Recovery posture | Cache behavior | On-call consequence |
|---|---|---|---|---|
| Lossy overwrite of upload | Lowest initial footprint | No pristine recovery source | Fewer object classes | Quality incidents become data incidents |
| Immutable master, bounded derivatives | Higher master footprint | Derivatives are rebuildable | Hot outputs can be cached independently | Rebuild load must be capacity-planned |
| Lossless-normalized master, bounded derivatives | Corpus-dependent | Original file bytes are no longer retained | Same as above | Normalization must be proven acceptable |
The third row is sometimes reasonable, but call it what it is: a migration from original-file retention to pixel-content retention. Metadata, embedded profiles, animation, and other format capabilities can matter, so the acceptance test must reflect the archive contract. MDN's format guide is useful here because it describes capabilities and browser support without pretending that one format is universally correct.
A preventative Go path with an immutable boundary
The write path should make the safety rule difficult to violate. This example uses generic interfaces, content-derived identity, conditional creation, and a manifest that points from one master to its replaceable variants. The storage implementation can be managed or self-hosted; the invariant stays in application code and tests.
package archive
import (
"context"
"crypto/sha256"
"encoding/hex"
"fmt"
)
type Store interface {
PutIfAbsent(ctx context.Context, key string, body []byte, contentType string) error
PutReplaceable(ctx context.Context, key string, body []byte, contentType string) error
}
type Encoder interface {
Encode(source []byte, width int, quality int) ([]byte, string, error)
}
type Variant struct {
Width int
Quality int
}
func Ingest(ctx context.Context, s Store, e Encoder, upload []byte, mediaType string, plan []Variant) (string, error) {
sum := sha256.Sum256(upload)
id := hex.EncodeToString(sum[:])
masterKey := fmt.Sprintf("masters/%s", id)
// Conditional creation prevents a retry or later transform from replacing the master.
if err := s.PutIfAbsent(ctx, masterKey, upload, mediaType); err != nil {
return "", fmt.Errorf("store master: %w", err)
}
for _, v := range plan {
encoded, outType, err := e.Encode(upload, v.Width, v.Quality)
if err != nil {
return id, fmt.Errorf("encode width %d: %w", v.Width, err)
}
key := fmt.Sprintf("derivatives/%s/w%d-q%d", id, v.Width, v.Quality)
if err := s.PutReplaceable(ctx, key, encoded, outType); err != nil {
return id, fmt.Errorf("store derivative: %w", err)
}
}
return id, nil
}
The code intentionally does not name a format. The encoder policy belongs in a versioned rendition plan, because browser support, transparency, animation, decode behavior, and photographic content all affect the choice. A derivative key should also include every transformation input that changes bytes; the abbreviated key above is safe only if width and quality are the complete plan. In a real system I would include encoder version, crop policy, color policy, and output format in the identity or manifest.
Retries deserve special attention. Master creation is conditional and idempotent. Derivative writes are replaceable because they are outputs, although publishing should still avoid exposing partial sets. Validate decode before acknowledging ingestion, record the master checksum, and periodically prove that a sampled master can be read and transformed. A checksum that is never checked is inventory, not assurance.
Where should storage and cache controls live?
Retention is the cheaper lever conceptually because it removes work with no audience. Do not generate every conceivable width at ingest. Keep a small, evidenced rendition set for predictable paths, generate uncommon variants on demand, and expire render intermediates after their defined recovery window. Cache keys must distinguish transformation inputs, while cache-control policy should allow immutable derivatives to remain hot without pinning obsolete plans forever.
This is also a buy-versus-build decision, but the codec is a small part of it. The platform boundary should be judged on durability controls, conditional writes, lifecycle rules, auditability, egress exposure, restore testing, and on-call burden.
| Question | Managed storage | Self-hosted storage |
|---|---|---|
| Who operates replication and repair? | Service boundary, verified through its contract | Platform team, verified through drills |
| Who owns lifecycle mistakes? | Application and policy owner still do | Application and platform teams do |
| How is lock-in limited? | Portable keys, manifests, checksums, and export tests | Standard interfaces and tested restore paths |
| What consumes engineering capacity? | Integration, policy review, and cost controls | Those tasks plus fleet operation and upgrades |
No column wins by default. If the archive is small but the team is smaller, operating storage may cost more attention than it saves in invoices. If access patterns make transfer dominate, locality and cache topology may outweigh raw retained bytes. Measure both, then reserve headroom for regeneration so an eviction wave cannot consume the capacity needed by live video jobs.
When is changing the master acceptable?
Preserving the uploaded file is not a religious rule. A controlled normalization can be appropriate when the contract defines the master as decoded visual content rather than the exact uploaded bytes, the accepted formats and metadata behavior are explicit, and a representative corpus passes visual and functional acceptance tests. Lossless compression can also reduce some files without the same information-discarding trade-off, although the result depends on format and content.
There is another exception: regulated deletion. An immutable application convention must not prevent an authorized retention or deletion workflow. Immutability means ordinary transforms cannot mutate the source; it does not mean data lives forever.
The decision rule I use is blunt: if a future output may need information that the proposed transform removes, keep the pre-transform object under a defined retention policy. Compress delivery variants aggressively enough to meet latency and cache objectives, but make them disposable. Then test the recovery path under load, because an archive that is theoretically reversible and operationally unable to regenerate before its SLO expires is only half designed.
Top comments (0)