Short answer: when an avatar crop cuts off heads, preserve one normalized source, reproduce the exact crop from recorded coordinates, and use center crop only after proving that the face is near the center; otherwise use a bounded smart crop with a manual adjustment as the final authority. For a logistics platform that removes backgrounds from staff, courier, or seller profile photos, run that decision at upload so every consumer receives the same reviewed avatar, while retaining enough source material and audit data to correct a bad framing decision.
The visible defect is a missing forehead. The expensive part is the asset policy behind it. If every upload produces a normalized foreground image plus three avatar sizes, the storage term is N x (S + F + V1 + V2 + V3), where N is the number of accepted uploads and each other variable is an observed byte size. An illustrative 100,000-upload calculation with a 2 MB source, a 700 KB foreground, and three 100 KB variants yields 300 GB of variants but 2 TB of sources. Those numbers are arithmetic inputs, not a benchmark; measure your own distribution before deciding what to keep.
Keep the source long enough to survive moderation, reconciliation, and a detector change. Then delete redundant intermediate files according to a written retention rule. The catch is immediate: shorter retention lowers stored bytes, but a disputed or badly framed avatar may become impossible to regenerate after its source expires.
What does the bill contain before an avatar is cropped?
Separate durable evidence from replaceable output. The accepted source is evidence of what the user submitted. The background-removed foreground is useful when removal itself is expensive or nondeterministic from the application's perspective. Square renditions are replaceable if the pipeline records the source checksum, crop rectangle, orientation decision, algorithm version, and encoder settings. Treat those fields like a small ledger: append a new transformation record rather than silently overwriting the previous decision.
This distinction changes the dominant term. If measurements show that sources account for most stored bytes, shrinking a 128-pixel rendition won't matter much. A lifecycle rule on expired sources will. If foreground assets dominate because transparent lossless files are large, retaining one normalized foreground and regenerating display sizes can be the better move. I'm not sure which term will dominate in your collection, because subject matter, transparency, dimensions, and file format change the answer; a histogram of byte size by asset role resolves that uncertainty. MDN's image-format guide is a useful starting point for format capabilities, but the decision still belongs to measurements from the actual upload population.
Measure first.
Don't count storage alone. Upload-time processing spends compute once per accepted revision and makes reads predictable. On-demand processing defers compute until a rendition is requested, can avoid work for unused uploads, and introduces cache misses and latency into a user-facing read. In a logistics application, where the same identity image may appear in dispatch, proof-of-delivery, support, and account review, consistency across those surfaces is usually more valuable than saving one early crop operation. A low-traffic internal directory with many never-viewed profiles can justify the opposite choice.
I would stop keeping superseded square renditions once a replacement is published and cache invalidation has completed. I would not discard the transformation record, because without it a later investigation cannot distinguish an incorrect detector result from an incorrect render or a stale cache. That audit trail has a privacy cost of its own: coordinates and checksums are still account-linked metadata, so access and retention must follow the organization's data policy rather than an engineer's indefinite-debugging instinct.
How should you debug an avatar crop that cuts off heads: center, smart, or manual?
Start with reproduction, not a different model. Fetch the transformation record for the affected asset, verify the source checksum, apply orientation normalization once, and render the recorded rectangle with the recorded encoder configuration. If the result differs from production, the fault lies after crop selection: stale rendition lookup, inconsistent orientation handling, or a cache key that omitted the revision are prime suspects. If it matches, inspect the rectangle against the normalized source. A center crop chooses geometry without understanding the subject. For normalized width W, height H, and square side L = min(W,H), its origin is ((W-L)/2, (H-L)/2). It works when the important region is central and has enough margin. It fails predictably when a portrait places the face high in the frame, when a tall head covering extends beyond the assumed face box, or when background removal tightened the visible foreground but the cropper still used coordinates from the pre-normalized image. That last mismatch is easy to miss: both coordinate sets look plausible, yet they refer to different pixel spaces.
Smart crop should mean a documented mapping from a detected focal region to an output rectangle, not a magic label. Record the detector or rule version and its confidence, expand the focal box by a policy margin, force the expanded box into the requested aspect ratio, clamp it to image bounds, and reject automatic publication when the required safety margin cannot fit. A detector can locate a face while the crop still cuts hair or headwear; face bounds and acceptable avatar framing are different concepts.
Manual adjustment is the authority when automation is uncertain or the user rejects its composition. Store the adjustment as normalized coordinates such as (cx, cy, scale) against a specific source revision, then derive pixels at render time. Never attach the adjustment only to a generated 256-pixel file. A new upload must create a new source revision and invalidate the old adjustment unless a reviewed migration rule says otherwise. Exactly-once publication is the mindset here — retries may execute more than once, but only one idempotency key may commit a given source revision, crop policy version, and output specification.
Use a compact triage table during review:
| Observation | Likely boundary | Next check |
|---|---|---|
| Reproduction differs from the published avatar | Rendering or delivery | Revision-aware object key and cache key |
| Reproduction matches and the focal region is high | Center-crop policy | Switch this revision to bounded focal placement |
| Smart rectangle contains the face but trims hair or headwear | Framing margin | Expand and aspect-fit the focal region before clamping |
| Automatic confidence is below policy | Automation limit | Require manual adjustment |
| Manual crop changes after a replacement upload | Revision binding | Bind coordinates to source checksum and revision |
No guesswork.
A small Go crop contract with idempotency and an audit trail
The useful implementation boundary is a pure planner followed by an idempotent publisher. The planner consumes dimensions and an optional reviewed focal point; the publisher records what it committed. Image decoding, orientation normalization, background removal, crop planning, encoding, object storage, and cache publication should have distinct trace events even if one worker performs all of them. That separation makes a 422 policy rejection, such as an out-of-range manual center, distinguishable from an infrastructure retry without claiming the image itself is corrupt.
The following planner deliberately avoids face detection. Detection belongs behind an interface whose output is tested separately. This code demonstrates the invariant that matters: a square crop remains inside the normalized image, and the same input produces the same rectangle.
package avatar
import (
"errors"
"image"
"math"
)
type Point struct {
X float64 // Normalized to [0, 1].
Y float64 // Normalized to [0, 1].
}
type Plan struct {
SourceSHA256 string
Revision int64
Policy string
Crop image.Rectangle
}
func SquareCrop(width, height int, focal *Point) (image.Rectangle, error) {
if width <= 0 || height <= 0 {
return image.Rectangle{}, errors.New("invalid normalized dimensions")
}
side := min(width, height)
cx, cy := float64(width)/2, float64(height)/2
if focal != nil {
if focal.X < 0 || focal.X > 1 || focal.Y < 0 || focal.Y > 1 {
return image.Rectangle{}, errors.New("focal point outside normalized range")
}
cx, cy = focal.X*float64(width), focal.Y*float64(height)
}
x0 := clamp(int(math.Round(cx-float64(side)/2)), 0, width-side)
y0 := clamp(int(math.Round(cy-float64(side)/2)), 0, height-side)
return image.Rect(x0, y0, x0+side, y0+side), nil
}
func clamp(value, low, high int) int {
return min(max(value, low), high)
}
The commit key can be the hash of source checksum + source revision + policy version + output dimensions + encoder configuration. A retry first looks up that key; if a completed record exists, it returns the existing asset, and if no record exists, it creates the rendition and commits the object reference plus audit fields atomically within the application's chosen consistency boundary. Do not call that delivery exactly once. The honest guarantee is idempotent effect under repeated execution, backed by a uniqueness constraint and reconciliation.
Tests should include landscape, portrait, square, focal points at all four corners, one-pixel dimensions, out-of-range manual input, orientation-normalized dimensions, and repeated commits with the same key. Property tests can assert that Dx == Dy, the rectangle is nonempty, and every edge stays within the source bounds. Golden images help reviewers judge composition, but they must supplement geometry assertions rather than replace them; a visual snapshot can hide a one-pixel drift that later changes a checksum and defeats deduplication.
Observe decisions with low-cardinality fields: policy version, manual-versus-automatic mode, output size class, result, and latency bucket. Keep source identifiers out of metric labels. Put per-asset detail in access-controlled traces or audit records, and reconcile published objects against committed records on a schedule appropriate to the service objective.
When should processing happen at upload, and when is on demand better?
Choose upload-time processing when an accepted avatar must look identical across several logistics workflows, manual approval is available, and read latency should not contain image transformation. The publish sequence should be explicit: accept a source revision, normalize it, remove the background, propose a crop, obtain manual input when policy requires it, encode declared renditions, commit the audit record, and expose the new revision. Consumers should never infer the latest rendition from a mutable filename alone.
Choose on-demand processing when requested aspect ratios vary, most uploads are never viewed, or source retention is already required for another legitimate purpose. The catch is that every request path now needs a revision-aware cache key and a deterministic fallback. It is not suitable when a reviewer must approve one canonical identity image before it reaches safety-sensitive or account-review screens; stick with upload-time publication in that case.
A hybrid is often defensible: publish one reviewed square avatar at upload, then derive noncanonical display sizes from that approved square on demand. This limits the semantic decision to one point while letting delivery adapt to device sizes. It also deliberately gives up the ability to create a later wide crop from the original once the original passes its retention deadline. When that goes wrong, support can reproduce the canonical square from its plan and checksum, but cannot recover pixels that policy required the system to delete.
The final decision rule is short. Center crop only when a dataset review establishes sufficient head margin for the intended population. Use smart placement when its confidence and safety margin meet a written threshold. Offer manual adjustment whenever identity, headwear, unusual composition, or user preference can invalidate automation. Bind every result to a source revision, and make deletion as deliberate and auditable as publication.
References
- MDN, "Image file type and format guide": https://developer.mozilla.org/en-US/docs/Web/Media/Formats/Image_types
Top comments (0)