TL;DR: Run a read-only inventory before changing a logistics image pipeline. Stream each source file once, record byte size and dimensions, and compute a content digest while reading. Then separate exact duplicates from merely oversized inputs and estimate the storage plus cache exposure of each class. The first useful change is usually a retention rule backed by evidence, not another crop preset.
For smart-cropping shipment photos into several aspect ratios, the bill is made of source bytes, derived-object bytes, cache copies, request traffic, and repeated transformation work. Count those terms independently. A duplicate detector can shrink the first two; a size threshold alone cannot. An oversized unique photograph may be the only evidence available when a crop must be regenerated.
Start read-only. No deletes, no rewrites, and no metadata mutation belong in the first run. The output should be a manifest that another process can review, version, and compare.
Bytes first.
How should a Node.js image audit find oversized sources?
Use bytes, not file counts, as the primary unit. Ten tiny icons and one loading-dock photograph are eleven objects, but they do not exert equal pressure on object storage or caches. For an object i, a useful accounting model is:
stored_i = source_i + sum(derived_i) + cache_i
Across the collection, group that value by operational status: active shipment, delivered but inside the retention window, expired, legal hold, and unknown. Unknown must be a real state. Treating missing retention metadata as expired is an unsafe shortcut.
Consider a hypothetical manifest row with a 12 MB source and four 1.5 MB crops. Its pre-cache footprint is 18 MB. If the same source bytes appear under three object keys and every key has its own derivatives, the upper-bound pre-cache footprint is 54 MB. This is arithmetic for prioritization, not a benchmark or a promise about any storage system. Measure cache residency separately because cache policy and request locality decide how much of that upper bound exists at one time.
That distinction changes the order of work. Flagging the 12 MB source as oversized identifies a review candidate. Proving that three sources have identical bytes identifies reclaimable copies, subject to ownership and retention rules. Those are different findings and should never share one Boolean column.
Why can an oversized file still be worth keeping?
A smart crop is a decision, not an archival master. A new aspect ratio, a corrected focal region, or a changed orientation policy may require pixels outside yesterday's derivative. Deleting every large source after the first crop saves source storage but makes later regeneration lossy or impossible.
The practical decision is a retention matrix. Keep an original while a shipment is active and for the documented evidence window afterward. Keep a digest and lineage record longer if policy permits, because a digest can establish that two observed byte streams matched without retaining either stream. Place holds above ordinary expiry. When policy is absent, report the object instead of guessing. Compliance failures often begin in that quiet else branch.
Formats complicate the word "oversized." The MDN image format guide documents that web image formats have different capabilities and browser support characteristics. A byte threshold should therefore be a triage rule, not a quality verdict. Record the detected format, dimensions, byte count, and intended use, then let a policy layer decide which combinations deserve review.
Keep the categories explicit:
-
exact_duplicate: the full content digest matches another complete object. -
oversized_candidate: bytes or pixel area exceed a policy threshold. -
both: a large object is also an exact duplicate. -
unknown: the file could not be read completely or its metadata was not established.
A perceptual match is not an exact duplicate. Two recompressed shipment photos may look alike while preserving different detail, metadata, or evidence value. Put perceptual similarity in a later, human-reviewed workflow; never feed it directly into deletion.
A read-only scanner contract
The audit process needs less authority than the crop service. Give it read access to the selected namespace and write access only to a separate report destination. Its contract is narrow: enumerate an immutable snapshot or a bounded prefix, stream bytes, extract metadata without decoding more than necessary, calculate a cryptographic digest, and append a result row. It does not rename or delete assets.
Although this architecture fits a Node.js worker, the core interface should stay language-neutral. The following Python reference expresses the scanner contract because the same fields and invariants can be implemented behind a Node.js queue consumer without coupling the manifest to a library:
from dataclasses import asdict, dataclass
from hashlib import sha256
from pathlib import Path
import json
@dataclass(frozen=True)
class AuditRow:
object_key: str
bytes_read: int
sha256: str | None
status: str
error: str | None
def inspect_file(path: Path, chunk_size: int = 1024 * 1024) -> AuditRow:
digest = sha256()
total = 0
try:
with path.open("rb") as source:
while chunk := source.read(chunk_size):
digest.update(chunk)
total += len(chunk)
return AuditRow(str(path), total, digest.hexdigest(), "complete", None)
except OSError as exc:
return AuditRow(str(path), total, None, "incomplete", type(exc).__name__)
def write_json_line(row: AuditRow) -> None:
print(json.dumps(asdict(row), separators=(",", ":"), sort_keys=True))
The complete state matters more than it looks. A digest from a partial read must not enter a duplicate group. Do not preserve the partial digest as if it represented the object; keep the byte count and error class for diagnosis, then retry under a bounded policy. This is the media equivalent of refusing to mark an OTP delivered when only the enqueue step succeeded.
Partial means unknown.
Dimension extraction belongs beside this loop, but behind a parser that recognizes the detected format. Do not infer format solely from a filename extension. The report can include nullable width, height, and format fields; null is safer than a fabricated zero.
For very large inventories, partition enumeration by a stable key range and checkpoint only after report rows are durably written. Rerunning a partition is acceptable because the operation is read-only and rows can be keyed by snapshot identifier plus object key. Concurrency should be capped against storage read capacity and report-writer throughput. More workers can raise cache misses and throttling without improving the decision quality.
From digest groups to a cost decision
After the scan, group only complete rows by digest. Before labeling bytes reclaimable, verify that the group members belong to the same content boundary and that their retention states permit consolidation. Cross-tenant equality is a security-sensitive observation; a report should not reveal one tenant's object keys to another.
A compact review table keeps the decision legible:
| Finding | Evidence | Storage action | Failure cost |
|---|---|---|---|
| Large, unique source | Bytes, dimensions, use | Retain or move by policy | Future crops may lose source pixels |
| Exact duplicate, same owner | Complete digest match | Select one canonical source after review | Broken references if lineage is incomplete |
| Similar-looking images | Perceptual signal only | No automatic removal | Distinct evidence may be discarded |
| Incomplete read | Error and partial byte count | Retry; exclude from groups | Audit coverage remains unknown |
Estimate a candidate group's gross removable source bytes as sum(member_bytes) - canonical_bytes. Then subtract nothing on faith. Derived objects may already be shared, may have independent retention, or may be needed by live references. Cache savings are even less certain: report them as a separately measured term rather than multiplying source savings by an assumed cache factor.
This is where storage and cache cost stay honest. Report gross candidate bytes, policy-eligible bytes, and confirmed removable bytes as three columns. A finance-friendly single total hides the engineering uncertainty that still needs resolution.
Deployment should preserve the same restraint. First run against a bounded snapshot, compare enumerated objects with completed and incomplete rows, and publish only the manifest. A second process can enrich ownership and retention data. Any later mutation job needs its own authorization, dry-run diff, reference validation, approval record, and rollback plan; it is outside the audit worker's contract.
What should we deliberately stop keeping?
Stop retaining redundant byte-identical sources after ownership, references, holds, and retention windows have been resolved. Stop retaining derivatives that have no live reference and can be reproduced from a retained source under a documented recipe. Do not stop keeping the manifest, lineage, policy decision, and audit timestamp.
The trade-off is recovery time and, sometimes, recoverability itself. If a deleted derivative can be regenerated, an incident consumes compute and delays image delivery while caches refill. If the original was also deleted, a new crop may be impossible. That is why the inventory should expose the dominant term before anyone optimizes it: a storage reduction is useful only when its failure cost is named.
The first run ends with evidence, not deletion. That boundary keeps the Node.js audit worker small, makes retries boring, and gives operations, compliance, and product owners the same set of facts when they decide what may disappear.
Further reading
- MDN, "Image file type and format guide": https://developer.mozilla.org/en-US/docs/Web/Media/Formats/Image_types
Top comments (0)