DEV Community

UriahHawkins5489
UriahHawkins5489

Posted on

Node.js Image Audit for Oversized and Duplicate Assets — A Read-Only First Run

Node.js Image Audit for Oversized and Duplicate Assets — A Read-Only First Run

Bandwidth is the constraint that changes this audit. A file can be valid, popular, and still too large for the delivery path; a byte-for-byte duplicate can be harmless in one product area and the wrong thing to remove in another. Start with a read-only Node.js inventory that identifies oversized and exact-duplicate images, then measure the report before any mutating job is allowed near storage.

Short answer: inspect metadata first, hash only the candidates that need proof, and write a review report with stable asset identifiers. The first run should change nothing.

How should a Node.js image audit find oversized and duplicate assets?

Define “oversized” before opening a directory. I use two independent thresholds: a byte budget for transfer cost and, when dimensions are available, a pixel budget for transformation cost. A 1.5 MB original may be acceptable for a download endpoint but excessive for a 320-pixel card. Mixing those meanings creates noisy tickets.

The cheap pass reads path, byte count, modification time, and a storage identifier. Extensions are only hints. A decoder or content-signature check must confirm that a file is an image before the report makes an image-specific claim. MDN's format guide explains why format choice changes compression and browser support; it does not make a filename trustworthy.

Exact duplicates need a digest of the bytes. Start with size grouping, then hash only files that share a size. This saves reads when a library has thousands of unrelated assets, while preserving the result for identical bytes. A digest will not identify a re-encoded JPEG, a crop, or a file whose metadata changed. Those are perceptual-similarity questions and belong in a separate review queue.

Here is a deliberately boring TypeScript scanner. It walks regular files, groups by size, hashes each member of a repeated-size group, and emits JSON. There is no delete, rename, upload, or rewrite operation hidden in the example.

import { createHash } from "node:crypto";
import { createReadStream } from "node:fs";
import { readdir, stat } from "node:fs/promises";
import { extname, join, relative } from "node:path";

type FileRecord = {
  path: string;
  bytes: number;
  modified: string;
  digest?: string;
};

const IMAGE_EXTENSIONS = new Set([".jpg", ".jpeg", ".png", ".webp", ".gif", ".avif"]);

async function walk(root: string, current = root): Promise<string[]> {
  const entries = await readdir(current, { withFileTypes: true });
  const files: string[] = [];
  for (const entry of entries) {
    const fullPath = join(current, entry.name);
    if (entry.isDirectory()) files.push(...await walk(root, fullPath));
    else if (entry.isFile()) files.push(fullPath);
  }
  return files;
}

function sha256(path: string): Promise<string> {
  return new Promise((resolve, reject) => {
    const hash = createHash("sha256");
    const input = createReadStream(path);
    input.on("data", chunk => hash.update(chunk));
    input.on("error", reject);
    input.on("end", () => resolve(hash.digest("hex")));
  });
}

async function audit(root: string, maxBytes: number) {
  const records: FileRecord[] = [];
  const errors: Array<{ path: string; message: string }> = [];

  for (const path of await walk(root)) {
    if (!IMAGE_EXTENSIONS.has(extname(path).toLowerCase())) continue;
    try {
      const metadata = await stat(path);
      if (!metadata.isFile()) continue;
      records.push({
        path: relative(root, path),
        bytes: metadata.size,
        modified: metadata.mtime.toISOString()
      });
    } catch (error) {
      errors.push({ path: relative(root, path), message: String(error) });
    }
  }

  const bySize = new Map<number, FileRecord[]>();
  for (const record of records) {
    const group = bySize.get(record.bytes) ?? [];
    group.push(record);
    bySize.set(record.bytes, group);
  }

  for (const group of bySize.values()) {
    if (group.length < 2) continue;
    for (const record of group) {
      try {
        record.digest = await sha256(join(root, record.path));
      } catch (error) {
        errors.push({ path: record.path, message: String(error) });
      }
    }
  }

  const digestGroups = new Map<string, FileRecord[]>();
  for (const record of records) {
    if (!record.digest) continue;
    const group = digestGroups.get(record.digest) ?? [];
    group.push(record);
    digestGroups.set(record.digest, group);
  }

  return {
    root,
    maxBytes,
    scanned: records.length,
    oversized: records.filter(record => record.bytes > maxBytes),
    duplicates: [...digestGroups.values()].filter(group => group.length > 1),
    errors
  };
}

const root = process.argv[2] ?? "./public";
const maxBytes = Number(process.argv[3] ?? 500_000);
console.log(JSON.stringify(await audit(root, maxBytes), null, 2));
Enter fullscreen mode Exit fullscreen mode

The example uses extensions to keep the sample readable. In a production bucket, put content-signature validation in front of the same record shape, and record “unreadable” separately from “not an image.” A failed decode is data-quality evidence, not permission to discard a file.

The first report should also make a race visible. Imagine uploads/team-a/banner.jpg is 480,000 bytes when stat runs, then an editor replaces it while the stream is being read. The digest now describes a different byte sequence than the metadata row, even though both operations completed normally. Capture metadata again after hashing and mark the record changed when size or modification time differs; do not silently present that digest as stable. On object storage, use the provider's immutable version or generation identifier when one exists, and keep it beside the relative path. If a read fails halfway through, retain the error, omit the digest, and continue scanning other files so one bad object does not erase the scope of the report. The review UI can then show three distinct states: confirmed exact duplicate, candidate needing another read, and unreadable item. That small distinction prevents a later cleanup script from treating missing evidence as permission.

No mutation.

What must a read-only first-run report prove?

Another engineer should be able to reproduce the scope without guessing. Include the root or bucket prefix, threshold, scan timestamp, number of files considered, and every permission or read error. Preserve the relative path and immutable source identifier; a digest with no location cannot drive a review.

I keep a decision ledger beside the JSON report. It records the reviewer, the approved action, and the identifier that a later job is allowed to use. That distinction matters when a CDN URL, database row, and object-store key all point at the same bytes. The URL is not automatically the deletion key.

Do not pick one file from each duplicate group and remove the rest on run one. Identical bytes do not prove identical business meaning. A hero image and an email attachment can be the same object with different retention rules. The report should expose the group and its owners; it should not invent the owner.

For oversized images, dimensions add useful context when a trusted decoder is available. Bytes answer the bandwidth question. Dimensions explain why a transform or thumbnail may be expensive. Keep both fields independent so a product team can change its delivery policy without rescanning every byte.

One short rule: prove the scope.

Where does this audit stop being a good fit?

This approach is conservative by design. A full-byte SHA-256 catches exact duplicates, not near-duplicates, animated-frame differences, or images with distinct color profiles. Add perceptual hashing only when reviewers can inspect false positives and the product actually benefits from merging similar images.

It is not suitable when a multi-terabyte, hot store must be enumerated by one process during peak traffic. Partition by prefix, checkpoint progress, and rate-limit reads. Keep the output schema stable so partitions can be merged without reinterpreting older reports. Your mileage may vary with the storage backend; verify its consistency guarantee before trusting a digest produced after a metadata read.

The simple size grouping also has a boundary. Two different files can share a size, so both are still hashed. That is extra I/O, but it is bounded by repeated-size groups and keeps the duplicate claim exact. If the read budget is tighter than the review budget, sample first and publish the sampling rule with the report.

How should teams turn the report into a bandwidth decision?

Before changing an asset, sample the edges: files just above the byte threshold, large-dimension files below it, several duplicate groups, and a handful of skipped or unreadable entries. Measure wall-clock duration, bytes read, read-error rate, and the fraction of candidates a reviewer accepts. Those numbers tell you whether the next run should optimize enumeration, hashing, decoding, or review throughput.

Only after that evidence exists should a separate mutating job consume approved identifiers. Give that job a dry-run mode, an idempotency key, and a bounded batch size. Run the read-only audit again afterward; the before-and-after reports are the audit trail.

I initially treated the byte threshold as the whole answer. It wasn't. A 500,000-byte limit (the sample default above) says nothing about a 6000 by 4000 image that will be resized on every request, and it says nothing about a small duplicate attached to two different records. The useful decision combines transfer budget, dimensions, ownership, and observed review cost.

That's the trade-off. Hashing spends I/O now; deleting on weak evidence spends trust later. For a small team shipping image features, the first-run report is usually the safer place to spend attention. I'm not sure a single universal threshold exists, and that uncertainty is a reason to publish the measurement with the policy.

Sources

Top comments (0)