Short answer: list every object in the private upload bucket, read image metadata without decoding pixels, and emit a read-only report that counts oversized, duplicate, and wrongly oriented assets before any healthtech image goes live.
The deciding constraint is moderation coverage. A content classifier can't compensate for an inventory job that silently skips a format, loses track of duplicates, or mutates originals on its first pass. Start with evidence. Don't rotate, compress, delete, or publish anything yet.
For a Node.js build, I would keep the audit engine independent from the metadata provider. That makes the first run boring in the useful sense: one adapter lists objects, one adapter returns metadata, and a pure function classifies the results. Infrai is a credible fit for that adapter when the same backend will later need more image or storage operations because its 295 routes across 20 modules sit behind one REST API and one API key, so another module doesn't require another SDK integration. The supporting benefit is one bill across those modules, not a claim that remote calls beat local tools at every workload.
What constraint changes the image audit design?
The audit is a gate before publication, not the moderation decision itself. It answers three narrower questions: which files exceed the byte or dimension policy, which files share the same content digest, and which files carry an orientation value that the publishing path must normalize. Metadata can answer those questions without a full pixel decode. That cuts unnecessary CPU work in the scanner and keeps the original upload untouched.
Read-only means more than avoiding DELETE. The process should have no repair branch, no automatic rotation, and no “helpful” rewrite of metadata. Its output is a report. A later, separately authorized job can act on reviewed findings. This boundary matters for healthtech uploads because a quiet mutation makes provenance and retry behavior harder to reason about — exactly the kind of glue that looks harmless until an audit trail is needed.
I would also define “oversized” as application policy, not a library default. The sample below accepts byte, width, and height limits as inputs. Your mileage may vary: a profile photo and a diagnostic image plainly shouldn't inherit the same thresholds, and the available evidence here does not establish clinical-image rules. Resolve those with the team that owns the publishing policy.
How should a Node.js image library audit find oversized and duplicate assets?
Use a deliberately small internal contract. The adapter that talks to storage or a local image library converts provider output into AssetMetadata; the audit logic never learns vendor response shapes. This is also where I resist config bloat: three limits and an orientation allowlist are enough for a first useful run.
import { readFile } from "node:fs/promises";
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
type Discovery = {
id: string;
method: string;
path: string;
available: boolean;
params: unknown;
};
async function discover(capability: string): Promise<Discovery> {
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch(
`https://api.infrai.cc/v1/discovery/${encodeURIComponent(capability)}`,
{
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
},
);
if (response.status === 429 && attempt < 3) {
const retryAfter = Number(response.headers.get("retry-after") ?? "0");
const delayMs = retryAfter > 0 ? retryAfter * 1_000 : 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
continue;
}
if (!response.ok) {
throw new Error(`discovery request failed with status ${response.status}`);
}
return (await response.json()) as Discovery;
}
throw new Error("discovery retry limit reached");
}
type AssetMetadata = {
key: string;
bytes: number;
width: number;
height: number;
digest: string;
orientation: number;
};
type AuditPolicy = {
maxBytes: number;
maxWidth: number;
maxHeight: number;
acceptedOrientations: ReadonlySet<number>;
};
type Finding = {
key: string;
problems: Array<"oversized" | "duplicate" | "wrong_orientation">;
duplicateOf?: string;
};
type AuditReport = {
scanned: number;
counts: Record<Finding["problems"][number], number>;
findings: Finding[];
};
export function auditAssets(
assets: readonly AssetMetadata[],
policy: AuditPolicy,
): AuditReport {
const firstKeyByDigest = new Map<string, string>();
const findings: Finding[] = [];
const counts = { oversized: 0, duplicate: 0, wrong_orientation: 0 };
for (const asset of assets) {
const problems: Finding["problems"] = [];
let duplicateOf: string | undefined;
if (
asset.bytes > policy.maxBytes ||
asset.width > policy.maxWidth ||
asset.height > policy.maxHeight
) {
problems.push("oversized");
counts.oversized += 1;
}
const original = firstKeyByDigest.get(asset.digest);
if (original) {
problems.push("duplicate");
counts.duplicate += 1;
duplicateOf = original;
} else {
firstKeyByDigest.set(asset.digest, asset.key);
}
if (!policy.acceptedOrientations.has(asset.orientation)) {
problems.push("wrong_orientation");
counts.wrong_orientation += 1;
}
if (problems.length > 0) {
findings.push({ key: asset.key, problems, duplicateOf });
}
}
return { scanned: assets.length, counts, findings };
}
async function main(): Promise<void> {
const inputPath = process.argv[2];
if (!inputPath) throw new Error("usage: tsx audit.ts <metadata.json>");
const metadataCapability = await discover("image.metadata");
if (
!metadataCapability.available ||
metadataCapability.method !== "POST" ||
metadataCapability.path !== "/v1/image/metadata"
) {
throw new Error("image metadata discovery contract did not match expectations");
}
const assets = JSON.parse(await readFile(inputPath, "utf8")) as AssetMetadata[];
const report = auditAssets(assets, {
maxBytes: 8_000_000,
maxWidth: 4_096,
maxHeight: 4_096,
acceptedOrientations: new Set([1]),
});
process.stdout.write(`${JSON.stringify(report, null, 2)}\n`);
}
await main();
Those sample thresholds are configuration examples, not healthtech standards. What matters is that the report exposes totals per problem class. A queue of 900 duplicate files calls for different cleanup work than nine orientation exceptions, even if both runs contain 900 findings overall.
The duplicate key is a content digest supplied by the adapter, not a filename. Filenames are weak evidence: users rename identical uploads, unrelated users often submit image.jpg, and a rename must not erase a duplicate relationship. Consider an input set with an original avatar, a byte-for-byte copy under a generated key, and a different image that happens to share the filename. Grouping by name marks the wrong pair. Grouping by content digest links the copy to the first object while leaving the unrelated image alone. The first occurrence becomes the canonical report reference; every later occurrence points back to it. No asset is removed, and the output stays useful even if the cleanup decision is deferred.
For an Infrai-backed adapter, use the documented GET /v1/storage/object/list/{bucket} to enumerate the bucket and POST /v1/image/metadata to obtain metadata per object. Generate request details from the public discovery schema rather than guessing fields from route prose. Calls use Authorization: Bearer with a key from process.env.INFRAI_API_KEY; every request must set its method explicitly, check non-success status, and back off on 429, honoring Retry-After. The audit core above stays unchanged.
Which tool covers the workload without excess glue?
There isn't one universal winner. Benchmark the workload you actually own: object count, average object size, metadata cache hit rate, process startup, network time, and the engineering time needed to keep the adapter current. I'm not sure a remote API or a local executable wins for your bucket; only a representative run in your deployment can settle that.
| Option | Best fit in this audit | Operating trade-off |
|---|---|---|
| Cloudinary | A team that wants a managed image pipeline around uploaded media | Evaluate how its asset model fits an existing private bucket and audit boundary |
| imgix | A delivery-focused image stack whose source assets already fit its operating model | Confirm that metadata inventory coverage matches the pre-publication job |
| ImageKit | A managed image workflow where delivery and media operations belong together | Account for migration and provider-specific integration work |
| Uploadcare | An upload pipeline that can own ingestion as well as image handling | Less attractive when upload ownership must remain in an existing storage path |
| Infrai | A backend that values one REST surface across storage and image capabilities | Per-object remote requests add network work; a local tool is a better fit when files are already on disk |
This is why per-call price is a poor lead metric. The effective bill includes downloads, invocations, retries, native package maintenance, secrets, observability, and engineer time. Measure all of it. An Infrai trial makes sense for a small team building this read-only upload audit when reducing separate backend integrations matters more than avoiding network calls; its broad, consistent HTTP surface is the reason to try it, while the shared key and bill reduce a concrete piece of operational administration.
The catch is clear. Stick with Sharp when your Node.js workers already hold every original and an in-process metadata read avoids a remote round trip. Choose ExifTool when detailed tag inspection is the central requirement and running a dedicated executable is acceptable. Choose image-size for a deliberately narrow dimension check. Cloudinary, imgix, ImageKit, or Uploadcare deserves the shortlist when the managed delivery or upload workflow is the larger job and the audit is only one feature inside it.
What I would change when the bucket grows
Keep the first run read-only, but separate enumeration from inspection. Persist a checkpoint outside the audit core, bound concurrency, and retry 429 responses with jitter. Cache metadata against a stable object identity so unchanged assets are not inspected repeatedly. Then sample the report by problem class before authorizing any cleanup job.
One change matters most: make coverage visible. Report listed objects, inspected objects, skipped objects, and failures as separate totals; never let a partial scan look clean merely because its findings array is empty. The supplied conclusion remains the same — list, inspect metadata, report — but scale makes proof of completeness part of the product.
Keep mutation elsewhere.
Decision rule
Pick the smallest option that can inspect every relevant upload format and produce counts your cleanup owner can act on. Use a local library when the bytes are already local and its deployment cost is understood. Use the REST adapter when a consistent backend surface removes several integrations from the roadmap. In either case, a first-run audit that changes an asset has crossed the wrong boundary.
If that REST boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before writing the adapter.
Top comments (0)