TL;DR: Put a deterministic content fingerprint in front of every expensive image rendition, but keep moderation decisions versioned separately. For event photos feeding short promo videos, the skip condition should be “this exact asset, rendition recipe, and moderation policy already reached an accepted terminal state,” not merely “this filename exists.” Log that decision as data. A cost spike then becomes a query over repeated keys instead of a tour through worker logs.
This is the smallest design I trust because it attacks duplicate work without weakening moderation coverage. It also keeps configuration small: one key function, one state record, one conditional claim. No sprawling workflow file.
What constraint changes the skip check?
The obvious check is usually wrong. An event-photo pipeline sees IMG_0042.jpg twice and skips the second copy, or sees two different object paths and renders both. Neither filename nor path answers the useful question: has this image already passed the required checks and produced this requested output?
The moderation policy changes the key. A photo accepted under one policy is not proof that it was checked under a later policy. Likewise, an approved original does not imply that every crop is suitable for a promo frame; a crop can change what is visible. So I would separate immutable asset identity from policy and rendition identity, then require all three at the reuse boundary.
That trade-off costs a little metadata. Good. Paying for explicit state is easier to debug than paying repeatedly for opaque processing.
Image format belongs in the intake record too. Event uploads may arrive in different file types, and format support varies across image ecosystems. MDN's image format guide is a useful baseline for recognizing common types and their characteristics. Still, a file extension is descriptive metadata, not a deduplication key. Two names can carry identical bytes; identical names can carry different bytes.
The smallest implementation I would ship
I start with raw-byte identity because it is deterministic and cheap to reason about. The rendition recipe includes every input that can change the output. The moderation policy version is explicit. If any one changes, reuse is denied.
import { createHash } from "node:crypto";
type RenditionRecipe = {
width: number;
height: number;
fit: "cover" | "contain";
format: "jpeg" | "png" | "webp";
};
type WorkIdentity = {
assetHash: string;
recipeHash: string;
moderationPolicy: string;
};
function sha256(value: Uint8Array | string): string {
return createHash("sha256").update(value).digest("hex");
}
function stableRecipe(recipe: RenditionRecipe): string {
return JSON.stringify({
fit: recipe.fit,
format: recipe.format,
height: recipe.height,
width: recipe.width,
});
}
export function identifyWork(
bytes: Uint8Array,
recipe: RenditionRecipe,
moderationPolicy: string,
): WorkIdentity {
return {
assetHash: sha256(bytes),
recipeHash: sha256(stableRecipe(recipe)),
moderationPolicy,
};
}
Sorting the recipe fields here is deliberate. Relying on whatever object shape happened to arrive at the worker makes the cache boundary depend on calling code. I want the function to own its serialization contract.
The work record needs a narrow state machine. accepted is reusable. rejected is terminal but not reusable for publishing. processing is neither success nor permission to enqueue another copy. A retry can claim stale work according to an explicit lease policy, but a second live worker must not win the same claim.
type WorkState = "processing" | "accepted" | "rejected" | "failed";
type WorkRecord = WorkIdentity & {
state: WorkState;
outputKey?: string;
claimedAt: string;
finishedAt?: string;
};
type ClaimResult =
| { action: "reuse"; outputKey: string }
| { action: "wait" }
| { action: "process" };
interface WorkStore {
find(identity: WorkIdentity): Promise<WorkRecord | undefined>;
claim(identity: WorkIdentity, claimedAt: string): Promise<boolean>;
}
export async function claimOrSkip(
store: WorkStore,
identity: WorkIdentity,
now: Date,
): Promise<ClaimResult> {
const existing = await store.find(identity);
if (existing?.state === "accepted" && existing.outputKey) {
return { action: "reuse", outputKey: existing.outputKey };
}
if (existing?.state === "processing") {
return { action: "wait" };
}
const claimed = await store.claim(identity, now.toISOString());
return claimed ? { action: "process" } : { action: "wait" };
}
The critical operation is claim, not find. It must allow only one winner for the composite identity. A read followed by an unconditional insert leaves a race: two workers can both observe nothing and both start processing. The store implementation should enforce uniqueness at the write boundary.
Keep the moderation call and rendition call outside the claim transaction. Long-running media work should not hold a database transaction open. The record represents ownership; completion updates the state afterward.
How should I debug duplicate image processing after costs spiked?
I would not begin by tuning codecs or shopping for cheaper processing. First, count decisions at the boundary where work is claimed. Each event should include the asset hash prefix, recipe hash prefix, policy version, prior state, action, and a correlation ID for the promo-video request. Do not log image bytes or user-facing filenames when a stable opaque key will do.
Start there.
Then group processing starts by the full composite identity. More than one start for the same identity is direct evidence that the claim boundary failed, the key changed unexpectedly, or a retry policy reclaimed live work. Group reuse decisions separately. This makes the diagnosis mechanical.
My first hypothesis would be a missing skip check. I would try to disprove it before changing the pipeline: pick one promo request with repeated starts, trace each start back to its identity fields, and compare those fields byte for byte. If the keys match but multiple workers report process, the atomic claim is suspect. If the asset hashes differ, inspect where bytes were read and transformed. If only recipe hashes differ, diff the serialized recipes rather than eyeballing application config. If only policy versions differ, the extra work may be intentional moderation coverage. This path matters because all four cases can look like “the same photo ran twice” on a dashboard, yet they demand different fixes. Collapsing them into one duplicate counter hides the boundary that actually changed.
The evidence decides.
I benchmark the boundary with a fixed corpus, but I would not publish made-up throughput numbers. The useful local test runs the same byte buffer concurrently through many callers and asserts exactly one process result. A second test changes only the moderation policy and expects new work. A third changes only the crop recipe. A fourth sends the same filename with different bytes, because filenames lie.
import assert from "node:assert/strict";
async function assertSingleWinner(
run: () => Promise<ClaimResult>,
callers: number,
): Promise<void> {
const results = await Promise.all(
Array.from({ length: callers }, () => run()),
);
const winners = results.filter((result) => result.action === "process");
assert.equal(winners.length, 1);
}
Ten concurrent callers are enough for a regression test fixture; they are not a capacity claim. For performance work, increase concurrency against the actual store and report the machine, corpus, image sizes, store latency, and configuration alongside results. A naked requests-per-second number tells me almost nothing.
There is another failure mode worth checking: metadata normalization. If one uploader rotates pixels before hashing while another preserves the original bytes, visually similar photos will get different raw hashes. That does not make the key incorrect. It means the team must decide whether the reusable unit is the uploaded file or a canonical decoded image. Canonicalization can merge more work, but it adds processing before the skip check and expands the code that must remain stable.
I would begin with uploaded bytes. Fewer moving parts win until observed duplicates justify more.
What I would change at scale
At higher volume, I would split the ledger from the rendered object store and retain compact decision records longer than transient worker logs. The ledger answers why work ran. The object store holds the result. Mixing those lifecycles makes an expired log look like missing evidence.
I would also add two explicit versions: one for the key algorithm and one for the moderation policy. Algorithm versioning matters when serialization or canonicalization changes. Without it, an apparently identical hash namespace can quietly mean two different things.
Retries need leases with observable expiry and a bounded attempt count. A crashed worker should not block an event gallery forever, while a slow worker should not be duplicated just because a queue message became visible again. Those time limits depend on measured processing duration, so there is no honest universal number to paste here.
For promo generation, keep provenance from the final clip back to every accepted rendition and its moderation decision. If a policy changes, the system can identify affected inputs and regenerate deliberately. Blindly clearing every cache is easy. It is also an expensive substitute for data modeling.
The trade-offs I would accept
Raw-byte hashing misses visually equivalent files that were re-encoded. Perceptual matching may catch them, but similarity thresholds can merge distinct event photos and create a moderation ambiguity. I would use perceptual signals for investigation before allowing them to authorize reuse.
Policy-aware keys produce more work after a policy update. That is correct behavior, not cache inefficiency. Moderation coverage is part of the output contract.
Finally, do not turn the ledger into a configuration platform. Store the inputs that actually affect output, the decision, timestamps, and correlation data. Every optional flag multiplies key combinations and test cases. The best developer experience here is boring: deterministic keys, one atomic claim, explicit terminal states, and logs that answer a cost question in a single grouping operation.
Further reading
- References: MDN, Image file type and format guide
Top comments (0)