Treat every video submission as a durable request record before starting OCR, and let retries retrieve that record instead of launching fresh work. The deciding constraint is not the HTTP timeout; it is whether the server can prove that two attempts represent the same user intent after either side loses the response. This matters in a developer tool that extracts text from uploaded video frames. A user can click once, the upload can finish, and the response can disappear on the way back. The client then retries. If each arrival creates a job, the same frames are decoded, stored, cached, and sent through OCR twice. Storage and cache cost grow even though the user asked for one result. The practical design is small: assign an intent key at the client, store one receipt under that key in the same transaction that accepts the work, and give each downstream step its own retry scope. Don't use a process-local set or a short cache entry as the authority. They can be useful accelerators, but neither is the request ledger.
How should retry boundaries use stable request records for duplicate video jobs?
A retry boundary should surround an operation whose outcome can be named and looked up. At the intake boundary, that outcome is a request record. Its key stays constant while the client is retrying the same submit action; its payload digest prevents accidental reuse for different content. The response can then say, in effect, "this intent was accepted and here is its current state," regardless of which attempt receives the answer.
Keep three identities separate:
-
request_keyidentifies the user's submit intent. -
input_digestidentifies the exact upload manifest or immutable object reference. -
job_ididentifies one internal execution created for the accepted intent.
Collapsing those into one hash looks attractive, but it changes the product semantics. Two deliberate requests for the same video may need separate settings, retention policies, or output formats. Conversely, a retry can carry the same intent even when transport details differ. The request key answers "did we already accept this action?" while the digest answers "is this retry still describing the same input?"
Use the same request_key only until the original action reaches a terminal receipt. If a user changes the OCR language, frame interval, crop, or output schema, that is new intent and needs a new key. This line is the first retry boundary.
Put the receipt before the expensive work
The simplest failed shape is also the most common: accept an upload, enqueue OCR, and only then write enough state to answer the caller. There is a gap between those operations. A retry landing in that gap cannot distinguish "nothing happened" from "work exists but its acknowledgement was lost." Starting another job feels defensive, yet it turns uncertainty into duplicate cost.
Put a uniqueness constraint in durable storage and make receipt creation the admission decision. The focused example below uses Python's standard library and SQLite so it can run locally. It does not pretend that one database fits every deployment; the important part is the transaction, the unique request key, and the payload mismatch check.
import hashlib
import json
import sqlite3
import uuid
def canonical_digest(video_uri: str, options: dict) -> str:
payload = json.dumps(
{"video_uri": video_uri, "options": options},
sort_keys=True,
separators=(",", ":"),
).encode("utf-8")
return hashlib.sha256(payload).hexdigest()
def open_db(path: str = "requests.db") -> sqlite3.Connection:
db = sqlite3.connect(path, isolation_level=None)
db.row_factory = sqlite3.Row
db.execute("PRAGMA journal_mode = WAL")
db.execute(
"""
CREATE TABLE IF NOT EXISTS video_requests (
request_key TEXT PRIMARY KEY,
input_digest TEXT NOT NULL,
job_id TEXT NOT NULL UNIQUE,
state TEXT NOT NULL CHECK (state IN ('accepted', 'running', 'done')),
result_json TEXT
)
"""
)
return db
def accept_video(
db: sqlite3.Connection,
request_key: str,
video_uri: str,
options: dict,
) -> tuple[dict, bool]:
digest = canonical_digest(video_uri, options)
proposed_job_id = str(uuid.uuid4())
db.execute("BEGIN IMMEDIATE")
try:
cursor = db.execute(
"""
INSERT OR IGNORE INTO video_requests
(request_key, input_digest, job_id, state)
VALUES (?, ?, ?, 'accepted')
""",
(request_key, digest, proposed_job_id),
)
created = cursor.rowcount == 1
row = db.execute(
"SELECT * FROM video_requests WHERE request_key = ?",
(request_key,),
).fetchone()
if row["input_digest"] != digest:
raise ValueError("request key was reused with a different payload")
db.execute("COMMIT")
return dict(row), created
except Exception:
db.execute("ROLLBACK")
raise
Call accept_video twice with the key submit_7f31, the same immutable video URI, and the same options. The first call returns created=True; the second returns created=False with the original job_id. Reuse submit_7f31 with a different crop and the function rejects it. Those are deterministic assertions for an eval harness, not timing-dependent guesses.
There is a deliberate limitation here. SQLite serializes the admission transaction and is a good executable model, but a service with many writers may need a database whose concurrency and availability characteristics match its deployment. Preserve the invariant when changing storage: one committed request key maps to one digest and one job. A cache can mirror the row for fast reads after that invariant exists.
Short gap, big consequence.
Give workers a narrower retry scope
Admission deduplication prevents two jobs from being born for one submit intent. It does not make every OCR step magically safe to repeat. Frame extraction, object writes, cache population, and result assembly each need an output identity derived from the accepted job and the step. For example, a frame artifact can be addressed by job_id, an extraction-plan version, and a frame index. A worker retry then targets that artifact rather than creating a second anonymous copy.
The state transition should be conditional as well. A worker claims an accepted job only if it is still claimable, records a lease or attempt token, and publishes the final result only if it still owns that attempt. The exact lease duration depends on observed frame extraction and OCR latency. I'm not sure which duration fits your traffic; a latency distribution from traces, including long videos, is what resolves it. Guessing a round number in code merely moves the duplicate boundary.
This is also where media format handling belongs. Validate the input container, codecs, and browser or toolchain compatibility before expensive extraction, and record the normalized inspection result with the request. Media support varies by container and codec combination, so an extension alone is not a sufficient execution plan. The MDN media formats guide is a useful starting map for those distinctions, while the formats accepted by your own decoder still need direct tests.
Keep retries local. An OCR timeout for frame 120 should not resubmit the whole video, and a result-write retry should not decode frame 120 again. Stable intermediate names make that separation inspectable. They also keep prompt or model calls inside the smallest failed unit, which matters when an OCR pipeline includes a later language-model cleanup pass.
Test the lost-response path, not only the happy path
A normal unit test that submits once proves almost nothing about duplicate work. The useful experiment forces ambiguity: commit the request receipt, discard the first response, and send the same request again. Then assert one request row, one job ID, and one set of named artifacts. Run the same test concurrently so the uniqueness constraint, rather than lucky ordering, decides the result.
The evaluation matrix should vary intent and payload independently. Same key plus same payload must return the existing receipt. Same key plus changed payload must be rejected. A new key plus the same payload must create new intent when the product allows deliberate reprocessing. Two keys arriving at nearly the same instant must remain two requests; deduplication should not merge separate users merely because their video bytes match.
Measure what the architecture claims to control: accepted requests, duplicate intake attempts, jobs created, worker attempts by step, bytes written by artifact class, cache writes and evictions, and OCR invocations. The most revealing ratio is jobs created per accepted request, segmented by client version and retry reason. It should reflect the product's explicit replay actions, not transport noise. Track mismatched-payload rejections separately because a rise there usually points to a client key-lifecycle mistake rather than worker behavior.
The tempting assumption is that a 200 response means the boundary worked. It doesn't. The client may never receive it, and the server may accept the same intent through another instance. Inspect the durable rows and side effects after induced response loss; status codes alone observe only one path through the system.
Know when a stable key is the wrong tool
Stable request records add writes, retention decisions, and a privacy surface. They are not suitable when every submission is intentionally independent, even if the bytes repeat, or when the operation is a cheap read with no lasting side effect. In those cases, keep ordinary transport retries and avoid a durable idempotency ledger.
The catch is retention. Delete request records too early and a late retry can become new work; retain them indefinitely and metadata keeps accumulating. Choose the window from real retry distributions and product replay semantics, then make expiry observable. Your mileage may vary because mobile uploads, command-line clients, and queued integrations have different retry tails.
Content-addressed deduplication is another tool, not a replacement. Use it for immutable artifacts when identical bytes should share storage. Stick with distinct request keys when two users submitting one video must retain separate authorization, settings, audit history, or deletion behavior. The safe decision rule is plain: deduplicate execution by intent, share artifacts only where ownership semantics permit it, and retry each expensive stage against a stable named output.
Before copying this design, measure response-loss retries, concurrent duplicate arrivals, request-record retention, artifact bytes, cache churn, and repeat OCR calls. If those signals cannot be joined by request key and job ID, fix the observability first. Otherwise the system can appear calm while duplicate work continues out of view.
Top comments (0)