CogniPrep has an interview practice feature. You record five video answers to behavioural questions in the browser, they go to Cloudflare R2, and a background worker picks them up, transcribes them, looks at frames from each answer and writes you feedback. You can see the user-facing half of it here:
The "How it works" block on that page is four steps. The worker behind it is seven, and the order of them is the only interesting thing about the code.
The constraint
Nothing we record of you is allowed to outlive the run that reads it.
That is not a nice-to-have, it is written into the privacy policy: raw video is deleted from storage as soon as processing completes. So the worker has to delete its own input before it finishes.
Here is the obvious way to write that, and it is wrong:
download videos -> extract -> generate feedback -> delete videos -> persist feedback
The task retries up to three times with exponential backoff. Step 3 is the step most likely to fail, because it is the one that talks to a third-party API over the network. And by the time step 3 has failed, step 4 has not run yet, so that ordering looks safe.
It is not, because any failure after the delete makes the next attempt fail at the download. The videos are the only input a retry can re-fetch. Delete them and attempt two has nothing to work with. A transient 500 from an API becomes a permanent failure, and the user loses a paid session to a blip.
So the delete moved down:
download -> extract -> generate feedback -> persist feedback -> delete videos -> email
Persisting the feedback is now the point of no return. Everything after it is cleanup.
Everything after the point of no return has to be unable to throw
This is the part that is easy to get wrong, because it looks like defensive noise until you work out what throwing actually does.
Trigger.dev gives you an onFailure hook that runs once when a run has permanently failed. Ours is where a failed interview is marked failed and the user's credit is refunded. Which means: if the delete throws, or the "your feedback is ready" email throws, the run is marked failed, and a user who has their feedback sitting in their dashboard gets a refund email telling them it did not work.
So both of those steps swallow their own errors. The delete uses Promise.allSettled across the five keys and logs failures rather than rethrowing:
const results = await Promise.allSettled(
Array.from({ length: N_QUESTIONS }, (_, i) => deleteObject(`${baseKey}/q${i + 1}`))
);
and the email is wrapped in its own try/catch with a log line that says, in as many words, that the feedback is saved. Leftover objects in R2 are recoverable by hand. A completed interview flipped to failed is not.
Two different kinds of failure, two different hooks
The retry story splits in two.
catchError runs on every failed attempt and decides whether retrying is worth it:
catchError: async ({ error }) => {
if (error instanceof InterviewError && !error.recoverable) {
return { skipRetrying: true };
}
return undefined;
}
A recording that is 400MB is not going to be smaller next time. Burning three attempts on it just delays the refund by a minute and costs us three worker runs. Returning undefined defers to the task's retry config, so the default for anything we have not classified is still "try again".
Two conditions skip even that and throw AbortTaskRunError immediately: the interview row does not exist, and the interview has no storage key because the upload never completed. Neither is fixable by waiting.
onFailure runs exactly once, only when the run is permanently dead. That single-run guarantee is what makes the refund correct. A run that fails twice and succeeds on the third attempt is never refunded. A run that fails all three is refunded once. There is no arithmetic anywhere that counts attempts, because the hook's contract does the counting.
onFailure also deletes the recordings, because the success path never got to its delete. It falls back to the deterministic key layout when the stored key was never set, and deleting a key that is already gone is a no-op in R2, so calling it on both paths is safe. The policy still offers a mail-us route for removal, as the backstop for the case where even this delete fails.
The refund is idempotent anyway
Relying on a framework's "runs exactly once" promise for something that touches a balance is a bit much, so the refund guards itself too. It opens a transaction, takes a row lock on the user's credit balance, then checks for an existing refund row:
await tx
.select({ balance: interviewCreditsTable.balance })
.from(interviewCreditsTable)
.where(eq(interviewCreditsTable.user_id, userId))
.limit(1)
.for('update');
The lock is not there for the retry case. It is there because the completion route has its own rollback path, and that path can race the worker. Without the lock both could pass the "already refunded?" check and each credit the balance. With it, the second one sees the row the first inserted and returns false.
It also refuses to refund an interview that has no matching spend row. A refund with no spend is a free credit, and a bug that mints currency is worse than a bug that loses it.
The bits that stay on disk
Frames and audio never reach R2 at all. They are written to the worker's temp directory, fed straight to the API, and removed in a finally block that runs whether the body succeeded or threw. Persisting them would mean keeping screenshots and voice recordings of our users for exactly no downstream reader.
That is the test I would apply to any media pipeline: for each artefact you write, name the thing that reads it later. If nothing does, you are not storing data, you are accruing liability.
See it
Open cogniprep.app/interview and read the four steps under "How it works", then read the Data Retention section of cogniprep.app/privacy and find the line about interview video recordings. The two together are the contract this job exists to keep. One-time access covers up to 20 sessions, and pricing is on cogniprep.app/pricing.
The summary, if you want one sentence: in a job that deletes its own input, the position of the delete is a retry policy, and every step after your point of no return has to be incapable of throwing.
Top comments (0)