DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Chrome records WebM with no duration in the header, and five HEAD requests decide whether you get charged

CogniPrep's interview practice records five video answers in the browser and sends them to object storage, where a worker picks them up. I wrote about the worker already. This is the other half, which turned out to have more sharp edges in it.

You can see the user-facing flow at cogniprep.app/interview.

The MIME type is a server-side decision made from a client-side fact

function detectMimeType(): 'video/mp4' | 'video/webm' {
  if (typeof navigator === 'undefined') return 'video/webm';
  const ua = navigator.userAgent;
  return ua.includes('Safari') && !ua.includes('Chrome') ? 'video/mp4' : 'video/webm';
}
Enter fullscreen mode Exit fullscreen mode

Yes, that is user-agent sniffing, and yes, MediaRecorder.isTypeSupported exists. The reason it is a UA check is that the answer has to be known before recording starts, because the client sends it to the server when it creates the session, and the server bakes it into the presigned upload URLs. A capability probe that disagrees with what the URL was signed for gives you five upload slots you cannot use.

Which is why the server validates it as an enum rather than coercing it:

mimeType: z
  .enum(['video/webm', 'video/mp4'], {
    error: 'Invalid MIME type. Supported types: video/webm, video/mp4',
  })
  .default('video/webm'),
Enter fullscreen mode Exit fullscreen mode

A default is right for an absent field and wrong for a junk one. video/quicktime arriving here should be a 400, not a silent fallback, because the fallback produces uploads the processing stage cannot read and the failure surfaces minutes later in a background job instead of immediately in the response. Default on missing, reject on invalid.

The container that does not know how long it is

This is the one I did not see coming.

Chrome's MediaRecorder, producing WebM, does not write the duration into the header. The spec allows an unknown duration for a live stream, and a stream being recorded in real time is exactly that: at the moment the header is written, the length genuinely is not known yet. Nothing goes back and patches it.

The result is a file that plays, because a player can just decode until it runs out, but whose duration reads as Infinity or 0 depending on who is asking. And the server-side pipeline asks, because it probes each video to validate it and to work out where in the timeline to sample frames from.

The fix is to repair the blob in the browser, where the length is known, because a timer was counting the whole time:

const rawBlob = new Blob(chunksRef.current, { type: mime.current });
// Chrome MediaRecorder WebM files have no duration metadata in the header.
// Fix it in-browser before upload so ffprobe can read it correctly server-side.
const blob =
  mime.current === 'video/webm'
    ? await fixWebmDuration(rawBlob, durationMs, { logger: false })
    : rawBlob;
Enter fullscreen mode Exit fullscreen mode

fix-webm-duration walks the EBML structure and writes the Duration element. Note the condition: only for WebM. Safari's MP4 output has its duration in the moov atom like a normal file, and running a WebM patcher over an MP4 would be a corruption, not a fix.

The general lesson is less about WebM than about where facts live. The duration was known in exactly one place in the entire system, a counter in a React component, and it was being thrown away at the moment the blob was built. The server could not recompute it, could not infer it, and had no one to ask. If a fact only exists on the client, the client has to write it down before the data leaves.

A ref that exists purely because of closure staleness

The duration passed to that call comes from a ref, not from state:

// Mirror of recSecs as a ref so stopRecording always reads the latest value
// regardless of closure staleness (needed for the auto-stop setTimeout path).
const recSecsRef = useRef(0);
Enter fullscreen mode Exit fullscreen mode

Recording stops two ways. The user presses stop, in a handler that React has re-rendered with current state. Or the cap is reached and a setTimeout fires, holding whatever recSecs was when that timeout was scheduled, which is zero.

So the same stopRecording is correct down one path and writes a zero-length duration down the other. A mirrored ref makes both callers read the same current value, and the comment names the specific path that needs it, because that is the thing that a future reader would otherwise "simplify" straight back into the bug.

Five presigned URLs up front, and no charge yet

Creating a session validates the user's credit balance and deliberately does not deduct from it. It mints five presigned PUT URLs, one per question, hands them back with the questions, and the browser uploads directly to storage without the bytes passing through our server at all.

The credit is taken on submit, and only after the bytes are confirmed to exist:

const baseKey = `interviews/${userId}/${interviewId}`;
const existence = await Promise.all(
  Array.from({ length: N_QUESTIONS }, (_, i) => objectExists(`${baseKey}/q${i + 1}`))
);
const missing = existence
  .map((present, i) => (present ? null : i + 1))
  .filter((n): n is number => n !== null);
Enter fullscreen mode Exit fullscreen mode

Five HEAD requests. If any object is missing, the response is a 422 that names the question numbers that did not make it, and nothing is charged. A dropped connection on answer four costs you nothing, and the error tells you which answer to redo rather than offering a generic apology.

Charging at session creation would have been one line simpler and would have billed people for their own flaky wifi.

The order of the last three steps is the whole thing

A submit that charges twice is the worst outcome available, and a double-click is the easiest input in the world to produce. So the route does this:

  1. A cheap status check, which is explicitly documented as not the real guard, because a concurrent duplicate can pass it. It exists to cheaply reject obvious repeats before paying for five HEAD requests.
  2. Confirm the five objects exist.
  3. Confirm the balance is sufficient, still without deducting.
  4. Atomically claim the interview, moving it from pending to processing. One of any number of concurrent submissions wins. The rest get a 409.
  5. Deduct the credit.
  6. Persist the storage key and queue the background job.

Step 4 before step 5 is the design. The claim is a conditional update, so concurrency is resolved by the database before any money is involved. Do it the other way around and two requests both deduct and both queue a job, and now you are reconciling a double charge against a duplicated run.

Both of the last two steps roll back, in opposite directions:

  • If the deduct fails, the claim is handed back to pending so the user can retry. Nothing was charged, so nothing is owed.
  • If queueing the job fails, the credit is refunded and the interview is marked failed, because a charged-but-unqueued interview is a state nobody will ever notice and the user paid for.

That second one is the condition worth naming out loud. It is not "the job failed", it is "the job was never created", and the symptom is a dashboard row that sits there forever looking like it is thinking. Any step sequence where you take payment and then schedule work needs an answer for the gap between them, and "it is a small window" is not one.

See it

Open cogniprep.app/interview and read the four steps under "How it works". Step 2 is the 60 second prep window with a skip, which is the prep stage above. Step 3 is the recorder. Step 4 is the worker from the other post.

Then start a session with DevTools open on the Network tab, and watch what the upload requests are. They are PUTs going straight to a storage host, not to cogniprep.app, each one signed for a single object key and a single content type. Our server sees the URLs being minted and later sees five HEAD responses. It never sees a byte of video.

One-time access covers up to 20 sessions and is listed with everything else on cogniprep.app/pricing.

Top comments (0)