"We have crash recovery" is an easy thing to claim. I wanted to know what happens if the process disappears right now.
The question is not specific to Electron. It applies to a media exporter, downloader, migration runner or AI workflow that writes files locally. Restarting an app is easy. Recovering enough state to finish the user's work is the part worth testing.
I used an Electron recorder as the example below. The mechanics are simple: persist something, kill the process, restart, find the leftovers, rebuild, then inspect the result.
For ProgressCut, a local-first desktop app that turns a work session into a short visual story, "right now" has two uncomfortable meanings:
- while screenshots are being captured;
- while FFmpeg is assembling the final MP4.
In both cases, raw frames can exist while the output is partial or absent.
I started isolated Electron processes, created real session folders, sent SIGKILL, then started fresh Electron processes and asked them to find and rebuild what remained. SIGKILL ends a process immediately, so shutdown handlers cannot quietly clean up the state first.
The test found a rendering bug too: a story budgeted for 3.00 seconds encoded to 3.46 seconds.
What had to survive
If there are captured frames, a later app process must be able to offer a rebuild without recording again.
The implementation is a per-session manifest stored under the user's local application-support directory. It contains the output folder, raw-frame folder, target duration, output format and timestamps. It does not contain screenshot data.
The record has to be per session. A single active-session.json works until two incomplete sessions exist and the second one overwrites the first.
session A → recovery/<hash(A)>.json
session B → recovery/<hash(B)>.json
Every manifest is written as a temporary file and then renamed into place. A successful recovery clears only the manifest for the recovered session.
What I actually killed
The smoke test uses synthetic PNGs instead of recording a real screen. Everything after that is production code: Electron startup, session orchestration, manifest storage, story selection and FFmpeg.
The harness runs two scenarios:
| Scenario | Kill point | What a restart must prove |
|---|---|---|
| Capture interruption | after the third PNG is persisted | the three frames are discovered and remain byte-for-byte unchanged |
| Encoding interruption | after the real FFmpeg process starts | a fresh process finds the session and can render a valid MP4 |
The parent launches each Electron app in its own process group and sends SIGKILL to that group. This matters: killing the Electron parent while leaving FFmpeg behind does not simulate an interruption cleanly.
After each kill, a fresh Electron process reads local recovery manifests. The harness rebuilds every discovered session, runs ffprobe (FFmpeg's inspection tool) on the output and verifies:
- H.264 video;
-
yuv420ppixel format; - encoded duration within 150 ms of the requested three seconds;
- raw PNG bytes unchanged;
- recovery metadata removed only for the session that completed.
The command is part of the repository:
pnpm test:crash-recovery
If you are testing a different app
Replace "frames" with whatever your app leaves behind: uploaded chunks, edited documents, downloaded files, generated reports or queued jobs.
- Choose the smallest irreversible user value. In my case, it is a persisted screenshot; in your app it might be a database checkpoint or a downloaded part file.
- Write a recoverable record before expensive work begins. Keep enough information to resume, but avoid storing sensitive content in that record.
- Interrupt the real process at a deterministic checkpoint. Do not mock the thing whose behaviour you are claiming to test.
- Start a fresh process with the same storage. A retry in the same memory space is not recovery.
- Verify the final artefact, then scope cleanup. Check bytes, schema, duration, checksum or domain result, not only that a file exists.
In pseudocode it looks like this:
await runUntil("checkpoint:persisted");
killProcessGroup();
const restartedApp = await launchWithSameDataDirectory();
const pending = await restartedApp.discoverRecoverableWork();
const result = await restartedApp.resume(pending[0]);
expect(await inspect(result)).toMatchObject(expectedArtefact);
expect(await sourceBytes()).toEqual(bytesBeforeCrash);
expect(await pendingRecords()).toEqual([]);
I check cleanup separately because it is easy to get wrong: stale records cause repeated recovery prompts, while broad cleanup can erase another unfinished job.
The failure I wasn't looking for
The first recovery run failed on a stricter check than "the file exists."
The story target was three seconds. ffprobe measured 3.458333 seconds.
That was annoying, because the video looked fine. It played, the codec was right, nothing seemed broken. A test that only checked for a file would have passed it, and I would have shipped the bug.
The renderer uses FFmpeg's concat demuxer because story moments have variable durations. The usual workaround repeats the last file so FFmpeg honours that last still frame. The side effect was extra tail time in the encoded video.
I capped the encoded output at the story's calculated duration and kept a regression test that creates weighted moments and measures the resulting file with ffprobe.
await runFfmpeg([
// concat input and video filter omitted
"-movflags", "+faststart",
"-t", (story.totalDurationMs / 1000).toString(),
"-y", outputPath,
]);
That is the whole reason to inspect the output rather than assert that a file exists. A report can exist with missing rows; a video can play and still have the wrong length.
What passed
On 2026-10-05, the smoke ran on macOS 26.3, an arm64 Mac17,2 and Node.js 22.14.0:
capture interruption → 3 frames discovered → recovered
encoding interruption → 3 frames discovered → recovered
source PNGs → unchanged
output → H.264 / yuv420p / 3.00 s
I also tested the unsigned packaged arm64 .app rather than only the development process. The app launched with a temporary profile, loaded Sharp from inside app.asar, accepted a production IPC re-render request and generated a three-second MP4.
CSC_IDENTITY_AUTO_DISCOVERY=false pnpm run pack
pnpm test:packaged
That catches a different class of failure: code that works in a workspace because it accidentally resolves a native dependency from the developer machine.
What this does not prove
This covers process crashes and nothing more.
- It does not test power loss or filesystem durability beyond atomic rename.
- It does not test a multi-hour recording, sleep/wake, monitor changes or a full disk.
- It does not test Screen Recording permission UX.
- It does not test Developer ID signing, notarization, Gatekeeper or a clean-machine install.
- It does not evaluate whether the story selector makes better summaries. That requires real held-out sessions and blind review.
Why I added this before release
For a local-first tool, the useful data can already be on disk when the UI is gone. In this app, throwing away an unfinished session because FFmpeg died would be a bad default.
Each run of the harness goes through the same loop:
Next I need a real two-to-five-hour session with resource measurements, sleep/wake and an intentional interruption. This smoke test narrows the gap; it does not replace that run.
Source: ProgressCut on GitHub, MIT licensed.




Top comments (0)