DEV Community

Cover image for Building a RAW photo editor in the browser: every number surprised me
Matthieu Audo
Matthieu Audo

Posted on AI-assisted

Building a RAW photo editor in the browser: every number surprised me

Rawnd develops RAW files in a browser tab. Five times, a measurement overturned what the code assumed — and one of them nearly deleted a feature.

Rawnd opens a folder of RAW files in a browser tab, culls the shoot, develops what survives, and writes XMP sidecars that Lightroom reads. No upload, no account, no install. Decoding is WebAssembly, rendering is WebGL2, and the photographs never leave the machine — not as a privacy policy, but because there is no server to send them to.

That part is not interesting. Plenty of things run in a browser now. What was interesting is that the whole product rests on one promise — nothing slows down moving from one photograph to the next — and every time I measured whether I was keeping it, I found out I had been reasoning about the wrong thing.

A 22 MB file to show a thumbnail

A filmstrip of 118 photographs was reading 2.6 GB off disk to display about 170 MB of pixels. The obvious fix is to decode faster. The actual fix was not to decode at all: every RAW carries a full JPEG rendering of the shot, and a TIFF-structured file says exactly where it is, in a directory tree that lives in the first couple of hundred kilobytes.

Reading that tree and then reading only the bytes it points at turns 22 MB per photograph into roughly 1.7 MB. The JPEG that comes out goes straight to createImageBitmap, which decodes it off-thread in the browser's own codec rather than in ours.

The camera already did the work. Seven formats keep their preview in a TIFF directory; Canon's CR3 is an ISO container and Fujifilm's RAF wraps a TIFF at an offset, so those two fall back to the decoder. Handling the common case as a shortcut, rather than making the general path faster, was worth more than any optimisation I could have written.

Two settings beat a week of code

Export was the next thing to be slow: a batch of eight 24-megapixel frames took about 53 seconds. I profiled it step by step expecting to find a hot loop.

Change Gain Kind
JPEG instead of WebP 7.2× one setting
Quality 90 instead of 92 (WebP) 1.83× one setting
Overlap encode with decode — real code

53 s → 13 s, measured at JPEG 90. Two thirds of that came from the first two rows. Quality 92 to 90 is a threshold in the encoder, not a curve — the visual difference is nil and the cost nearly doubles.

The app still ships at JPEG 92, though. That step was measured on WebP, and I have not yet measured whether JPEG has the same one. Moving a default on a number taken from another codec is the mistake the rest of this article is about.

Four hypotheses going in, four of them wrong, including one about the network. The profiler stayed in the codebase rather than being deleted after the session, which is the only reason I can tell you that.

The feature that was tuned four times, blind

Assisted culling groups a shoot into moments, names the sharpest frame of each group, and lists what came out soft, blown or too dark. Grouping, sharpness, clipping and exposure use no model and no inference — what sells as intelligent culling is three quarters arithmetic over data the app already reads. The last quarter is faces: a 15 MB landmark model, run in the tab, finds the eyes so that sharpness is judged where the subject is, and flags a blink.

Grouping was the part that never quite worked. It compared perceptual signatures of each frame, and it had been re-tuned four times: a hash widened from 64 bits to 128, then dropped for a continuous luma grid, a threshold moved, a distance made trimmed rather than plain. Every round was argued from a screenshot of the browser console, because the measurement lived inside a worker inside a browser inside a folder nobody else could open.

So the fifth time, I stopped tuning and built a bench: extract a real folder from a terminal, label the truth by hand in a local page, and score the algorithm against it. It imports the real grouping code from src/, so there is no second copy to drift.

Two folders, 232 photographs, labelled by hand in twenty minutes. The first thing it printed was the baseline.

Folder Precision Recall AUC Recall at 95% precision
chalet-dune 93.1% 26.5% 0.925 6.9%
sunday-catch 60.0% 14.0% 0.831 1.8%

Three to six duplicates out of seven were never seen. And the last column is the one that mattered: no threshold reached a usable trade. The best available anywhere on the curve was 52% precision for 51% recall.

That last column is what a bench buys you. Recall of 26% looks like a threshold set too tight — something a dial fixes. An AUC of 0.83 with 1.8% recall at usable precision is not a dial problem: the distributions overlap, and no dial moves an overlap. I had spent four rounds adjusting a number that could never have been right.

Then one pair explained everything

I sorted the pairs the photographer had called duplicates by how far apart the algorithm thought they were, and opened the worst one.

A wide frame of a child far down a dark forest track. Eight seconds later, a close frame of the same child, face on, in full sun.

— DSC_0502 and DSC_0503, signature distance 1.057

1.057 is further apart than two photographs with nothing in common. Framing, scale, exposure and the subject's position had all changed. The photographer had grouped them without hesitating, because to a human they are obviously one moment — and no comparison of pixels was ever going to agree. Asking it to was the mistake.

So I tried the dumbest possible thing: group on capture time alone, without looking at a single pixel.

Signal Recall at precision
Image only 14% 60%
Clock only 46% 100%
Clock, then image 54% 100%

The clock beat the pixels outright. And the signal had been sitting in the file the whole time: DateTimeOriginal, sub-seconds included, is read during the same TIFF walk that already looks for the preview — 0.8 ms per file, no extra disk read, present on all 732 RAW files I tried.

An earlier version had tried a timestamp and abandoned it, correctly, because it used File.lastModified — which is not when the shutter fired but whenever the last tool to touch the file wrote it. The right idea had been thrown out along with the wrong source.

Where the image still earns its place

Inside a group the photographer drew, the median gap between frames is 2.5 to 3.1 seconds. Between two of their groups, it is 22 to 27. Those two distributions overlap in a narrow band, and that band is the only place appearance is now consulted — where it scores an AUC of 0.78 to 0.89, rather than everywhere, where it did not.

Under five seconds is one moment, no questions asked. Past twenty-five it is not, however alike two frames look. In between, the signature decides.

Folder Before After False positives
chalet-dune 93.1% / 26.5% 96.6% / 55.9% 2
sunday-catch 60.0% / 14.0% 100% / 54.4% 0

Recall doubles on one folder and quadruples on the other, while precision goes up. The remaining misses are pairs minutes apart and visually unalike — neither signal reaches them.

The bench also deleted work I had just done

While the grouping was still weak, I had added colour and a gradient histogram to the signature — the reasoning being that three descriptors failing in different places would beat one failing everywhere. It is a good argument. The bench said it was worth exactly nothing:

Arbiter in the ambiguous band chalet-dune sunday-catch
Luma only 96.6% / 55.9% 100% / 54.4%
Luma and colour 96.6% / 55.9% 100% / 54.4%
Luma and gradients 96.6% / 55.9% 100% / 54.4%

Identical to the digit. Every pair colour or edges would have rejected, luma had already rejected — for 88% more time per photograph and double the storage.

Both descriptors came out. A signature is 256 bytes again, and the whole measuring worker compiles to 2.63 kB. Deleting your own work an hour after writing it is the cheapest thing a bench does, and the easiest to skip without one.

Then colour came back, and five more folders sent it away

A month later the complaint was "lots of near-identical ones". The cautious setting found 55% of the labelled pairs, and the misses had a pattern: the same scene, taken a step closer, 5 to 25 seconds later. A luma grid compares cell by cell, so a step closer looks like another photograph. A chromaticity histogram has no cells.

On the 88 pairs in that band, colour scored an AUC of 0.85 and 0.81 on the two folders, against 0.83 and 0.73 for luma. A MobileNet embedding, run in Chrome, scored 0.84 and 0.84 — for 4 MB and a GPU pass. The 144-byte histogram went in; the model did not.

Fitted on those two folders, colour promised 97% precision for 78% recall. Five more folders were labelled from contact sheets before anything was scored on them, and colour delivered 65%. A birthday shot from one chair, the camera never moving, has the same colours from the cake arriving to the cake being eaten — and the whole party folded into one group.

So the bench learned to choose its settings on six folders and score them on the seventh, once for each of the seven:

Setting, on all seven folders Before After, each chosen without the folder it is scored on
Cautious 98% / 55% unchanged — no colour
Balanced 86% / 67% 94% / 64%
Wide 76% / 78% 86% / 83%

At 95% precision, no colour weight bought more than two points of recall. The cautious default never looks at colour; the looser settings do, where mistakes are already accepted and shown with a one-click undo. It is the previous section's lesson one level up: a number fitted on the data it is scored on is a hope.

What I would rather you knew before trying it

Chrome, Edge and Opera get the full thing, because the File System Access API is what lets Rawnd reopen a folder without asking you to find it again. Firefox and Safari work too — pick a folder or drop it on the page — without that.

There is no defringe — the violet edge on high-contrast contours is longitudinal aberration, it does not follow the radius, so the lateral correction cannot touch it by construction. Camera rendering is learned rather than ported: each body's tone curve is fitted against the JPEG the camera embedded in its own files, from the second photograph on and without a second decode — on a Fuji X100F, that takes the mean gap to the camera's own JPEG from 10.8 to 2.5 levels. Colour itself is not matched yet.

And recall on grouping tops out near 55% at the cautious setting; the looser ones reach 64 to 83% by accepting mistakes you can see and undo. The frames the cautious setting misses are minutes apart and look nothing alike, which is where a learned embedding was supposed to earn its keep. The bench said it does not, yet: MobileNet did no better than a 144-byte colour histogram. Rawnd does ship one model, for faces, and its thresholds were set by reasoning and one real folder rather than by the bench, because the model will not run in Node. That is the one number in this article I would trust least.

If you take one thing

Four rounds of tuning cost more than the instrument would have. Not because the reasoning was sloppy — each round fixed something real — but because none of them could see whether the thing being adjusted was capable of the job at all. The bench answered that in an afternoon, and its first useful output was the news that four rounds of work had been spent on the wrong question.


Rust · WebAssembly · WebGL2 · Vue — 732 RAW files measured, 659 tests. rawnd.app

Top comments (0)