For a few days our analytics showed a class of failure that made no sense: images rejected by the decoder in well under a second. Not slow failures, not out-of-memory on a huge photo — instant refusals. Between 2026-09-06 and 09-10, 60 % of every decode_failed event came from the same shape of file: an iPhone photo called photo.jpg, reporting image/jpeg, whose actual bytes were HEIC.
We had a HEIC path. It was never tried, because the gate that decides whether to load the HEIC decoder was reading the filename.
Disclosure: this is from LensUp, which is ours. The fix is the boring one — read the bytes — but the interesting part is where naive byte-sniffing then went wrong, and the precedence rule we ended up needing.
How a HEIC ends up called .jpg
Nobody renames these on purpose. The path is mundane: an iPhone shoots HEIC, you share the photo through a messenger or copy it to a PC, and somewhere along the way the transport relabels it. The extension becomes .jpg, file.type becomes image/jpeg, and the container is still HEIC.
From the browser's side there is no hint of this. File.type is not derived from content — it is whatever the OS or the transport asserted. So:
// Cheap, synchronous gate so the page can decide whether to lazy-load the HEIC decoder
// module at all. Byte-level truth lives elsewhere — names and MIME lie often enough
// that this is only a candidate check, never a verdict.
export function isHeicCandidate(file) {
const mime = file?.type?.toLowerCase();
if (mime === 'image/heic' || mime === 'image/heif') return true;
const extension = file?.name?.split('.').pop()?.toLowerCase();
return extension === 'heic' || extension === 'heif';
}
That function is fine — as a hint for whether to lazy-load a wasm decoder you do not want to ship to everyone. It is not fine as the thing that decides which decoder runs. That was the bug: a hint had been promoted to a verdict.
Sniffing: twelve to sixteen bytes, never the file
The fix reads a small slice. Not the file — a slice:
const FTYP_READ_BYTES = 64;
export async function sniffInputSignature(file) {
let head;
try {
head = new Uint8Array(await file.slice(0, FTYP_READ_BYTES).arrayBuffer());
} catch { return null; }
if (head.length < 4) return null;
const ascii = (from, to) => String.fromCharCode(...head.subarray(from, Math.min(to, head.length)));
if (head[0] === 0xff && head[1] === 0xd8 && head[2] === 0xff) return 'jpeg';
if (head[0] === 0x89 && ascii(1, 4) === 'PNG') return 'png';
if (ascii(0, 4) === 'GIF8') return 'gif';
if (ascii(0, 4) === 'RIFF' && head.length >= 12 && ascii(8, 12) === 'WEBP') return 'webp';
if (ascii(0, 2) === 'BM') return 'bmp';
if (ascii(0, 4) === '%PDF') return 'pdf';
// … TIFF (II*\0 / MM\0*), then the ISO-BMFF family below
}
Sixteen bytes covers every format the importer accepts, and file.slice() means the read is one small range request against the blob rather than a load of a 12-megapixel photo.
Where "just read the bytes" is not enough
ISO base media files — HEIC, HEIF, AVIF, MP4 and friends — all start the same way: bytes 4–8 spell ftyp, bytes 8–12 carry the major brand. So the naive version is: read the major brand, map heic → HEIC decoder, avif → native.
Then review found the hole. Two of those brands are generic:
const HEIC_BRANDS = new Set(['heic', 'heix', 'hevc', 'hevx', 'heim', 'heis', 'hevm', 'hevs']);
const AVIF_BRANDS = new Set(['avif', 'avis']);
const GENERIC_HEIF_BRANDS = new Set(['mif1', 'msf1']);
mif1 and msf1 mean "this is a HEIF container" and say nothing about the codec inside. An AVIF file may carry mif1 as its major brand and declare avif only further down, in the compatible-brands list. Treat mif1 as HEIC and you hand a file the browser decodes natively — losslessly, instantly — to a lossy wasm transcode. You have recompressed somebody's photo for no reason.
So a generic major brand gets resolved from the compatible brands, and the ftyp box is walked using its own declared size:
if (head.length >= 12 && ascii(4, 8) === 'ftyp') {
const major = ascii(8, 12).trim().toLowerCase();
if (HEIC_BRANDS.has(major)) return 'heic';
if (AVIF_BRANDS.has(major)) return 'avif';
if (GENERIC_HEIF_BRANDS.has(major)) {
// ftyp box: [size:4]['ftyp'][major:4][minor version:4][compatible brands:4 each …]
const boxSize = ((head[0] << 24) | (head[1] << 16) | (head[2] << 8) | head[3]) >>> 0;
const end = Math.min(head.length, boxSize || head.length);
const compatible = [];
for (let offset = 16; offset + 4 <= end; offset += 4) {
compatible.push(ascii(offset, offset + 4).trim().toLowerCase());
}
if (compatible.some((b) => AVIF_BRANDS.has(b))) return 'avif';
if (compatible.some((b) => HEIC_BRANDS.has(b))) return 'heic';
return 'heif'; // a HEIF container that names no codec
}
}
Two details in there that are easy to skip. The box size is clamped with Math.min(head.length, …) so a bogus length cannot walk past the slice we actually read. And when no compatible brand is specific, the answer is the honest 'heif' — "a HEIF container, codec unknown" — rather than a guess.
The precedence rule: bytes win, the label breaks ties
"Always trust the bytes" is the slogan, and it is almost right. The generic-HEIF case is where it breaks: a file whose bytes only say heif could be a HEIC the browser cannot open, or something it can. There is no byte that settles it.
So the rule that shipped is narrower and, I think, the actually correct one: bytes win over the label; the label breaks ties only when the bytes say nothing.
export async function classifyInput(file) {
const signature = await sniffInputSignature(file);
const labelPdf = isPdfCandidate(file);
const labelHeic = isHeicCandidate(file);
const labelTiff = isTiffCandidate(file);
const isPdf = signature === 'pdf' || (labelPdf && signature === null);
const isHeic = !isPdf && (signature === 'heic' || (labelHeic && (signature === null || signature === 'heif')));
const isTiff = !isPdf && !isHeic && (signature === 'tiff' || (labelTiff && signature === null));
return { signature, isPdf, isHeic, isTiff, isImage: !isPdf && !isHeic && !isTiff,
mislabelled: labelPdf && IMAGE_SIGNATURES.has(signature) };
}
Read the three interesting rows:
-
scan.pdfwhose bytes are a JPEG is an image, not a broken PDF. Before this, those went to the PDF parser and producedpdf_invalid: 39 of them between 09-06 and 09-11, most from Windows desktops retrying the same file over and over — the user's file was fine, our routing was not. -
IMG_0001.jpgwhose bytes start with%PDFis a PDF. Same rule, other direction. -
A generic HEIF goes to the HEIC decoder only if the label agrees.
signature === 'heif' && labelHeic→ HEIC. A.jpgwith those bytes stays on the native path, because that is the case where the label is the only evidence that exists.
One more lie: application/octet-stream
Worth its own line, because it is not a type:
// application/octet-stream is what browsers say when they do not know — treat it as
// undeclared so the extension map (and downstream byte sniffing) can decide.
if (declared && declared !== 'application/octet-stream') return declared;
Windows file pickers hand over .pdf files with an empty or application/octet-stream type often enough that treating it as a real MIME type poisons every branch downstream. It means "no opinion", and the code should say so.
The same file also needs a small alias table, because the same format arrives under several names: image/pjpeg, image/x-ms-bmp, image/x-png, image/jpg. Normalise first, decide second.
What I would take from this
A cheap label check and an expensive byte check are different functions and should be named differently. Our bug was one function doing both jobs — a lazy-load hint named like a verdict. isHeicCandidate and isHeicFile now sit next to each other with comments saying which is which, and that naming is doing more work than the sniffer is.
"Read the bytes" needs a tie-break rule, not just a preference. Some formats genuinely do not identify their own contents. Write the precedence down explicitly — bytes, then label, then a documented default — rather than letting it emerge from the order of your ifs.
Failures that are too fast are a routing smell. A decoder that gives up in 200 ms did not struggle with your file; it was handed the wrong file. That timing was the clue that pointed at the label, and it is the one I will look for first next time.
If you want to see the byte path do its thing, drop a HEIC — renamed or not — into a browser-based document scanner. It runs in the tab and files are never uploaded.
Top comments (3)
The tie-break rule is the piece worth stealing: "bytes win" is not a policy until you say what happens when the bytes are silent. Your generic-HEIF case is exactly where a strict sniffer makes things worse — sending a natively decodable photo through a wasm transcode because the label was the only evidence left is a regression dressed as a fix.
Treating application/octet-stream as undeclared rather than as a type matches what I see on upload paths: Windows pickers and a couple of Android share intents hand over octet-stream or an empty type often enough that any branch which bails on an unknown MIME bails on healthy files. The alias table (image/pjpeg, image/x-png, image/jpg) is not cosmetic either — I have had a real PNG rejected by a validator purely because it arrived as image/x-png.
The "too fast is a routing smell" line belongs in a runbook. Sub-second failures are almost never decoding; they are a decision made before any byte was read. Tagging failures by which gate rejected them, instead of by exception type, would have collapsed your five-day tail into one bucket on day one.
The generic-HEIF case bit us before release: our first draft treated mif1 as HEIC, and review caught it sending AVIF files through the wasm transcoder. Checking the ftyp compatible brands fixed it. The gate tagging is what we ended up doing too, just later than we should have. Since last week our processing failures carry the step that failed and the error name, so a sub-second failure shows up as a routing decision right away instead of looking like a decoder problem. Those five days of guessing would have been one query.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.