DEV Community

Cover image for A contract.pdf that was secretly a DOCX broke my file-type auto-detection
InApp
InApp

Posted on Originally published at imapp.blogspot.com

A contract.pdf that was secretly a DOCX broke my file-type auto-detection

When my PDF, DOCX, XLSX, PPTX and EPUB converters ran as separate endpoints, agents kept picking the wrong one. A universal "send anything, we'll figure out the type" endpoint was the obvious fix. My first detection logic: file extension first, Content-Type header as backup. It survived about a day of real traffic.

The file that broke it: a contract named invoice.pdf. My PDF parser opened it and found no %PDF header at all — it was a DOCX somebody had renamed during a mail-merge export. And I couldn't lean on names even in general, because a large share of my uploads arrive as raw base64 with no filename, and most clients stamp anything ambiguous as application/octet-stream. Extensions and headers are rumors. Bytes are facts.

So I rebuilt detection from the bytes up:

  • Starts with %PDF- — easy, it's a PDF.
  • Starts with PK — it's a zip container, unzip and peek inside. A word/document.xml means DOCX, xl/workbook.xml means XLSX, ppt/presentation.xml means PPTX, and a mimetype entry containing application/epub+zip means EPUB. Four of my formats turned out to be zip files; stopping at "it's a zip" would have collapsed four converters into one confidently wrong answer.
  • Otherwise heuristics: a BOM, a doctype, tag density suggests HTML; mostly printable bytes suggests plain text.

Extensions and MIME types got demoted to tiebreakers, not evidence.

Honest residuals: an HTML file stripped of its tags still fools my printable-ratio check occasionally, and sniffing is a 99% solved problem, not 100%. That's also why the per-format converters still exist — sometimes you know the answer better than the bytes do.

The lesson generalizes past file parsing: any pipeline that keys identity off names or headers will eventually meet a renamed file. The first few bytes can't lie, because nobody thinks to edit them.

I ended up packaging this as the Universal Document-to-Markdown API at https://x402.freeq.one/tools/document_to_markdown.html — the endpoint itself is mostly boring glue over those sniffing rules, which is where all the interesting bugs lived.

Top comments (0)