DEV Community

Cover image for My PDF converter errored on a file every reader could open — owner passwords aren't open passwords
InApp
InApp

Posted on Originally published at imapp.blogspot.com

My PDF converter errored on a file every reader could open — owner passwords aren't open passwords

A legal team once sent my PDF-to-Markdown converter a 40-page services contract. Every desktop reader opened it fine. My converter threw File has not been decrypted.

I assumed corruption. It wasn't. PDFs carry two independent passwords: a user password, needed to open the file, and an owner password, which only restricts printing, copying, and text extraction. Publishers and legal departments love owner passwords — the document opens everywhere but refuses to give up its text. My extraction library saw the encryption dictionary and stopped, even though the user password was empty.

The fix turned out to be one deliberate line: when extraction fails with a not-decrypted error, call decrypt with an empty string before giving up. An empty user password is how these files ship — the viewer does the same trick silently, it just never mentions it. After `decrypt(""), every page extracted normally. The permissions flag turned out to be advisory, not security.

But the deeper lesson was about failure modes. Before this incident, an unreadable PDF returned empty Markdown. That's the worst possible output for a RAG pipeline: it looks like success, gets chunked, embedded, and quietly poisons the index with nothing. Now the converter returns a structured error that distinguishes between 'not a PDF at all', 'image-only scan with no text layer', 'open password required', and 'extraction failed'. Callers can route a permissions-locked contract to a human instead of indexing a void.

Two small things I'd tell anyone building the same thing:

  1. Treat PDF encryption as a tri-level state, not a boolean: no encryption, empty-user-password encryption, real encryption. Handle the middle state explicitly.
  2. Never let a document pipeline return an empty-string success. Empty output should always be distinguishable from 'nothing was wrong'.

I ended up packaging all of this as my PDF-to-Markdown API (https://x402.freeq.one/tools/pdf_to_markdown.html) — direct URL or base64 bytes in, headings, lists and tables out. If you're feeding contracts or reports into an agent's context, check what your converter does with an owner-password file before you trust its empty outputs.

Top comments (1)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The empty-user-password case is a useful distinction because an encryption flag alone does not tell the caller whether user input is needed. Your structured error categories also make it possible to separate an extraction problem from a file that should be sent through OCR.

I would keep a small regression corpus covering the three encryption states, plus a mixed document containing both scanned and selectable-text pages. A document-level success could otherwise hide missing scanned sections. Reporting extracted-page coverage alongside the error category would help callers decide whether the resulting Markdown is complete enough to index.