DEV Community

Lank_M
Lank_M

Posted on

A PDF permission password is one integer, and the reader decides

A PDF can open with no password prompt and still be encrypted. I generated seven two-page placeholder PDFs to see where that contradiction lives, and it comes down to one signed integer in the encryption dictionary. Nothing in the file forces a reader to respect that integer. Every viewer and library I tried made its own call, and a processing tool has to make one too.

My samples were generated by a script and contain only placeholder text. One has no encryption. Five have only a permission password, meaning the user password is empty and an owner password is set. One needs a password to open. This is what their /Encrypt dictionaries hold:

Sample V / R Cipher /P Denies
S1 V5 R6 AES-256 (/AESV3) -3392 everything except accessibility extraction
S2 V2 R3 RC4 128-bit -3392 same as S1
S3 V4 R4 AES-128 (/AESV2) -20 copying only
S4 V5 R6 AES-256 (/AESV3) -2056 printing and high-quality printing
S5 (open password) V5 R6 AES-256 (/AESV3) -3392 same as S1
S6 V1 R2 RC4 40-bit 0 everything, accessibility included

Reading the P value bit by bit

V and R pick the algorithm and the revision of the security handler. Length is the key size, and CFM names the crypt filter for V4 and V5. The part people actually care about is P. It is a 32-bit field stored as a signed number, and each bit that is set grants one action: bit 3 print, bit 4 modify, bit 5 copy, bit 6 annotate, bit 9 fill forms, bit 10 accessibility extraction, bit 11 assemble, bit 12 high-quality print. A cleared bit is a denial. So -3392, which is 0xFFFFF2C0, keeps bit 10 and clears the other seven. -20 clears only bit 5. -2056 clears bits 3 and 12. Zero clears all of them. I wrote a small check that reads the dictionary from the trailer without decrypting anything and names the cleared bits:

import sys, pypdf

FLAG_BITS = {3: "print", 4: "modify", 5: "copy", 6: "annotate",
             9: "fill-forms", 10: "accessibility", 11: "assemble", 12: "hq-print"}

def describe_encrypt(path):
    trailer = pypdf.PdfReader(path).trailer
    if "/Encrypt" not in trailer:
        return "no /Encrypt"
    enc = trailer["/Encrypt"].get_object()
    p_val = int(enc["/P"])
    cfm = enc.get("/CF", {}).get("/StdCF", {}).get("/CFM", "-")
    denied = [n for b, n in FLAG_BITS.items() if not p_val & (1 << (b - 1))]
    return f"V{enc['/V']} R{enc['/R']} {cfm} P={p_val} denies={','.join(denied) or 'none'}"
Enter fullscreen mode Exit fullscreen mode
S3 V4 R4 /AESV2 P=-20 denies=copy
S4 V5 R6 /AESV3 P=-2056 denies=print,hq-print
S6 V1 R2 - P=0 denies=print,modify,copy,annotate,fill-forms,accessibility,assemble,hq-print
Enter fullscreen mode Exit fullscreen mode

It only reports. It never writes a file and never asks for the owner password.

Why an empty password still encrypts the file

The streams in S1 really are AES-256 encrypted, even though anyone can open it. The key comes from the user password, and the user password is empty, so every reader derives it silently. My reading is that the encryption here is mostly the container the Standard security handler needs to carry P, not a lock on the content. Two libraries even disagree on what to call it. PyMuPDF reports is_encrypted False for S1 once it has opened, while pypdf reports True because the dictionary exists. If your code checks a single boolean, find out which meaning it has.

Who enforces the flags

Once the content is readable, obeying P is a choice each reader makes. I ran the same files through Chromium 149 (an open-source build, not Chrome), Firefox 151, and two libraries.

Grid of five PDF handlers against print, copy and edit on the same permission-only file

Read each row as one handler facing the same restricted file. Chromium's viewer blocks print and copy, yet still lets you draw on the page. Firefox on default settings lets all three through. With pdfjs.enablePermissions turned on it blocks print and editing, and copying still works. The two library rows are grey because they hand the flags to the caller and render as usual. The cipher made no difference. S1, S2 and S6 share a P value and behaved identically in both browsers.

Why a processing tool refuses the file

A compressor has to write a new file, which puts it in the enforcer's seat. The PDF engine in ImgIng is not my part of the product (I work on the on-device codecs), so I only know what the interface shows. I used ImgIng's PDF compression (https://imging.ai/) in its Chinese interface, with my plain file, three permission-only files and the open-password file as input. The plain one shrank by 25.7%. The other four all failed with the same message, 「这份 PDF 有加密或权限保护,需要先解除保护」, which says the file has encryption or permission protection and the protection has to be removed first. It refused S3 too, even though S3 allows modifying and assembling. The whole run made zero non-GET requests. The progress card still said "DONE" next to a line saying 0 files had been compressed, which is a small UI mismatch.

My guess at the reasoning, and it is only a guess, is that a new file has to be either re-encrypted or written without encryption. Re-encrypting means rebuilding the owner's security setup without the owner password. Writing it plain would quietly drop the restrictions someone chose. Refusing every encrypted file is the one option that is easy to defend.

I did not test Acrobat, macOS Preview or the released Safari. If you build anything that reads PDFs, decode P, show the user what the file declares, and point them to the file's author for an unrestricted copy.

Top comments (0)