
The admin console my team owns exports statements as PDF. I've sat through a compliance review that went over every data flow we had, so I don't sign off on an export pipeline because a screenshot looks right. So when we started shipping fonts with the page, I wanted proof that the PDF actually used them. I wrote a placeholder statement page myself, with a made-up company and invented amounts. The body is set in Noto Sans SC and the title in ZCOOL KuaiLe, a rounded display font that's easy to spot. Both are OFL fonts that were already on my machine. Then I stopped judging the PDF by eye and read its font table.
Three font sources gave three font tables
I printed the same page with page.pdf() in an open-source Chromium 149 build on an M4 Mac, once per font source. With the system font, the PDF held 39 Type3 fonts and weighed 154,603 bytes, and the PingFang names only showed up inside the font descriptors. With the Google Fonts version, printed after document.fonts.ready, the page fetched 16 woff2 slices (412,984 bytes) and the PDF held 15 Type0 subsets, roughly one per slice, at 117,799 bytes. With both full TTFs shipped next to the page, Chromium loaded 13,880,960 bytes of font and wrote a PDF of 71,102 bytes.
Blue is what the page loaded and orange is the PDF. Look at the bottom pair. The biggest font input produced the smallest file, because Chromium subsets fonts when it prints. The six-letter prefix in AAAAAA+ZCOOLKuaiLe-Regular marks a subset, and the KuaiLe program inside was 48,088 bytes, down from 3,274,128. Two things almost fooled me here. Type3 fonts carry no usable BaseFont, so the real name has to come from FontName in the FontDescriptor. My first script skipped that and nearly told me the server PDF had no font names at all. And the Google Fonts Noto is a variable font, listed as NotoSansSCThin, even though the page renders at regular weight.
The server kept the subsets but not my full fonts
The setup I actually worry about is a generator I don't control. For that I used ImgIng (https://imging.ai/) HTML to PDF, with the same placeholder statement imported once as a single file and once as a whole folder. Import and preview run in the browser. On convert, it sends one POST with a script-free snapshot to a server-side Chromium 151 on Linux. I counted 0 non-GET requests before the click and exactly 1 after. That request body is the whole snapshot, so every font you inline travels with it. Worth knowing if someone reviews what leaves your network. The folder import inlined both TTFs as data URLs, and the snapshot came to 17.65 MB. The preview title was in KuaiLe. The PDF title was Noto Sans CJK SC, and the quality report said it had waited for document.fonts and the page images, with no font warning.
The top two rows had the full KuaiLe (3.27 MB) and the full Smiley Sans (2.64 MB), and both came out in the server's Noto Sans CJK SC. The bottom two had the 48 KB KuaiLe subset, and that same subset padded with a dummy table to the size of the full file. Both rendered in KuaiLe. The padded one is why I don't think upload size is the cause. A full 2.05 MB pixel font also worked. I can't see the reason from outside. The upload cost is plain, though. The two full-font runs sent about 18.5 MB each for fonts that never showed up, while a two-font subset sent 501,152 bytes and both fonts made it into a 44,084-byte PDF.
My export check reads the table first
This check now runs on every PDF we generate. It lists every font per type, pulls the name from the descriptor for Type3, and reports which CSS families never got embedded.
PREFIX = re.compile(r"^[A-Z]{6}\+")
def font_inventory(pdf_path):
pdf, inventory = fitz.open(pdf_path), {}
for page in pdf:
for xref, _, ftype, basefont, *_ in page.get_fonts(full=True):
name, size = basefont, 0
if ftype == "Type3": # Type3 keeps its real name in the descriptor
kind, ref = pdf.xref_get_key(xref, "FontDescriptor")
if kind == "xref":
name = pdf.xref_get_key(int(ref.split()[0]), "FontName")[1].lstrip("/")
else:
size = len(pdf.extract_font(xref)[3])
family = PREFIX.sub("", name)
hits, biggest = inventory.get((ftype, family), (0, 0))
inventory[(ftype, family)] = (hits + 1, max(biggest, size))
return inventory
def families_not_embedded(pdf_path, wanted):
got = {fam.lower().replace("-", "") for (ftype, fam) in font_inventory(pdf_path) if ftype != "Type3"}
return [w for w in wanted if not any(g.startswith(w.lower().replace(" ", "")) for g in got)]
full-fonts.pdf
Type3 NotoColorEmoji x2 0 B
Type3 NotoSansCJKsc-Bold x4 0 B
Type3 NotoSansCJKsc-Regular x41 0 B
not embedded: ['ZCOOL KuaiLe', 'Noto Sans SC']
subsets.pdf
Type0 NotoSansSC-Regular x1 326272 B
Type0 ZCOOLKuaiLe-Regular x1 48088 B
not embedded: []
Code notes: full-fonts.pdf is the folder-import server PDF and subsets.pdf the two-subset one, and I cut the Type3 lines from the second listing. Server system fonts come back as Type3, so I only count Type0 fonts as embedded web fonts. The table can't catch everything, though. On the server, 𪚥 (U+2A6A5) was drawn as a crossed box, but text extraction still returns 𪚥. So the check also renders page one to PNG for a person to look at.
In the bottom row, 𠀀 and 𪚥 are crossed boxes and the other five rare characters are fine. I haven't tested a Linux box with no CJK fonts at all, or WOFF2 fonts shipped with the page. Before you trust an HTML-to-PDF export, compare its embedded font names with your @font-face families, and subset your fonts to the characters on the page before you hand them to a generator you don't run.


Top comments (0)