I made a PDF lose its Chinese text by removing one font mapping. That looked like a satisfying explanation of the problem, until I built another PDF without that mapping and its text extracted correctly. The second file was useful precisely because it did not fail. It forced me to turn a confident diagnosis into a narrower statement: the first file depended on information that the second file could obtain another way.
This is a learning note from a controlled experiment on September 30, 2026. I generated the files with ReportLab 4.5.1 and inspected them with PyMuPDF 1.26.5, pypdf 6.10.2, and PDF.js 5.6.205. None of the damaged files came from a customer or a production incident. Having the source text was the point. I could tell whether an extraction was right without trusting another extraction as my answer key.
The first test supported too broad a guess
The starting document was a one-page holiday notice with five lines of Chinese and English. I retained an untouched copy and removed the Chinese font's ToUnicode entry from another copy. In the original, each extractor returned 56 Han characters. In the modified file, all three returned zero. English and numbers remained, along with unwanted characters. The missing mapping was a strong lead for this particular document, but I had not established a rule about every PDF.
I also rendered the original and modified files with MuPDF 1.26.10, using the same 1.5 scale and RGB settings. Both produced identical 893 by 1263 pixel images. I could still read the Chinese notice in the rendered page. This was the part I wanted to understand: displaying a glyph and recovering the character it represents are different jobs. My experiment changed the extraction result without changing the rendered image under those conditions.
The screenshot is a report made from the saved extraction results, not the interface of a PDF application. The middle column has no Han characters even though the source page remains readable. Notice that the replacement-character count is also zero. Searching for U+FFFD would not have detected this failure. I needed a known sentence from the source, not just a search for a familiar symbol associated with broken text.
A counterexample changed the question
For the second test, I generated a small English file using Helvetica and WinAnsiEncoding. It had no ToUnicode entry. Its single line was “Invoice A-1024: total 1280.50 CNY”, and all three extractors recovered that text correctly. This did not make the first test wrong. It showed that the presence of one dictionary entry was not the whole explanation. The font and encoding context mattered as well.
I checked the distinction against a PDF Association technical presentation about text extraction. It describes several ways to obtain character information, including standard encodings and explicit Unicode mappings. I am keeping that explanation separate from my measurements: the presentation explains the available mechanisms, while my two files show one failing case and one successful counterexample. I did not test every fallback path, and I would not infer one from the output alone.
A practical detour helped me keep that separation. I used the Chinese-language interface of ImgIng as a browser-based comparison on the damaged notice. The native content extraction result still lost the Chinese text. Enabling its option for recognizing text inside images did not change that result. The option's name was not evidence that the existing text had been discarded and the entire page recognized again. I could describe what came out; I could not claim to know the internal decision that produced it.
A present mapping can be wrong too
Removing information was only half the exercise. In another copy of the notice, I kept ToUnicode but changed the entry for 假 to 真. The altered mapping affected two occurrences. The page still rendered as the original notice, yet all three extractors returned the wrong Chinese character in those positions. The count stayed at 56 Han characters. No replacement characters appeared, and agreement between the extractors did not establish correctness.
That example changed how I would write a test for my own small PDF project. I would keep the visible sentence and the expected extracted sentence as separate fixtures. Checking that extraction returns a nonempty string is useful for detecting an empty result, but it cannot detect a plausible wrong character. Even matching the expected character count leaves a gap. For this fixture, I can compare the actual words because I know what was written into the source document.
I would also avoid silently normalizing the output before looking at it. Spaces and line breaks differ across extractors; that deserves its own decision. Replacing or removing control characters immediately can hide clues about the failure I am trying to understand. Keeping the raw string next to a cleaned copy lets me inspect both. A cleaned result should not become the only record of what the extractor actually returned.
My next diagnostic step for an unfamiliar file would be modest: choose a short passage I can read in the page, save the raw extracted passage, and trace the font used by that passage. A missing ToUnicode entry would be a reason to investigate, not a complete verdict. I still need a real-world broken PDF to see how far this exercise transfers. For now, the successful Helvetica counterexample is staying beside the failing notice in my test folder. It stops the simplest explanation from becoming an unjustified rule.
References
- PDF Association technical presentation on text extraction, used to check the distinction between glyph selection and text semantics.
Top comments (0)