After trying three PDF extractors on the same damaged notice, I still did not have the Chinese text I needed. The outputs looked different, but none contained the original Chinese title. That distinction matters when I am deciding whether another dependency is worth adding to a small tool. A different-looking failure is not yet a better deliverable. Before trying another converter, I want a reason to expect it to obtain information that the previous route could not recover.
This was a controlled experiment, not a client incident. On September 30, 2026, I used a synthetic, one-page holiday notice containing five lines of Chinese and English. I kept the original, made a copy without the Chinese font's ToUnicode mapping, and made another copy with one incorrect mapping. Having the source text let me judge the outputs directly. I did not need to treat one extractor's answer as the standard that every other extractor should match.
The original notice contained 56 Han characters. PyMuPDF 1.26.5, pypdf 6.10.2, and PDF.js 5.6.205 all recovered them. In the copy with the mapping removed, all three recovered zero Han characters. Numbers and English survived, mixed with unwanted characters. Some of those unwanted characters differed between implementations. None of that variation restored the words I needed. At that point, adding a fourth implementation would need a specific diagnostic purpose rather than a hope that another button might help.
The other damaged copy gave me a stronger reason to keep the source involved. Its mapping changed 假 to 真, affecting two occurrences in the notice. All three extractors returned the same wrong characters. Their agreement was real, but it was agreement with an incorrect mapping. The rendered page still showed the original wording under my MuPDF test conditions. This was not a case where choosing the most common extraction result would establish what the document actually said.
This is a report assembled from saved experimental outputs, not a product screen. The last column keeps the expected Han-character count while failing the word-level comparison. For a small project, that changes what I would ask a document provider for. A statement that “the PDF opens correctly” does not answer the question. A short, confirmed source passage does. It gives me something concrete to compare before I promise that an import is ready to use.
I also tried the Chinese-language interface of ImgIng as a browser comparison. On the notice with the mapping removed, enabling the option to recognize text inside images produced the same content-extraction output as leaving it off. Converting the file to HTML did not restore the Chinese text either. The HTML task reached a completed state, but completion described the conversion, not the correctness of the text I wanted to hand over.
There was a working recovery step, with an important condition. Image OCR recovered readable content from a clear PNG that had already been rendered externally. The Chinese spacing differed from the source, so the result still needed review. ImgIng's own PDF-to-image operation failed on this sample at the tested single-page PNG setting. I could not describe the recovery as one successful chain inside a single application. The external image was a separate input with a separate origin.
If I were accepting a similar document-import request, this is where I would ask whether an editable source or a confirmed text export was available. That is a proposed working choice based on this experiment, not a claim that I measured the cheapest process across customer projects. I did not measure support time or calculate a monetary saving. My reason for asking is narrower: another output container has not supplied missing semantic information, while source content may let me check the exact fields the requester needs.
The request can be small. I would send the page number, the readable passage, and the extracted result, then ask for the corresponding source text. I would not automatically ask for an entire project archive when a title and a few dates are enough to evaluate the current problem. If the provider sends a freshly exported PDF, I would rerun the same comparison. A new file timestamp tells me the file changed; it does not establish that the text now matches.
If the source no longer exists, I would describe the next deliverable as reviewed OCR text from selected page images. That wording leaves room for the work still required: obtaining a correct image, recognizing it, and checking the fields that matter. It also avoids claiming that the original PDF has been repaired. I would keep the failed direct extraction alongside the reviewed result so that a later correction can be traced to the route that produced it.
The stopping rule I took from this test is simple enough to use without building a larger system: after a second extraction route reproduces the same important error, pause and identify what new evidence another attempt could provide. If I cannot name it, I would ask for source content or agree on a reviewed image-based route. I want to hand over text whose origin and remaining uncertainty I can explain.
Top comments (0)