DEV Community

hao jia
hao jia

Posted on

Keep the original text when a PDF page falls back to OCR

A PDF extraction failure becomes harder to investigate when the fallback result replaces the first output. You eventually have readable text, but cannot tell which page needed help or where the replacement came from. I tested a small recording approach using controlled PDF samples. The useful change was keeping each result attached to its page and source, including results that looked wrong. This was a local experiment, not a production deployment.

My main sample was a notice generated with ReportLab. I kept the source wording before producing the PDF, then made separate versions with a missing Unicode mapping and an incorrect mapping. Neither was a document supplied by a customer. Under the same MuPDF rendering conditions, both altered versions looked identical to the original. Their extracted text did not. That gave me a way to examine the extraction result without also changing the visible wording.

The missing mapping removed all 56 Chinese characters from the extracted output. With pypdf, the result still contained 152 characters, exactly the same total as the normal version. A nonempty string would have passed a basic import check. Counting replacement characters would not have helped either: this output contained no U+FFFD. I kept the actual returned string rather than deleting unusual characters to make the report easier to read.

Next I combined a normal page and a damaged page into one PDF. This made the problem with a document-level result obvious. The complete extraction contained Chinese text because the first page was fine. The second page still needed investigation. Once both outputs were flattened into one string, that distinction was harder to maintain. Saving separate page records preserved it without requiring an OCR scheduler or a new document-processing service.

Each record in my experiment carried the file name, page number, raw checks, font observations, and reference comparison. The reference was the original generated wording, not another extractor's output. A page without obvious character anomalies received no_signal, which deliberately did not mean correct. Pages matching their complete known reference were marked sample_checked. Pages without a complete reference remained unverified. These names describe evidence, rather than promising that every downstream use is safe.

The distinction mattered for a second failure. I changed one mapping entry so that two occurrences of a Chinese character were extracted as another valid character. The output still contained 56 Chinese characters and no replacement symbol. PDF.js, pypdf, and PyMuPDF all returned the same wrong substitutions. My basic character checks missed them. Comparing against the independent source text found the mismatch. Agreement between extractors would have hidden it if I had used a majority vote as ground truth.

I also needed a counterexample to avoid sending every missing mapping to recovery. A simple Helvetica document using WinAnsiEncoding contained an English invoice-style sentence and no ToUnicode mapping. All three extractors recovered that sentence correctly. Rejecting it merely because the mapping was absent would have introduced a false alarm. I therefore recorded missing mappings as an observation, while leaving the text evidence and the final decision separate.

Across the eight pages I checked, the records ended with three needing review, three checked against their known source, and two unverified. The last two were pages from an existing copy of RAG-Safety-Bench. Selected passages had been inspected, but I did not have a complete manually checked reference for either page. I did not count those pages as failures or passes. The controlled set also reused the notice, so these numbers are not an accuracy benchmark.

I then tested ImgIng in the browser on its Chinese site, imging.cn. The link to imging.ai is its overseas homepage; these trials did not retest that deployment. For the damaged text PDF, enabling and disabling image-text recognition inside PDF content extraction produced identical text. Converting that file to HTML finished but left the Chinese text damaged. Both outcomes belong in the record. A completed operation is useful evidence about execution; it does not tell us whether the resulting characters match the page.

There was a successful image-recognition path, with an important boundary. The clear PNG came from an external renderer. ImgIng's PDF-to-image attempt on this sample had failed with SVG render failed, so the PNG was not a successful output of that step. Feeding the existing clear image into the separate image OCR tool recovered the notice's five lines. Spaces inside the Chinese text differed from the original PDF. The interface count of 134 characters was not an accuracy measurement.

Actual image OCR result using an externally rendered page

The screenshot shows the image OCR result with an English interface. The notice itself remains Chinese because that was the tested input. It documents a different extraction source, not a repair to the original PDF's mapping. If I kept only this readable result, the failed direct extraction and failed conversion attempt would disappear from the explanation.

For a fallback integration, I would retain the original page record and attach the OCR output as another candidate with its own source and comparison scope. That integration is a proposal; this experiment did not implement retries, queues, or automatic merging. Start by checking whether your current pipeline overwrites the first text result. Preserve it, keep the page number, and compare the replacement against a known passage before allowing the new output to stand in for the original.

Top comments (0)