DEV Community

NullPointerZen
NullPointerZen

Posted on

I tested the OCR switch instead of trusting its label

An OCR checkbox is easy to overinterpret. I see it next to a PDF extractor and assume it will recognize the page again when the existing text is broken. That assumption deserves a test. In a controlled example this week, enabling the checkbox produced exactly the same damaged text as disabling it. The tool reached its completed state both times. The Chinese sentence I needed did not return.

The useful question turned out to be narrower than whether the application “supports OCR.” I wanted to know whether this particular option replaced an unreliable text interpretation with a fresh reading of the visible page. Those are different claims. A feature label cannot tell me which one happened, but comparing the actual output can. This post follows that comparison and the separate image-recognition step that eventually gave me usable text.

The input was a synthetic one-page notice, created for testing rather than taken from someone’s company documents. It contained five lines of Chinese and English, including a holiday interval, a return date, and a deadline. One copy retained its font’s ToUnicode mapping. In another copy, that mapping was removed without changing the page’s drawing instructions. This let me compare two known inputs rather than guess what an unfamiliar PDF originally contained.

Under the same MuPDF rendering conditions, both copies produced identical 893 by 1263 pixel page images. Text extraction told a different story. PyMuPDF, pypdf, and PDF.js each recovered 56 Chinese characters from the normal copy and none from the missing-mapping copy. English and digits survived. I therefore had a precise failure to test: could the browser tool recover the missing Chinese content, rather than merely produce another nonempty result?

I chose a sentence containing both Chinese words and dates as the comparison target. Checking only the year would have been misleading because the broken output still contained numbers. I also kept the unedited output from each attempt. That made it possible to distinguish an actual improvement from a nicer preview, a different line break, or a newly enabled export button. I wanted changed evidence, not a changed interface state.

For the browser portion I used ImgIng, an image and document utility with a public homepage at imging.ai. The tested deployment was its Chinese site, imging.cn, on September 30, 2026. I have not repeated the entire PDF workflow on the international domain. The result screenshot below came from switching that tested tool to its English interface and rerunning image recognition on the same test image.

In the PDF content extractor, the option specifically referred to recognizing text inside images. I processed the missing-mapping PDF once with it enabled and once with it disabled. The resulting text was identical. The title retained its year, parts of the body retained dates and odd characters, and the English line remained readable. Whatever the tool’s internal decision was, this option did not replace the damaged native Chinese text in this sample.

That finding does not establish that the checkbox never works. I did not test every combination of scanned pages, embedded images, and native text. It does establish a stopping point for this input. Repeating the same operation with the same setting would not be evidence of a new recovery route. I needed either a different extraction result or a correctly rendered page image that an independent recognition step could actually inspect.

I tried the PDF-to-image route in the same product next: one page, PNG, and 216 DPI. It failed with “SVG render failed.” This matters because the later OCR result would be easy to present as one uninterrupted success story. It was not. The usable PNG came from an existing external rendering step. It was 105,697 bytes and recorded as a 216 DPI page image. It was not downloaded from the failed online conversion.

I then imported that clear image into the separate Image OCR tool. The default Professional tier returned five lines, with 134 characters reported by the interface. The holiday dates, return date, deadline, and Chinese wording were available for inspection again. That count is not an accuracy percentage. I compared the actual lines against the visible source, and spacing within the Chinese text differed from the PDF’s original spacing.

Actual English interface after recognizing the controlled page image

The distinction between these two OCR operations is the recommendation I would keep. If native extraction is already correct, I would use it. If it is wrong and a clear page image exists, a separate image-recognition step is worth testing. Before relying on an OCR option inside a PDF workflow, I would check whether the disputed sentence changed. The presence of a checkbox, a completion badge, or an export format does not answer that question.

There is also a limit to judging by readability. A separate controlled variant changed the mapping for the Chinese character 假 to 真 while leaving the visible glyph unchanged. All three extractors agreed on the wrong text, with two affected occurrences. The result still contained Chinese and had the expected Chinese character count. Agreement between tools was therefore useful diagnostic information, but the visible source remained necessary for checking meaning.

My browser automation captured the result previews and their text. An attempted save action did not produce a download event in that automated run, so I am not claiming that every local save dialog was verified. Likewise, obtaining OCR text did not repair the original PDF’s mapping or create a new searchable PDF. Those would be additional output requirements. Keeping the original file, the recognition image, and the checked text separate made the limits easier to see.

For a quick test on your own material, choose one disputed sentence before changing tools. Preserve the original extraction, run the proposed alternative, and compare the same words rather than overall character counts. If the route switches to images, confirm that the image is actually readable before recognizing it. Finally, check the text where you intend to use it. That last check catches the unglamorous mistakes, including copying an older result after the newer one finally became correct.

Top comments (0)