A PDF screenshot can stay identical while the extracted words change. I tested that deliberately on September 30, 2026, using a small Chinese holiday notice. The page still said 放假, meaning a holiday or time off. Three text extractors returned 放真 after I changed one character mapping. That is a useful failure case for a document test suite: the output looks like ordinary text, so checks for empty strings and replacement characters will not catch it.
These were controlled files, not documents from a customer incident. I kept a normal notice, removed its text mapping in a second version, and changed one mapping destination in a third. The original file was 29,878 bytes. Its five lines included Chinese, dates, and English, giving me several kinds of text to compare without needing a large fixture. I wanted to know which existing checks would approve a result that a reader would reject.
The screenshot comparison passed
I rendered all three versions with MuPDF 1.26.10 at a scale of 1.5, using the same color settings. Each produced an 893 by 1263 image, and their raw pixel hashes matched. This is a claim about one renderer and configuration. It does not establish that every viewer behaves identically. It does establish that, in this test, a pixel comparison could not distinguish the intact notice from either damaged variant.
The changed mapping affected two occurrences of 假. Their glyphs still appeared correctly on the rendered page, but the mapping now pointed to 真. PyMuPDF 1.26.5, pypdf 6.10.2, and PDF.js 5.6.205 all extracted the wrong word. Agreement between the engines therefore did not settle correctness. They were given the same misleading information, and the visible original remained necessary to decide what the text should say.
This image is a screenshot of my experiment report, not a product interface. The missing-mapping version lost its Chinese text. The incorrectly mapped version retained 56 Chinese characters, matching the normal version. Both the normal and incorrectly mapped outputs contained zero replacement characters. A test that accepts readable Chinese with no replacement symbols would pass the wrong words shown on the right.
The counterexamples narrowed the rule
The missing-mapping variant also exposed a weak length check. PyMuPDF returned 152 characters for both the normal and damaged notices, although the damaged output contained no Chinese characters. Some original positions became other characters or control codes. That makes total length useful as a diagnostic observation, but insufficient as an acceptance condition. The incorrect-mapping variant goes further: even the Chinese count remains unchanged while two words are wrong.
I added a simple English counterexample before turning the mapping check into a gate. A Helvetica document using WinAnsiEncoding had no ToUnicode entry, yet all three engines recovered Invoice A-1024: total 1280.50 CNY correctly. A rule that rejected every missing ToUnicode field would reject that valid result. In my test set, the structural observation belongs beside the text assertion; it cannot replace it in either direction.
A two-page fixture supplied another boundary: its first page used the normal text mapping, while its second used the damaged one. Checking only the first page would miss the problem. I would keep page-specific references for such a fixture and report the failing page directly. This does not establish a statistically valid sampling policy for large collections. It provides a reproducible reason not to treat one good page as proof about the entire document.
Checking the product output
I also tried the normal and missing-mapping notices in ImgIng. The extraction experiment used its Chinese site; the international entry point is ImgIng. The normal result was readable. On the damaged notice, enabling and disabling the option for recognizing text inside images produced identical broken text. That observation concerns the returned output. Without internal traces, it does not tell me which recognition path the service executed.
I did not build the PDF engine in that product. The checks described here are independent fixture tests, and the proposed additions to a regression suite are recommendations derived from them. They do not demonstrate a deployed monitoring system or guarantee arbitrary PDF correctness. In particular, I have not solved how to detect every plausible wrong word without a trusted reference, and changing extraction libraries did not resolve the deliberate wrong-mapping case.
What I would change in the test suite
For an existing visual regression suite, I would start with a short, manually checked reference for a specific page. The holiday title is a useful assertion here because the deliberate error changes its meaning. A date by itself is weaker: the digits survived even when Chinese extraction failed. The reference should therefore include the title and a sentence connecting the dates to the holiday arrangement, with the expected text checked against the rendered original.
The test should record where that reference came from. In this fixture, I have the source text because I generated the document. For an external PDF, someone would need to verify the selected passage against a trustworthy original. Merely copying text from the extractor under test would make the assertion circular. Saving the exact input file or its digest alongside the page number also prevents a later document revision from silently invalidating the reference.
I would compare the expected words before accepting broader cleanup rules. Removing layout-only line breaks may be appropriate for one output contract, but replacing similar-looking Chinese characters would erase the failure this fixture is designed to expose. Punctuation, spacing, and paragraph order can have separate rules. The important choice is to write down which changes are permitted instead of treating any readable sentence as an acceptable reconstruction.
For the next document test I add, I would keep the screenshot assertion and place a page-specific word assertion beside it. A failure report should show the expected passage, the actual passage, and the source page, so a reviewer can inspect the disagreement directly. The screenshot can answer whether the page still looks right. The reference text answers whether the selected words survived. This experiment needed both answers before I could call its output correct.

Top comments (0)