DEV Community

Lank_M
Lank_M

Posted on

Split your OCR error rate in two before you touch preprocessing

On September 30 I ran an OCR preprocessing benchmark with 14 versions of each crop: the original, the built-in Scan enhancement, 2x upscaling, grayscale, four fixed thresholds, Otsu, adaptive thresholding, two denoisers and two JPEG qualities. One row barely moved. A small block of vertical Chinese text sat between 89.7% and 94.9% character error rate in every version on the Professional tier, and all three tiers scored 92.3% on its 1x original. That told me the number was tracking something other than pixel quality, so I went back to what the metric actually counts.

The samples were crops from a web page I wrote myself (a made-up notice about community service hours), captured at 1x and 2x, plus two crops from a holiday notice template I found online. For recognition I used ImgIng (https://imging.ai/) on its default Professional OCR tier, plus Fast OCR and Ultimate OCR, on an Apple M4 in a Chromium 149 open-source build with the WebGPU backend. The control was tesseract.js 5 (default parameters, chi_sim+eng). On ImgIng my own work is on-device codecs and model loading. I didn't build the PP-OCRv6 recognition pipeline, so what I say about its internals is inferred from inputs and outputs. Recognition ran in a local worker with zero non-GET requests; the first run of each tier only downloads its model.

Character error rate mixes two failures

CER is edit distance over reference length, and edit distance is positional. If every column is read correctly but the columns come out in the wrong order, the metric sees a block deleted at the front and reinserted at the back, so nearly every character costs an edit. So I added a second number that ignores order: treat both texts as bags of characters, count what they share, and call the rest errors. When CER is high and the order-free rate is low, the characters are right and the sequence is wrong. When both are high, text was really lost or misread.

def triage(ref, hyp):
    a, b = canon(ref), canon(hyp)
    common = sum(min(a.count(ch), b.count(ch)) for ch in set(a))
    cer = edit_distance(a, b) / len(a)
    order_free = (max(len(a), len(b)) - common) / len(a)
    if order_free >= 0.05:
        kind = "characters lost or misread"
    elif cer - order_free >= 0.03:
        kind = "reading order"
    else:
        kind = "clean"
    return f"CER {cer:.1%}  order-free {order_free:.1%}  -> {kind}"
Enter fullscreen mode Exit fullscreen mode
vertical 1x            CER 92.3%  order-free 0.0%  -> reading order
synthetic tilt         CER 4.6%  order-free 0.0%  -> reading order
synthetic shadow+Otsu  CER 31.0%  order-free 31.0%  -> characters lost or misread
12px sidebar 2x        CER 0.0%  order-free 0.0%  -> clean
Enter fullscreen mode Exit fullscreen mode

canon applies NFKC, strips all whitespace and folds the dash variants into one; edit_distance is plain Levenshtein. The 5% and 3% cut-offs are my own picks for this sample set. Each line above is a real Professional OCR output scored against its reference.

The vertical crop needed a rotation

On the vertical block the preview showed a box around each column, and the tiers read the characters inside with an order-free rate of 0% to 7.7%. Judging from the output, each column became one line and lines were emitted left to right, while vertical Chinese reads right to left. The whole passage came out reversed. Rotating left 90° in the interface fixed it: on the 2x crop all three tiers made zero errors, and on the 1x crop Fast and Professional dropped 2 punctuation marks (5.1%) while Ultimate made none. tesseract.js stayed high on both rulers, but I never gave it a vertical model.

Tilted lines swapped places in a synthetic photo

I have no real phone photos yet, only a programmatically synthesized "photo" of an A4 notice: about 2.5° of perspective tilt, uneven warm light, blur, noise, and an extra shadow on sample b. Professional OCR scored 4.6% on a and 15.5% on b, and the order-free rate was 0.0% for both originals and for the upscaled, grayscale, denoised, JPEG and adaptive versions. No character was wrong. Short tail lines were sorted ahead of their own line. Warping the page flat from its four known corners took all three tiers to 0.0% on both samples, apart from 0.6% for Fast on b's adaptive version.

Synthesized photo of a printed notice: original with shadow, Otsu, adaptive threshold

This is a programmatically synthesized simulated photo, not a camera shot. In the middle panel, global Otsu (it picked 135 and 137 on the two samples) turns the shadowed region on b solid black, and both rulers read 31.0%: those characters are gone. On the right, adaptive thresholding keeps every stroke, with 19.0% CER and 0.0% order-free.

Pixel fixes go last

On clean screenshots the pixel side barely matters until you get it wrong. Across the 10 screenshot crops, Professional OCR scored 0.1% on the originals and 25.8% after a threshold of 128. The three tiers on originals were 0.8%, 0.1% and 0.7%, under one point apart. Scan enhancement, off by default, is described as "Gentle grayscale, contrast and sharpening; no forced binarization." On my samples it matched the originals.

Character error rate across 14 preprocessing versions, three tiers plus tesseract.js

Look at the long bars for thresholds 100 and 128 and the 3×3 median filter, against the original's bar near zero. Vertical text is left out of this chart.

Every combination ran once, so a one or two point gap means nothing, and real phone photos are still untested. Now when I see a high error rate I compute the order-free number first. If it is low I fix the geometry first, rotation and tilt, before recognition. Only if it is also high do I look at lighting and contrast, and I binarize last, after measuring how dark the text actually is.

Top comments (0)