DEV Community

AsyncMonk
AsyncMonk

Posted on

I almost shipped binarize-at-128 before OCR. Gray text vanished

One of my small tool sites is getting a "paste a screenshot, get the text" box, and before wiring it up I wanted to squeeze a bit more accuracy out of it for free. Many OCR tutorials say: convert to grayscale, binarize at 128, then recognize. It's one line of OpenCV and costs nothing on the client, so I planned to make it the default. This is the test I ran before shipping it, not a post-mortem of anything users saw. On one sample, that line deleted every character.

The samples are screenshots of a web page I wrote myself, a made-up "community service hours" notice in Chinese, cut into blocks: body text at 16 px, a sidebar at 12 px, a small fee table, and an 11 px footer in #999 gray on a #f5f5f5 background. Each block was captured at 1x and 2x device pixel ratio, which gives eight crops. I added two more from a holiday-notice template I found online. I scored everything with character error rate, the edit distance to the known text divided by its length. For the OCR side I used ImgIng (https://imging.ai/) with its default Professional OCR tier, and ran Fast and Ultimate OCR over the same ten crops as a check, on an Apple M4 Mac in an open-source Chromium 149 build with the WebGPU backend. The samples are Chinese, so the runs went through its Chinese page. As a second opinion I ran tesseract.js 5 (default parameters, chi_sim+eng).

The 11 px gray footer from my test page: original on top, binarized at 128 in the middle, binarized at 200 at the bottom, shown at 3x

The middle strip, thresholded at 128, is a completely white image. The darkest pixel anywhere in that gray text is 153, so with the cutoff at 128 every pixel counts as background. All three tiers, at both 1x and 2x, returned 0 characters, a 100% error rate, and so did tesseract.js. The result panel read 0 lines, 0 characters, with a hint to try rotating the image, turning off scan enhancement, or switching to a higher tier. Nothing was left to read.

Moving the threshold up didn't fix it cleanly either. At 160 the cutoff sits just above those darkest pixels, so only the core of each stroke survives. Professional OCR read the 1x crop as garbage, 75.5% wrong. The 2x crop has thicker strokes and got 1.0% at the same threshold. At 200, and with Otsu (which picked 208 for this block), the text survived at 3.1% and 5.1%. The untouched original was 0.0%. So the best a threshold did on this block was a little worse than doing nothing.

Across all ten crops the pattern held. Professional OCR on the originals had a combined error rate of 0.1%. Binarizing at 128 pushed that to 25.8%, and 100 made it 32.2%. Otsu came in at 1.1% and adaptive thresholding at 1.4%. Grayscale stayed at 0.1%. The tier gap was smaller than any of that: 0.8%, 0.1% and 0.7% on the originals, less than one point apart. Picking the wrong preprocessing cost about 25 points. Tesseract.js went from 6.2% on the originals to 37.5% at 128, so the old advice didn't rescue it either.

ImgIng has its own toggle for this, "Scan enhancement", described in the English UI as "Gentle grayscale, contrast and sharpening; no forced binarization." It's off by default. With it on, Professional OCR measured 0.2% against 0.1% without, so I wouldn't call it a gain. What I liked is that it doesn't binarize. Recognition runs in a Worker on the machine, and across the whole session I saw zero non-GET requests. The first run downloads the model, 29.8 MB for Professional OCR, which is a download, not an upload of my screenshots. For me that also means no OCR server on my bill, and my cat, whom I named Server, stays the only server I pay to feed.

If you still want a preprocessing step, this is the guard I would put in front of it. It reads the darkest pixel and refuses a fixed threshold that sits below the ink:

import cv2

def safe_to_binarize(path, cutoff=128):
    gray = cv2.imread(path, cv2.IMREAD_GRAYSCALE)
    ink = int(gray.min())
    return ink < cutoff, ink
Enter fullscreen mode Exit fullscreen mode

On the footer crop it prints darkest ink 153, binarize at 128: skip, text would vanish. It only catches the worst case. The 160 threshold that keeps half a stroke passes it.

In the end I didn't ship the guard either. The upload box sends the screenshot untouched. The one fix I did add came from a different sample. A four-column vertical text block scored 92.3% even though the characters were right, because the columns came out left to right instead of right to left. After ImgIng's "rotate left 90°" button, the 2x crop was read with zero errors, so my box gets a rotate button instead of a filter. I only tested screenshots here. The only paper samples I had were synthetic ones made by a script, and real phone photos are still untested. Before you add a binarize line to your own OCR pipeline, check the darkest pixel of your lightest text.

Top comments (0)