For my capstone I want to pull text out of screenshots that people forward through chat apps, and almost every OCR tip I found says the same thing first: upscale the image 2x before you recognise it. Those screenshots are usually already shrunk by the app, so the tip sounded exactly right. I didn't want to add a resize step on faith, so I tested it. The samples are crops from a web page I wrote myself (a made-up "community service hours" notice in Chinese, with 16 px body text, a 12 px sidebar, an 11 px grey footer and a 13 px fee table), plus one crop of poster text from a holiday-notice template I found online. I shrank each crop to 75%, 60% and 50% to fake what a chat app does. At 50% the 12 px characters are only about 6 px tall.
Then I ran every shrunk crop as-is, upscaled 2x and upscaled 3x (bicubic) through two engines. The reference engine is tesseract.js 5 with default parameters and chi_sim+eng. For the other one I used ImgIng (https://imging.ai/) OCR, which runs PP-OCRv6, with all three tiers (Fast OCR, Professional OCR, Ultimate OCR) on the same crops, in an open-source build of Chromium 149 on an Apple M4 with the WebGPU backend. The score is character error rate: edit distance between the output and the true text, divided by the length of the true text. The core of my checker is below. The real one also folds the different dash characters into one and skips a decorative star, which I left out here.
def squash(s):
return ''.join(unicodedata.normalize('NFKC', s).split())
def char_error_rate(truth, got):
a, b = squash(truth), squash(got)
dist = list(range(len(b) + 1))
for i, ca in enumerate(a, 1):
diag, dist[0] = dist[0], i
for j, cb in enumerate(b, 1):
diag, dist[j] = dist[j], min(dist[j] + 1, dist[j - 1] + 1, diag + (ca != cb))
return dist[-1] / len(a)
bigger = cv2.resize(img, None, fx=2, fy=2, interpolation=cv2.INTER_CUBIC)
The chart combines all 15 shrunk crops, and the two groups go in opposite directions. The tip holds for tesseract.js. Its error rate went from 53.0% as shrunk to 25.8% after a 2x upscale, and 3x only brought it to 25.0%. For ImgIng it didn't hold. Professional OCR went from 2.2% to 2.5%, Fast OCR stayed at 6.6%, and Ultimate OCR moved from 3.1% to 3.3%. One crop got clearly worse. The 12 px sidebar shrunk to 50% went from 14.7% to 23.5% on Professional OCR after bicubic 2x, and 29.4% with Lanczos. I saw the same split on the 1x screenshots I hadn't shrunk at all. On the 11 px and 12 px crops combined, tesseract.js dropped from 7.8% to 2.4% with a 2x upscale, while Professional OCR went from 0.0% to 0.6%.
I honestly don't know why ImgIng doesn't gain anything. My guess was that it resizes every image to some fixed size before detection, but I didn't check that, so all I can report is what came out. My reading is only that the upscale tip was written for engines like tesseract, and a newer model may not need the help.
Two more things changed what I'll do. First, nothing rescues the 50% level. Professional OCR on the grey 11 px footer went from 12.2% to 11.2%, and tesseract.js on the 12 px sidebar from 100% to 92.7%, because those pixels are gone. Second, I compared a real 2x screenshot (taken at device pixel ratio 2, like a Retina screen) with a 1x screenshot upscaled 2x. On the grey footer, tesseract.js got 1.0% on the real one and 2.0% on the upscaled one. On the sidebar both came out at 2.9%, and Professional OCR read both at 0.0%. The gap is small, so I wouldn't lean on it much.
While I was at it I also tried the other classic, a denoise pass. The picture shows why that was a bad idea for screenshots. The top is the 12 px sidebar as captured, and the bottom is the same crop after a 3x3 median filter, and the strokes have melted into blobs. On Professional OCR that crop went from 0% to 32.4%. Recognition runs in a local Worker and made zero non-GET requests the whole time. The first run downloads the model, but that's a download, not an upload of my screenshots. The Scan enhancement toggle says "Gentle grayscale, contrast and sharpening; no forced binarization.", and on my unshrunk crops Professional OCR scored 0.2% with it on and 0.1% with it off, which I count as no change.
So my capstone pipeline now sends the screenshot as it is, with no resize and no filter. When I control the capture I take it at 2x instead of enlarging later. I only keep a 2x upscale as a branch for the tesseract.js fallback, and I check the error rate on my own crops before trusting any preprocessing tip again.

Top comments (0)