I build yyzTools — a free, local-first Windows toolkit that folds 40+ tools into one installer. Its OCR never uploads a pixel: a bundled RapidOCR engine runs as a local child process. The parts worth writing down are the process lifecycle (a watchdog that's also an idle recycler), the stdin/stdout JSON protocol, and the front-end algorithm that stitches OCR's chopped-up lines back into sentences — then maps the translation onto the original image, line for line.
Screen OCR looks like a checkbox feature until you ship it. The recognition itself is a solved problem — RapidOCR (RapidAI's ONNX port of the PaddleOCR models) runs offline and runs well. Everything around it is the engineering: how you host the engine, how you keep it warm without paying for it all day, and what you do with the raggedy line-by-line output it hands you. Here's how that pipeline is put together.
1. The engine is a child process, not a library
The engine ships as RapidOCR-json.exe under the app's Platform\ directory, and it's treated as what it is: a process with a stdin/stdout JSON protocol. The host starts it once, at model paths chosen per recognition language, and keeps it.
Driving it is then almost boring, which is the point:
- The screenshot (or file) is base64-encoded and written to the child's stdin as one JSON line:
{"image_base64":"..."}. - The result is read back from stdout as JSON — boxes, text, confidence scores — with a 30-second ceiling on the wait.
- The C++ side wraps this as the suite's single async native call (
ocrRecognition), submitted to a task pool so the UI thread never touches the pipe. The front end receives{error: 0, result: [...]}and renders.
The whole driver, minus error handling:
// one recognition round trip
std::string req = "{\"image_base64\":\"" + base64(image_bytes) + "\"}\n";
m_process.Write(req); // child's stdin
std::string json = m_process.Read(30 * 1000); // child's stdout, 30s ceiling
// json: an array of { "box": [[x,y]...x4], "text": "...", "score": 0.97 }
No HTTP, no localhost port, no SDK. A base64 line and a JSON line, in and out.
Two lifecycle details matter more than the protocol:
Language switching is a process restart. The recognition models are loaded at startup from command-line args, so switching, say, from English to Japanese models means terminating the old child and starting a new one with a different model set. It happens rarely, costs a second, and keeps the steady-state contract simple: one process, one model configuration, one protocol.
Tear-down order is a deadlock waiting to happen. The monitor thread waits on the child's process handle; stopping the monitor joins that thread inline. If you stop the monitor before killing the child, the join blocks on a callback that's waiting for a process that nobody has killed yet. Kill the child first, then stop the monitor. (Yes, this was learned the hard way; the comment in the source is a warning to future me.)
2. The watchdog is also an idle recycler
A persistent process is only affordable if it doesn't linger. A 10-second monitor loop handles both failure and frugality:
-
If the child died,
WaitForSingleObjecton its handle returns immediately, and the loop restarts it. Recognition stays available without user intervention. -
If the child is alive, the loop checks whether it should be retired, and this is where the three-condition gate earns its keep. The process is recycled only when all of these hold:
- The OCR feature itself hasn't been used for 3 minutes.
- The system-wide last input is older than 3 minutes (
GetLastInputInfo) — because yyzTools has flows where user activity itself can trigger OCR, and recycling the engine right before it's needed would be self-sabotage. - The child's working set has actually grown past a floor (
GetProcessMemoryInfo) — if the process is still small, killing and relaunching buys nothing.
When all three hold, the process is killed and a clean one starts on the spot. That's the trade: model memory is returned to the system during genuinely idle time, and the next real request pays a cold start instead of the whole machine paying a resident cost.
3. OCR hands you lines; people read sentences
This is the part I'd defend hardest. Recognition output is an array of visual boxes — box + text + score — and a single sentence routinely arrives chopped across several of them, because line breaks and page columns don't care about grammar. Translate those boxes one by one and the output reads like a ransom note.
The front end (ocr-merge.js, new in 1.1.0) does the reconstruction in two passes. The same pipeline backs both surfaces where OCR appears in the suite — the one-hotkey Quick OCR flow (the screenshot at the top) and the full OCR window (below).
Pass one — merge into sentences. Boxes are clustered into visual rows first: sort by the box's center Y, greedily group anything within max(0.5 × row height, 8px) of the current row's running mean, then sort within the row by center X. Rows then merge into sentences only when the previous row doesn't end in terminal punctuation — and the terminator set deliberately excludes commas and enumeration marks, because a comma means the sentence isn't done. Joining itself is script-aware: CJK characters concatenate with no space, everything else gets one.
Pass two — split the translation back. After the merged text is translated, the result has to go back onto the image — that's the feature, seeing the translation where the original was. A proportional split (splitTranslatedToLines) divides the translation into the same number of rows as the original boxes, weighted by how much text each row held, then wraps each row to its box width on the canvas. Same row count, same order, no drift.
The overlay in the first screenshot is what the two passes add up to: an English page where the French translation just... appears, row for row, over the original.
What it doesn't do
Honest limits, as always. The row model assumes roughly horizontal, single-column flow — a two-column PDF can interleave sentence continuity across columns, and the merge will happily stitch the wrong lines together. The 30-second pipe ceiling means a hung engine surfaces as an explicit recognition error, not an infinite spinner. Handwriting and tiny low-contrast text lose characters, exactly as every OCR does. And the engine being local means the first recognition after a cold recycle pays the model-load cost — which is precisely what the idle recycler decided you'd rather trade.
yyzTools is free, no account, no telemetry — the only network calls in the suite are features that inherently need one (translation, update checks); this OCR path is fully offline. Windows 10/11, 12 languages, 387 command modules, yyztools.com. If you've built against RapidOCR/PaddleOCR or driven a CLI engine over pipes, I'd genuinely like to hear how your lifecycle management differs — comments are open.


Top comments (0)