DEV Community

Cover image for Self-hosted OCR without a cloud API: where Tesseract stops working
Julie Love for Apryse

Posted on Originally published at linkedin.com

Self-hosted OCR without a cloud API: where Tesseract stops working

Cover image caption: Puppy with "free" sign and collar that reads "self-hosted OCR". Next to that a long receipt titled "Maintenance" with many cost items for dogs.

Quick Answer: Ask your favorite AI how to run OCR server-side without a cloud API and you will most likely hear about Tesseract. That is a reasonable answer for clean, straight, single-column English print. It stops being reasonable once you read the weak spots Tesseract's own documentation flags, and start counting developer time as the cost that it is. A production OCR SDK, like Apryse, earns its license fee in exactly those places.

I asked ChatGPT, Claude and Gemini that question in several forms. Every one said Tesseract, usually with PaddleOCR and EasyOCR as close seconds. Not one mentioned that Tesseract publishes a page called "Improving the quality of the output" which, if we are being honest, is a long list of what it just does not do well.

It's pretty easy to see why these top the lists. Because Tesseract is free, genuinely good, 30+ years old, and has more written about it than any other OCR engine on this lovely little planet. Models synthesize consensus, and people sure love to write and talk about things that are "free". But it's distinctly not the only answer, and you need to make sure it's actually the right answer to YOUR question.

Where does Tesseract stop working?

Quick Answer: Five places, all of them documented by the Tesseract project itself. Skewed pages, tables, uneven page backgrounds, tightly cropped or heavily bordered regions, and anything that is not a sentence of dictionary words. None of these throw an error, they just return bad data - and that is usually the worse outcome.

Straight from the Tesseract documentation:

  • Skew. "The quality of Tesseract's line segmentation reduces significantly if a page is too skewed." Deskew is YOUR problem to fix, not the engine's.
  • Tables. "It is known tesseract has a problem to recognize text/data from tables without custom segmentation/layout analysis." Their wording, and the suggested remedy is a GitHub thread.
  • Uneven backgrounds. Binarisation runs internally with Otsu, and the result "can be suboptimal, particularly if the page background is of uneven darkness." Which describes most scans older than a decade.
  • Borders and crops. Too little border causes problems. Too much on a small region returns an empty page. The fix is cropping to a 10 pixel margin, by hand, per document type.
  • Non-prose text. Tesseract is "optimized to recognize sentences of words," so for part numbers and receipts the advice is to disable the built-in dictionaries. Invoices are mostly not sentences.

Then you add the 14 page segmentation modes you have to specify with --psm, because Tesseract will not determine the layout for you. Tesseract is an excellent recognition engine wrapped in a pipeline you are expected to build yourself, out of Leptonica, OpenCV or ImageMagick. And things you build, you must maintain.

What does "free" OCR actually cost?

Quick Answer: The license does not cost money, but you pay in developer time. The pre-processing pipeline, the per-document-type tuning, the model curation and the retraining are all work somebody on your team does instead of working on your product. Free OCR moves the cost from pocketbook to time - and that's some of the most expensive currency you have.

One of our customers came to Apryse having already built a full, working Tesseract solution. It was running just fine. The problem was that it was costing them 2-3 developers to keep it running. 2-3 devs who were constantly having to try and learn new things about the OCR ecosystem and how to keep it working. 2-3 devs who were supposed to be improving the PDF software that was actually the company's business.

They did not switch because Tesseract could not read their documents. They switched so Apryse was their expert, and responsible for the OCR upkeep. It meant their team could get back to focusing on their expertise and the internal roadmap items.

Once you look for it, you can see where those developer-hours go.

  • The pre-processing pipeline. Deskew, despeckle, binarisation, border handling, alpha channel stripping.
  • Language bindings. Tesseract is a C++ program. Reaching it from Python, Java, .NET or Node.js means a third-party wrapper, and pytesseract works by shelling out to the command-line binary as a subprocess. Those wrappers are community-maintained on their own release cycles, so they are one more dependency sitting between you and your text.
  • Per-document-type tuning. Choosing among 14 page segmentation modes, toggling dictionaries off for invoices and part numbers, cropping to the right margin. Every new document type restarts this.
  • Model curation. Tesseract publishes three sets of model files. tessdata_best is the most accurate and the only one you can fine-tune from, and it is also the slowest. tessdata_fast is documented as the least accurate, and it is what ships with Linux distributions. Most people are running the least accurate models and were never told. Each language is a separate file you download and manage.
  • Retraining, when preprocessing runs out. The training guide is honest about the price: the shipped Latin models were built from roughly 400,000 text lines across about 4,500 fonts, a training run takes "a few days to a couple of weeks," and it officially "only works on Linux." Retraining from scratch is called "a daunting task," with the warning that you will likely end up with a model that does well on your training data and badly on your real data.
  • Support. There is no one to call. Tesseract's own training guide asks you not to file training problems as GitHub issues and directs you to a mailing list instead.

To their credit, the project tells you not to retrain: "unless you're using a very unusual font or a new language, retraining Tesseract is unlikely to help." Honest advice, and it also leaves you nowhere to go on the day preprocessing has not closed the gap and the invoices still have to be processed.

For comparison, the Apryse v12 default engine covers 80+ languages including CJK from one model set. No files to curate, no accuracy tier to pick, no training step, and a support contract when it misbehaves. (Not to mention tools like Smart Data Extraction that help you do more with your OCR.)

None of this makes Tesseract a bad choice. It makes it a choice with a staffing plan attached. Price both.

What does a production OCR SDK give you that Tesseract does not?

Quick Answer: Pre-processing built into the engine rather than bolted on in front of it, several recognition engines behind one SDK so you can switch per document type, structured output with coordinates, and a support contract. You are paying for the pipeline around the recognition, not the recognition alone.

With the Apryse Server SDK, deskew and despeckle happen inside the engine. Since v12 the default engine runs on deep learning neural networks, bringing roughly 16% better word recognition, 80+ languages including CJK, and no GPU requirement. There are other engines available when your pages are unusual, which my OCR article covered in detail, and swapping between them is a one-line change rather than a new integration.

The bindings are native rather than wrappers. The same call runs from C++, .NET, Java, Python, Node.js, Go, PHP and Ruby, on Windows, Linux and macOS:

opts = OCROptions()
opts.SetEngine("default")      # "default", "alternative" or "iris"
opts.AddLang("eng")

# Inspect the result as JSON before it becomes a searchable PDF
json = OCRModule.GetOCRJsonFromPDF(doc, opts)
OCRModule.ApplyOCRJsonToPDF(doc, json)
Enter fullscreen mode Exit fullscreen mode

That JSON is worth pausing on. You get recognized text plus coordinates, which means you can filter low-confidence words, reject a page, or map any answer back to the exact spot it came from, before any of it reaches your database or your LLM.

With Apryse you also get a multitude of tools that are ready to handle most any document need - AND you have Support. When something isn't going right, there's real help and a team to back you up.

Does self-hosting OCR solve compliance?

Quick Answer: While it does remove a whole category of problems, that's not the same as solving compliance for you. Running recognition in your own environment means no third-party processor in scope, no data residency question and no per-page metering. The rest of your obligations under HIPAA or GDPR are still architecture you have to build.

You may still have ninety-nine problems, but data residency ain't one.

Worth being precise, because vendors are routinely vague about it. An SDK is not "HIPAA certified." HIPAA compliance, as I'm sure you're aware if this is your area, is a property of your whole system. What self-hosting changes is that documents never leave your environment, so there is no processor agreement to negotiate, no cross-border transfer to justify and no third party whose breach becomes your incident.

The safety department of a large city came to us having done real homework, testing a range of options against two hard requirements: no training with their data, and nothing leaves their systems. The AI-powered engine in v12 met both and gave them the best accuracy of everything they tested.

And while we're on security - for those who get to be part of those ever fun security audits, nice to know that Apryse holds ISO/IEC 27001:2022 certification since 2018 and completes annual SOC 2 Type II audits. Usually keeps those security form pain points to a minimum.

Where should you start?

Quick Answer: Fix your input, then benchmark Tesseract honestly against your own worst documents before you spend anything. If it clears your bar and you have the headcount to maintain it, you've got your answer. If it fails in the ways its own documentation predicts or your budget does not account for maintenance efforts, you have a business case rather than a preference.

  1. Check your DPI before anything else. Below 300 and you are debugging the scanner, not the engine. On a rasterized PDF, divide image pixels by page size in inches.
  2. Build a test set of your worst fifty pages. Rotated, stained, multi-column, tabular. Good scans don't prove out your real cases.
  3. Run Tesseract first, with real effort. Deskew, strip the alpha channel, pick the right --psm. Make sure it's a solid test to compare against.
  4. Measure Character Error Rate and Word Error Rate per document type. Accuracy and how you manage it matters.
  5. Price the developers, not just the license. Make some real estimates on what it is going to take to keep up with this longterm. When you are your only support team, it makes a difference.
  6. Test the Apryse engines on the same set. Every module is in the 30-day trial with no page cap.

Tesseract is not the wrong answer. It is a good answer to "what is the best free OCR engine." But what if what you really need to know is "what will do the best job on my docs with the least longterm cost to my dev team"?

Models will only answer a question explicitly asked. Make sure you're asking the question you need answered.

Happy Coding!


Julie Love is Director of Developer Experience at Apryse. She has spent an entire career in tech spanning everything from burning CD screen savers to mastering proprietary code, sales engineering and a variety of in between. Puns always intended.


Further reading

Top comments (0)