<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Julie Love</title>
    <description>The latest articles on DEV Community by Julie Love (@julie_love).</description>
    <link>https://dev.to/julie_love</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4106394%2F82c532e7-7d26-4b08-94ce-9c2de3d16fce.jpg</url>
      <title>DEV Community: Julie Love</title>
      <link>https://dev.to/julie_love</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/julie_love"/>
    <language>en</language>
    <item>
      <title>Can you vibe code a DOCX editor for $200?</title>
      <dc:creator>Julie Love</dc:creator>
      <pubDate>Thu, 10 Sep 2026 20:44:10 +0000</pubDate>
      <link>https://dev.to/apryse/can-you-vibe-code-a-docx-editor-for-200-1hfp</link>
      <guid>https://dev.to/apryse/can-you-vibe-code-a-docx-editor-for-200-1hfp</guid>
      <description>&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; A weekend with Claude will get you a browser-based editor that opens a .docx, renders the paragraphs, lets you type, and saves it back out. That part is real and it is genuinely impressive. What the $200 does not cover is the rest of the OOXML spec, the security work that comes with accepting files, the round-trip problem, and someone who understands the whole thing for the 2am fire drill. Generating code got cheap. Owning code did not.&lt;/p&gt;

&lt;p&gt;I recently found this comment under a demo video for the &lt;a href="https://apryse.com/capabilities/docx-editor" rel="noopener noreferrer"&gt;Apryse DOCX editor&lt;/a&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"You can vibe code one better than that one with Claude so it will probably cost you 200 and you are better off doing that way since people are definitely gate keeping code that is technically free to make."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I think there are a lot of people excited about the possibilities of where we are with vibe coding. I hate to ruin a good buzz but can't resist bringing a bit of sobriety into this conversation. Sadly, the reality of owning code is never quite as fun as the imagination phase.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does a weekend and $200 actually get you?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Real code that runs. A .docx is a ZIP file full of XML, so pulling paragraph text out of document.xml is genuinely an afternoon's work, and what comes back opens real files and saves them again. But there is quite an expanse between "working" and great code. The distance between them is ECMA-376: several thousand pages, of which your weekend covered whatever your test file happened to use. Almost none of the rest is exotic.&lt;/p&gt;

&lt;p&gt;Here is what is sitting past that quick vibe output, and none of it is unusual. This is just Tuesday for anyone who works on documents:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Numbering.&lt;/strong&gt; Abstract definitions, concrete instances, per-level overrides, and restart semantics that nobody agrees on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Styles.&lt;/strong&gt; Inheritance running from document defaults, through linked and latent styles, down to direct formatting on a single run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Section properties.&lt;/strong&gt; Which can change halfway down the file, because of course they can.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tables.&lt;/strong&gt; Vertical merges expressed as continuation flags on the cells below rather than on the cell doing the merging.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fields.&lt;/strong&gt; Storing both a cached result and the recipe to recompute it, so you have to decide which one you believe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Everything else.&lt;/strong&gt; Footnotes, content controls, equations, bidirectional text, embedded objects.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that is hidden from you. It is just a whole lot more work than people are expecting to do when they "open a .docx".&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The first 90 percent of the code accounts for the first 90 percent of the development time. The remaining 10 percent of the code accounts for the other 90 percent of the development time. - Tom Cargill, Bell Labs&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why does every "can it also do..." request cost more than it sounds?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Because the cost of saying yes collapsed but the cost of the resulting system did not. Version 0.1 gets shown around, the requests come back, each one sounds like an afternoon, and each one lands somewhere much larger than the person asking realizes.&lt;/p&gt;

&lt;p&gt;Where those small requests actually land:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"Can it do numbered lists?"&lt;/strong&gt; Abstract versus concrete numbering, level overrides, restart rules, list styles.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"The indents look wrong."&lt;/strong&gt; Indent resolution and tab stops. Congratulations, you are now maintaining a layout engine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Print it exactly like Word."&lt;/strong&gt; Line breaking, font metric substitution, widow control, keep-with-next, rows splitting across pages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Keep the tracked changes."&lt;/strong&gt; Insertion, deletion, move and format revisions, and every single edit your editor makes now has to be expressible as one of them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Where did the comments go?"&lt;/strong&gt; See the round-trip section. This one is the mean one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Six weeks later there are twenty thousand lines in the repo and you have read maybe three thousand of them. The three thousand that broke.&lt;/p&gt;

&lt;h2&gt;
  
  
  What if you don't control what is uploaded?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; The moment you accept uploads, you have an endpoint consuming arbitrary compressed XML sent by strangers. That is one of the more hostile input classes in computing.&lt;/p&gt;

&lt;p&gt;A few of the vulnerability classes to consider:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;XXE.&lt;/strong&gt; Your XML parser will read files off your server unless somebody explicitly told it not to.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decompression bombs.&lt;/strong&gt; A few kilobytes of ZIP that expands to gigabytes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zip slip.&lt;/strong&gt; Archive entries containing &lt;code&gt;../&lt;/code&gt;, writing wherever they like the second you extract to disk.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every one of those appears in the CVE history of every mature document library, found the expensive way - by someone else, years ago.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is the round-trip problem?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; The failure people brace for is a crash. The failure that actually happens is silent. If you parse a document into your own model and serialize it back out, it drops everything your model did not know about - with no error report in sight.&lt;/p&gt;

&lt;p&gt;Someone opens a contract in your editor, changes one word, saves. The tracked changes are gone. So are the comments, the custom XML bindings a downstream system reads, the content controls, and the numbering that survived four rounds of legal review.&lt;/p&gt;

&lt;p&gt;The file opens fine. Word does not complain. It simply is not the same document anymore, and nobody finds out until the moment it matters most.&lt;/p&gt;

&lt;p&gt;Preserving the parts you do not understand is much harder than parsing the parts you do. No prompt is going to tell you which parts you did not understand, because by definition, you didn't ask.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can you maintain code you have never read?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Software cost lives in maintenance, and maintenance cost is a function of comprehension. Code generation drives the writing cost way lower but it does nothing for the comprehension cost. In fact it most likely makes comprehension worse, because writing the code is how you would otherwise have learned the code.&lt;/p&gt;

&lt;p&gt;The question at the point of failure is never "can Claude fix this." It is "do I understand this system well enough to know whether that fix is right, or whether it just moved the bug somewhere I am not currently looking."&lt;/p&gt;

&lt;p&gt;Writing scales with the model. Debugging still scales with you. Even after coding for many years, I still struggle to remember exactly what I meant in every line of code that I actually wrote - much less trying to decipher where the AI went wrong on something I had nothing to do with.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is anyone actually gatekeeping this?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; No, and this is the part of the comment that holds up the least. Plenty of sources out there. &lt;a href="https://www.libreoffice.org/" rel="noopener noreferrer"&gt;LibreOffice&lt;/a&gt; is open. So are &lt;a href="https://www.docx4java.org/" rel="noopener noreferrer"&gt;docx4j&lt;/a&gt;, &lt;a href="https://python-docx.readthedocs.io/" rel="noopener noreferrer"&gt;python-docx&lt;/a&gt; and a dozen readable OOXML implementations. Nobody is hiding the code, because the code was never the scarce part.&lt;/p&gt;

&lt;p&gt;Everyone deserves to be paid for their work. When you make the call to buy vs build, it's rarely because no one is capable of building it. You pay to keep your dev resources focused on solving your own business problems. You pay for that other company to be the expert for your team.&lt;/p&gt;

&lt;p&gt;That expert brings years of experience with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The zillions of real documents produced by 20 years of Word versions and every third-party generator that emits technically invalid XML which Word renders anyway.&lt;/li&gt;
&lt;li&gt;The bug report from the customer whose merged table cells collapse only at 120% zoom.&lt;/li&gt;
&lt;li&gt;Somebody to call when a filing renders wrong just before the deadline.&lt;/li&gt;
&lt;li&gt;The unglamorous rest of it: security review, indemnity, accessibility conformance.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is not about a goblin wildly guarding lines of code while sitting on a pile of your money. This is just the part you cannot prompt into existence because it is not code. It is lived, often traumatic, experiences that you are paying to skip entirely. ...and it's often worth every saved minute of your time and avoided suffering.&lt;/p&gt;

&lt;p&gt;Toss this problem back over the fence? Yes please, and thank you.&lt;/p&gt;

&lt;h2&gt;
  
  
  So when should you build it yourself?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Frequently. For a great many jobs, spending $200 and a weekend is exactly the right call. Not every tool needs to reach perfection or hold up under long term testing and maintenance stress.&lt;/p&gt;

&lt;p&gt;Build it yourself when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You control the whole workflow,&lt;/strong&gt; use your documents, and you know exactly how the in and out points need to function.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your company or workload is small,&lt;/strong&gt; no need to over-engineer for something that doesn't need to scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Wrong" costs a redo,&lt;/strong&gt; not a lawsuit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nobody outside the building is uploading anything,&lt;/strong&gt; or more specifically, you can entirely trust the users of the system.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You would be happy to delete the whole thing next quarter.&lt;/strong&gt; Throw it up, try it out, and no one is upset when you scrap it for the next iteration.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Reach for something battle-tested when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The inputs are other people's documents,&lt;/strong&gt; arriving from anywhere, in any state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The data is other people's data,&lt;/strong&gt; with residency or compliance rules attached.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The deadline is other people's deadline,&lt;/strong&gt; and the failure lands on them. ...or worse, lands on you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You need enterprise level scaling.&lt;/strong&gt; There is no point in starting over on trying to account for every edge case an enterprise-level roll out will turn up.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where should you start?
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Build the weekend version anyway.&lt;/strong&gt; It is the fastest way to find out how much of the format your use case actually touches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Round-trip a really complicated document early.&lt;/strong&gt; Tracked changes, comments, content controls. Save it, reopen it in Word, and see what quietly vanished.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the code you shipped.&lt;/strong&gt; All of it, or at least enough to know which parts you have not read.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide which 90% you are in for.&lt;/strong&gt; The first 90% is a weekend now, which is legitimately worth being excited about. The second 90% is the whole rest of your job.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you get to step 4 and only have interest in that first 90%, then that is when a document SDK earns its keep. You can try ours for 30 days at &lt;a href="http://docs.apryse.com/guides/get-started" rel="noopener noreferrer"&gt;docs.apryse.com/guides/get-started&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you're on step 4 and think, "I've got this," go forth and prompt away! Then come back and tell me what broke. I have no doubt the answers here will keep changing as things advance - I'm super interested to see where we are.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Disclosure:&lt;/strong&gt; I work for a company that sells document SDKs, so weigh all of the above accordingly. The vulnerability classes and the round-trip problem are real either way. Go check them against whatever you build.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Self-hosted OCR without a cloud API: where Tesseract stops working</title>
      <dc:creator>Julie Love</dc:creator>
      <pubDate>Wed, 09 Sep 2026 18:49:17 +0000</pubDate>
      <link>https://dev.to/apryse/self-hosted-ocr-without-a-cloud-api-where-tesseract-stops-working-224o</link>
      <guid>https://dev.to/apryse/self-hosted-ocr-without-a-cloud-api-where-tesseract-stops-working-224o</guid>
      <description>&lt;p&gt;&lt;em&gt;Cover image caption: Puppy with "free" sign and collar that reads "self-hosted OCR".  Next to that a long receipt titled "Maintenance" with many cost items for dogs.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Ask your favorite AI how to run OCR server-side without a cloud API and you will most likely hear about Tesseract. That is a reasonable answer for clean, straight, single-column English print. It stops being reasonable once you read the weak spots Tesseract's own documentation flags, and start counting developer time as the cost that it is. A production OCR SDK, like Apryse, earns its license fee in exactly those places.&lt;/p&gt;

&lt;p&gt;I asked ChatGPT, Claude and Gemini that question in several forms. Every one said Tesseract, usually with PaddleOCR and EasyOCR as close seconds. Not one mentioned that Tesseract publishes a page called "Improving the quality of the output" which, if we are being honest, is a long list of what it just does not do well.&lt;/p&gt;

&lt;p&gt;It's pretty easy to see why these top the lists. Because Tesseract is free, genuinely good, 30+ years old, and has more written about it than any other OCR engine on this lovely little planet. Models synthesize consensus, and people sure love to write and talk about things that are "free". But it's distinctly not the only answer, and you need to make sure it's actually the right answer to YOUR question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where does Tesseract stop working?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Five places, all of them documented by the Tesseract project itself. Skewed pages, tables, uneven page backgrounds, tightly cropped or heavily bordered regions, and anything that is not a sentence of dictionary words. None of these throw an error, they just return bad data - and that is usually the worse outcome.&lt;/p&gt;

&lt;p&gt;Straight from the &lt;a href="https://tesseract-ocr.github.io/tessdoc/ImproveQuality.html" rel="noopener noreferrer"&gt;Tesseract documentation&lt;/a&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Skew.&lt;/strong&gt; "The quality of Tesseract's line segmentation reduces significantly if a page is too skewed." Deskew is YOUR problem to fix, not the engine's.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tables.&lt;/strong&gt; "It is known tesseract has a problem to recognize text/data from tables without custom segmentation/layout analysis." Their wording, and the suggested remedy is a GitHub thread.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uneven backgrounds.&lt;/strong&gt; Binarisation runs internally with Otsu, and the result "can be suboptimal, particularly if the page background is of uneven darkness." Which describes most scans older than a decade.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Borders and crops.&lt;/strong&gt; Too little border causes problems. Too much on a small region returns an empty page. The fix is cropping to a 10 pixel margin, by hand, per document type.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Non-prose text.&lt;/strong&gt; Tesseract is "optimized to recognize sentences of words," so for part numbers and receipts the advice is to disable the built-in dictionaries. Invoices are mostly not sentences.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then you add the 14 page segmentation modes you have to specify with &lt;code&gt;--psm&lt;/code&gt;, because Tesseract will not determine the layout for you. Tesseract is an excellent recognition engine wrapped in a pipeline you are expected to build yourself, out of Leptonica, OpenCV or ImageMagick. And things you build, you must maintain.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does "free" OCR actually cost?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; The license does not cost money, but you pay in developer time. The pre-processing pipeline, the per-document-type tuning, the model curation and the retraining are all work somebody on your team does instead of working on your product. Free OCR moves the cost from pocketbook to time - and that's some of the most expensive currency you have.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;One of our customers came to Apryse having already built a full, working Tesseract solution. It was running just fine. The problem was that it was costing them 2-3 developers to keep it running. 2-3 devs who were constantly having to try and learn new things about the OCR ecosystem and how to keep it working. 2-3 devs who were supposed to be improving the PDF software that was actually the company's business.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;They did not switch because Tesseract could not read their documents. They switched so Apryse was their expert, and responsible for the OCR upkeep.  It meant their team could get back to focusing on their expertise and the internal roadmap items.&lt;/p&gt;

&lt;p&gt;Once you look for it, you can see where those developer-hours go.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The pre-processing pipeline.&lt;/strong&gt; Deskew, despeckle, binarisation, border handling, alpha channel stripping.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Language bindings.&lt;/strong&gt; Tesseract is a C++ program. Reaching it from Python, Java, .NET or Node.js means a third-party wrapper, and pytesseract works by shelling out to the command-line binary as a subprocess. Those wrappers are community-maintained on their own release cycles, so they are one more dependency sitting between you and your text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-document-type tuning.&lt;/strong&gt; Choosing among 14 page segmentation modes, toggling dictionaries off for invoices and part numbers, cropping to the right margin. Every new document type restarts this.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model curation.&lt;/strong&gt; Tesseract publishes three sets of model files. &lt;code&gt;tessdata_best&lt;/code&gt; is the most accurate and the only one you can fine-tune from, and it is also the slowest. &lt;code&gt;tessdata_fast&lt;/code&gt; is documented as the &lt;em&gt;least&lt;/em&gt; accurate, and it is what ships with Linux distributions. Most people are running the least accurate models and were never told. Each language is a separate file you download and manage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retraining, when preprocessing runs out.&lt;/strong&gt; The &lt;a href="https://tesseract-ocr.github.io/tessdoc/tess5/TrainingTesseract-5.html" rel="noopener noreferrer"&gt;training guide&lt;/a&gt; is honest about the price: the shipped Latin models were built from roughly &lt;strong&gt;400,000 text lines across about 4,500 fonts&lt;/strong&gt;, a training run takes "a few &lt;em&gt;days&lt;/em&gt; to a couple of &lt;em&gt;weeks&lt;/em&gt;," and it officially "only works on Linux." Retraining from scratch is called "a daunting task," with the warning that you will likely end up with a model that does well on your training data and badly on your real data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Support.&lt;/strong&gt; There is no one to call. Tesseract's own training guide asks you not to file training problems as GitHub issues and directs you to a mailing list instead.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To their credit, the project tells you not to retrain: "unless you're using a very unusual font or a new language, retraining Tesseract is unlikely to help." Honest advice, and it also leaves you nowhere to go on the day preprocessing has not closed the gap and the invoices still have to be processed.&lt;/p&gt;

&lt;p&gt;For comparison, the Apryse v12 default engine covers &lt;strong&gt;80+ languages&lt;/strong&gt; including CJK from one model set. No files to curate, no accuracy tier to pick, no training step, and a support contract when it misbehaves. (Not to mention tools like Smart Data Extraction that help you do more with your OCR.)&lt;/p&gt;

&lt;p&gt;None of this makes Tesseract a bad choice. It makes it a choice with a staffing plan attached. Price both.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does a production OCR SDK give you that Tesseract does not?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Pre-processing built into the engine rather than bolted on in front of it, several recognition engines behind one SDK so you can switch per document type, structured output with coordinates, and a support contract. You are paying for the pipeline around the recognition, not the recognition alone.&lt;/p&gt;

&lt;p&gt;With the &lt;a href="https://docs.apryse.com/core/guides/ocr" rel="noopener noreferrer"&gt;Apryse Server SDK&lt;/a&gt;, deskew and despeckle happen inside the engine. Since v12 the default engine runs on deep learning neural networks, bringing roughly &lt;strong&gt;16% better word recognition&lt;/strong&gt;, 80+ languages including CJK, and no GPU requirement. There are other engines available when your pages are unusual, which &lt;a href="https://www.linkedin.com/pulse/how-does-ai-text-recognition-work-ocr-icr-python-pipeline-julie-love-k2nbc/" rel="noopener noreferrer"&gt;my OCR article&lt;/a&gt; covered in detail, and swapping between them is a one-line change rather than a new integration.&lt;/p&gt;

&lt;p&gt;The bindings are native rather than wrappers. The same call runs from C++, .NET, Java, Python, Node.js, Go, PHP and Ruby, on Windows, Linux and macOS:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;opts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OCROptions&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;opts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;SetEngine&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;default&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="c1"&gt;# "default", "alternative" or "iris"
&lt;/span&gt;&lt;span class="n"&gt;opts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;AddLang&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eng&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Inspect the result as JSON before it becomes a searchable PDF
&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;OCRModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;GetOCRJsonFromPDF&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;opts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;OCRModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;ApplyOCRJsonToPDF&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That JSON is worth pausing on. You get recognized text plus coordinates, which means you can filter low-confidence words, reject a page, or map any answer back to the exact spot it came from, before any of it reaches your database or your LLM.&lt;/p&gt;

&lt;p&gt;With Apryse you also get a multitude of tools that are ready to handle most any document need - AND you have Support. When something isn't going right, there's real help and a team to back you up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does self-hosting OCR solve compliance?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; While it does remove a whole category of problems, that's not the same as solving compliance for you. Running recognition in your own environment means no third-party processor in scope, no data residency question and no per-page metering. The rest of your obligations under HIPAA or GDPR are still architecture you have to build.&lt;/p&gt;

&lt;p&gt;You may still have ninety-nine problems, but data residency ain't one.&lt;/p&gt;

&lt;p&gt;Worth being precise, because vendors are routinely vague about it. An SDK is not "HIPAA certified." HIPAA compliance, as I'm sure you're aware if this is your area, is a property of your whole system. What self-hosting changes is that documents never leave your environment, so there is no processor agreement to negotiate, no cross-border transfer to justify and no third party whose breach becomes your incident.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The safety department of a large city came to us having done real homework, testing a range of options against two hard requirements: no training with their data, and nothing leaves their systems. The AI-powered engine in v12 met both and gave them the best accuracy of everything they tested.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And while we're on security - for those who get to be part of those ever fun security audits, nice to know that Apryse holds &lt;strong&gt;ISO/IEC 27001:2022&lt;/strong&gt; certification since 2018 and completes &lt;strong&gt;annual SOC 2 Type II&lt;/strong&gt; audits. Usually keeps those security form pain points to a minimum.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where should you start?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Fix your input, then benchmark Tesseract honestly against your own worst documents before you spend anything. If it clears your bar and you have the headcount to maintain it, you've got your answer. If it fails in the ways its own documentation predicts or your budget does not account for maintenance efforts, you have a business case rather than a preference.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Check your DPI before anything else.&lt;/strong&gt; Below 300 and you are debugging the scanner, not the engine. On a rasterized PDF, divide image pixels by page size in inches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build a test set of your worst fifty pages.&lt;/strong&gt; Rotated, stained, multi-column, tabular. Good scans don't prove out your real cases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run Tesseract first, with real effort.&lt;/strong&gt; Deskew, strip the alpha channel, pick the right &lt;code&gt;--psm&lt;/code&gt;. Make sure it's a solid test to compare against.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure Character Error Rate and Word Error Rate per document type.&lt;/strong&gt; Accuracy and how you manage it matters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Price the developers, not just the license.&lt;/strong&gt; Make some real estimates on what it is going to take to keep up with this longterm.  When you are your only support team, it makes a difference.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test the Apryse engines on the same set.&lt;/strong&gt; Every module is in the &lt;a href="https://docs.apryse.com/guides/get-started" rel="noopener noreferrer"&gt;30-day trial&lt;/a&gt; with no page cap.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Tesseract is not the wrong answer. It is a good answer to "what is the best free OCR engine." But what if what you really need to know is "what will do the best job on my docs with the least longterm cost to my dev team"?&lt;/p&gt;

&lt;p&gt;Models will only answer a question explicitly asked.  Make sure you're asking the question you need answered.&lt;/p&gt;

&lt;p&gt;Happy Coding!&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Julie Love&lt;/strong&gt; is Director of Developer Experience at Apryse. She has spent an entire career in tech spanning everything from burning CD screen savers to mastering proprietary code, sales engineering and a variety of in between. Puns always intended.&lt;/p&gt;




&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/server-side-ocr-tools-features-formats-architecture" rel="noopener noreferrer"&gt;Server-side OCR: how to process documents at scale without third-party APIs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.apryse.com/core/guides/ocr" rel="noopener noreferrer"&gt;OCR module overview and supported languages&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/ocr-in-python" rel="noopener noreferrer"&gt;How to build OCR in Python with Apryse&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/api-vs-sdk-differences-use-cases" rel="noopener noreferrer"&gt;Local SDKs vs cloud APIs: what developers need to know&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.linkedin.com/pulse/how-does-ai-text-recognition-work-ocr-icr-python-pipeline-julie-love-k2nbc/" rel="noopener noreferrer"&gt;How does AI text recognition work? OCR, ICR and the Python pipeline that feeds your LLM&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.linkedin.com/pulse/icr-handwriting-recognition-python-when-do-you-need-self-hosted-love-dpadc/" rel="noopener noreferrer"&gt;ICR handwriting recognition in Python: when do you need self-hosted instead of a vision LLM?&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
    </item>
    <item>
      <title>ICR handwriting recognition in Python: when do you need self-hosted instead of a vision LLM?</title>
      <dc:creator>Julie Love</dc:creator>
      <pubDate>Wed, 09 Sep 2026 18:41:44 +0000</pubDate>
      <link>https://dev.to/apryse/icr-handwriting-recognition-in-python-when-do-you-need-self-hosted-instead-of-a-vision-llm-i56</link>
      <guid>https://dev.to/apryse/icr-handwriting-recognition-in-python-when-do-you-need-self-hosted-instead-of-a-vision-llm-i56</guid>
      <description>&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Self-hosted ICR is a recognition engine trained on handwriting that runs inside your own network and can be set to return text plus coordinates. A vision LLM reads handwriting using an enormous training set and usually returns a block of text from someone else's server. Pay close attention to your needs around deployment, cost and audit-ability to make the best choices for you.&lt;/p&gt;

&lt;p&gt;I've heard learning cursive in school is a thing of the past, much like my spiral notebook and my handwritten to-do lists.&lt;/p&gt;

&lt;p&gt;Yes, there are usually doodles in my notebooks as well. It is great for those of us with visual and spatial memories who need to put ideas together on paper before they make sense anywhere else. (At least that is my justification and I am sticking to it.)&lt;/p&gt;

&lt;p&gt;Luckily, kids these days do not need to learn cursive while we have so many cool digital toys now.&lt;/p&gt;

&lt;p&gt;Let's break it down. (Due to the limitations of the written word, you will need to mentally cue a great late-80s hip hop track and a dance of your choice.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is handwriting so much harder to read than printed text?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Printed text follows solid, definable, repeatable rules. Handwriting breaks five of them at once, which is why handwriting support is not just a quick add-on to a printed-text engine. It has to be a separate model trained on separate data.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Everyone writes differently.&lt;/strong&gt; This may be one of the world's largest understatements. Fonts are finite. Handwriting variation is not. (in the uncountable infinity category, for my fellow math nerds.) Even the same person writes the same letter differently depending on the number of lattes consumed and how annoying the traffic was on their way to fill out the form.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cursive has no gaps.&lt;/strong&gt; No need to mind the gap here, because there are none. No reliable boundary between characters means nowhere obvious to cut.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Words wander.&lt;/strong&gt; Print sits on the line. Handwriting drifts, slants, and takes a little trip to any nearby region that looks fun or available.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Training data is scarce.&lt;/strong&gt; Labeled handwriting collections are far smaller and rarer than printed text. Side quest, handwriting sample set: &lt;a href="https://fki.tic.heia-fr.ch/databases/iam-handwriting-database" rel="noopener noreferrer"&gt;IAM Handwriting Database&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every source looks different.&lt;/strong&gt; A doctor's note has nothing in common with an 18th-century census record. And let's be real, you thought teaching a computer to beat a human at Go was hard? There is no quantity of doctors' notes large enough to solve that great mystery of life.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then there are forms, which pile all of the challenges into one place. Typed labels that have to relate to handwritten answers, deciphered separately and reassembled correctly.&lt;/p&gt;

&lt;p&gt;One practical note before you start building: If you have any influence over the forms themselves, use it. Grid-style boxes force people to space out their letters and define a usable area for each field. If that's too limiting, keep your field boxes black (non-scan colors like red, pink or cyan can drop out completely in the pre-process), and your field labels above or to the left of the fields (not inside or touching the boundaries). Blank lines invite the sprawl that breaks everything downstream. Design with scaffolding wherever you can.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is ICR, and how is it different from OCR?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Intelligent Character Recognition (ICR) is a neural network trained on handwriting examples rather than font definitions, so it can read writing it has never seen before. Optical Character Recognition (OCR) is based on matching the shape of a letter and expected characteristics - based on printed fonts.&lt;/p&gt;

&lt;p&gt;Also of note, this is ever evolving.  AI-powered OCR can sometimes pick up handwriting pretty well. As AI continues its expected exponential growth, will be interesting to see where this all ends up.&lt;/p&gt;

&lt;p&gt;But as of 2026 as I write this, OCR and ICR each are the right tool for their different jobs, and there are many cases where you may need both. When you are building your self-hosted solution, make sure to review which engines are available. Tools like the Apryse Server SDK have lots of options (a few flavors of OCR and ICR), which make it easier to test multiple options out at one time.&lt;/p&gt;

&lt;p&gt;Let's take a look at a simple Apryse SDK ICR implementation.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;apryse_sdk&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;

&lt;span class="n"&gt;PDFNet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Initialize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;LicenseKey&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;PDFNet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;AddResourceSearchPath&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;../../../ICRModule/Lib/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;HandwritingICRModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;IsModuleAvailable&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Handwriting ICR module not available. Download it from https://docs.apryse.com/core/guides/info/modules#handwriting-icr-module&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;doc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;PDFDoc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;input_path&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;intake_form.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;opts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;HandwritingICROptions&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;opts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;SetPages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2-3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Run ICR and apply the result as hidden text
&lt;/span&gt;    &lt;span class="n"&gt;HandwritingICRModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;ProcessPDF&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;opts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Save&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output_path&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;intake_form_searchable.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SDFDoc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;e_linearized&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;PDFNet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Terminate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Full and beautiful sample, complete with C#, C++, Java, PHP, and all the favorites: &lt;a href="https://docs.apryse.com/core/samples/icrtest" rel="noopener noreferrer"&gt;https://docs.apryse.com/core/samples/icrtest&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;SetPages&lt;/code&gt; is doing quiet work here. Running ICR only on the pages that contain handwriting is faster. Skip sections like the signature box and use OCR on typed pages. (Although, it might be an entertaining exercise to see if what signatures come out as. lol)&lt;/p&gt;

&lt;p&gt;Developer note: Apryse ICR is a separate module you download from &lt;a href="https://docs.apryse.com/core/guides/info/modules#handwriting-icr-module" rel="noopener noreferrer"&gt;https://docs.apryse.com/core/guides/info/modules#handwriting-icr-module&lt;/a&gt;, and it maps to its own add-on to your Server SDK license. If you're using this, make sure to have it on your list of topics for sales.&lt;/p&gt;

&lt;h2&gt;
  
  
  When should you use a vision LLM instead?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; When volume is not huge and your documents are allowed to leave your network, then you can't beat the big players in the game. Vision LLMs are part of the AI ecosystem that is supposed to keep improving exponentially. Everything is subject to change quickly. Keep testing.&lt;/p&gt;

&lt;p&gt;Because if the growth and change in the industry, nothing I type can be counted on for any real length of time. But here are a few high level considerations:&lt;/p&gt;

&lt;p&gt;Reach for a &lt;strong&gt;vision LLM&lt;/strong&gt; when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Legal discovery, a few hundred pages.&lt;/strong&gt; Low volume contains the cost, and the documents are public.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Historical archives and freeform cursive, where searching beats structuring.&lt;/strong&gt; Broad training data copes with wildly unpredictable inputs, and you want a searchable layer rather than structured fields.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Reach for &lt;strong&gt;self-hosted recognition&lt;/strong&gt; when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Medical intake forms under HIPAA.&lt;/strong&gt; Data cannot leave the network. That is the whole decision.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;10,000 forms a night.&lt;/strong&gt; Per-page costs become significant enough to be worth the extra setup.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anything an auditor will review.&lt;/strong&gt; Recognition is deterministic. Same document, same output, every time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is one comparison that needs no benchmark. A recognition engine is deterministic: same document, same output, every time. A generative model is not. If you have an auditor involved with your output, making sure you know the output is repeatable is important.&lt;/p&gt;

&lt;p&gt;Healthcare keeps showing up for a reason. Handwritten intake forms, prescriptions and clinical notes are where patient data workflows pile up, and the time saving is measurable. A &lt;a href="https://link.springer.com/article/10.1186/s13054-025-05347-1" rel="noopener noreferrer"&gt;prospective multi-center study in &lt;em&gt;Critical Care&lt;/em&gt;&lt;/a&gt; found OCR cut data entry time in intensive care units by a mean of 43.9%. Reported figures for ICR in claims and e-prescription workflows run &lt;a href="https://apryse.com/blog/intelligent-character-recognition-icr" rel="noopener noreferrer"&gt;as high as a 70% reduction&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;In the end, there are a lot of numbers to throw around for these tools, and they are only as good as the time period they were printed. Technology is changing almost daily - usually improving these numbers for the better.&lt;/p&gt;

&lt;p&gt;To sum up: manual data entry stinks. Make it easier, and both time needed and accuracy are better.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you get from a scanned form to structured JSON?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; ICR gets you a searchable PDF and a JSON file full of words and coordinates. It does not get you &lt;code&gt;{"date_of_birth": "1985-03-14"}&lt;/code&gt;. Turning recognized words into named fields can take a few more steps.&lt;/p&gt;

&lt;p&gt;That position data is the part people undervalue. Text alone is a wall of words. Working in the Apryse Server SDK, here's one approach to that work:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Find the fields.&lt;/strong&gt; The Data Extraction module, run with the form engine, returns each field's position, type and a confidence score.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Convert those positions into zones.&lt;/strong&gt; Field detection measures from the top-left. PDFs measure from the bottom-left. Subtract each Y value from the page height, or ICR will confidently read the wrong part of the page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run ICR on just those zones.&lt;/strong&gt; Faster than the whole page, and it keeps the engine away from the signature box.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Apply the results to a clean template.&lt;/strong&gt; Key-value extraction does not work on scanned documents, which is the gotcha that derails people. Apply the ICR output onto a blank copy of the form instead, at zero opacity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extract the pairs, then check the doubtful ones.&lt;/strong&gt; Route anything below your confidence threshold to a person.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Roger Dunham's &lt;a href="https://apryse.com/blog/handwritten-form-data-extraction" rel="noopener noreferrer"&gt;full worked example&lt;/a&gt; walks all of it in JavaScript, including the coordinate flip and the template trick.&lt;/p&gt;

&lt;p&gt;It is a lot of steps for one form. That is the honest trade. You get a pipeline that runs entirely inside your own network, on documents that are not allowed to leave it. You pay for that in set up, but you also get to dial into exactly what you need.&lt;/p&gt;

&lt;h2&gt;
  
  
  What confidence threshold should you set?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Every recognition engine returns a per-field confidence score alongside the text. Your threshold is the line below which a field stops going straight into the database and goes to a person instead. Around 90% is a common starting point, tuned against your own documents.&lt;/p&gt;

&lt;p&gt;This is the mechanism behind human-in-the-loop, and it is worth being concrete about, because "we'll add review later" is how review never gets added.&lt;/p&gt;

&lt;p&gt;Nothing is going to be 100% accurate. Set the threshold too high and you drown your reviewers in fields that were fine. Set it too low and wrong data lands in the record silently. The only way to find your number is to run a few hundred of your own documents, compare against ground truth, and look at where the errors actually cluster.&lt;/p&gt;

&lt;p&gt;Two things that make the number easier to live with.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Set it before you build the pipeline, not after you have discovered the problem in production.&lt;/li&gt;
&lt;li&gt;Measure your reviewers too, because people are not 100% accurate either.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Remember the goal is always "better", not perfect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where should you start?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Test your worst handwriting first, decide the data residency question early because it usually settles the engine choice, and pick your confidence threshold before you build.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Start with your worst examples.&lt;/strong&gt; The messy intake forms, not the tidy ones. Good scans are easy once the bad ones work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure Character Error Rate and Word Error Rate on your own documents.&lt;/strong&gt; Doesn't matter what anyone else's tests showed, only what you see on what you're working with at the time you are starting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check whether your data is legally allowed to leave your network.&lt;/strong&gt; That usually decides between ICR and a vision model on its own.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide your confidence threshold before you build the pipeline&lt;/strong&gt;, and design the review queue at the same time. Everyone will thank you later for planning this upfront.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Redesign the form if you can.&lt;/strong&gt; Boxes beat blank lines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Get a trial key.&lt;/strong&gt; Why not start with testing Apryse? Every module is included in the 30-day trial with no page cap: &lt;a href="https://docs.apryse.com/guides/get-started" rel="noopener noreferrer"&gt;docs.apryse.com/guides/get-started&lt;/a&gt;. Sample code for ICR is at &lt;a href="https://docs.apryse.com/core/samples/icrtest" rel="noopener noreferrer"&gt;docs.apryse.com/core/samples/icrtest&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;With good recognition in place, all of those pieces of paper become sources your systems can actually use. Every visit your oldest client made is exhaustively documented and, better still, searchable. Your filing cabinets turn into a nostalgic history of useful data you can reach in seconds.&lt;/p&gt;

&lt;p&gt;Happy Coding!&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Julie Love&lt;/strong&gt; is Director of Developer Experience at Apryse. She has spent an entire career in tech spanning everything from burning CD screen savers to mastering proprietary code, sales engineering and a variety of in between. Puns always intended.&lt;/p&gt;




&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/icr-vs-ocr-differences-use-cases" rel="noopener noreferrer"&gt;ICR vs OCR: what's the difference and when does it matter?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/handwriting-ocr-guide" rel="noopener noreferrer"&gt;A simple guide to handwriting ICR&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/intelligent-document-processing-vs-traditional-ocr" rel="noopener noreferrer"&gt;Intelligent document processing vs traditional OCR&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/handwritten-form-data-extraction" rel="noopener noreferrer"&gt;From paper to patient records&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/tutorial-auto-recognize-process-form" rel="noopener noreferrer"&gt;Auto-recognize and process a form&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.linkedin.com/pulse/how-does-ai-text-recognition-work-ocr-icr-python-pipeline-julie-love-k2nbc/" rel="noopener noreferrer"&gt;How does AI text recognition work? OCR, ICR and the Python pipeline that feeds your LLM&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;alt text for top image:  handwritten to do list, complete with doodles.  Unchecked - research the status of cursive education; checked - write about OCR vs ICR and stuff; unchecked - [empty lines]; unchecked - remember what else I was supposed to do; checked - recall my 2nd grade teacher used to call me "messy bessy" because of my handwriting; checked - realize I still remember the name Elliot Lim because he sat next to me and had OCR legible handwriting; checked - remember the gross injustice of our alphabetical seating.&lt;/p&gt;

</description>
      <category>python</category>
      <category>icr</category>
    </item>
    <item>
      <title>How does AI text recognition work? OCR, ICR and the Python pipeline that feeds your LLM</title>
      <dc:creator>Julie Love</dc:creator>
      <pubDate>Wed, 09 Sep 2026 18:27:13 +0000</pubDate>
      <link>https://dev.to/apryse/how-does-ai-text-recognition-work-ocr-icr-and-the-python-pipeline-that-feeds-your-llm-2h9h</link>
      <guid>https://dev.to/apryse/how-does-ai-text-recognition-work-ocr-icr-and-the-python-pipeline-that-feeds-your-llm-2h9h</guid>
      <description>&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; AI text recognition converts images of characters into machine-readable text, and it runs before anything an LLM does. Classic OCR matches shapes against known fonts. AI-powered OCR replaces that template matching with a neural network. ICR is trained on handwriting. This is step 1 of 5, and arguably your most important step. Downstream tools can't repair what step 1 gets wrong.&lt;/p&gt;

&lt;p&gt;AI is currently center stage, in the spotlight, on every single space anyone can half-correctly call a stage right now. The most shiny of new toys even us geeks have seen in ages.&lt;/p&gt;

&lt;p&gt;So as is quite fitting with the time, the AI step in a document workflow gets all the attention. The recognition step gets glossed over, that is, until the answers start coming back wrong and it's time to dig into the why.&lt;/p&gt;

&lt;p&gt;Here is what catches people up: they blame it on hallucinations. But the recognition step is just doing its job - reporting what the characters look most like on an imperfect document.&lt;/p&gt;

&lt;h2&gt;
  
  
  How does AI text recognition actually work?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; There are three major players in the game. Classic OCR compares pixels against known letter shapes and picks the closest match. AI-powered OCR swaps that template matching for a neural network, still aimed at print but far more forgiving of bad scans. ICR is trained on real handwriting, so has more capacity to correctly analyze every new handwritten letter instance.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Classic OCR.&lt;/strong&gt; Optical Character Recognition assumes a standard font on a copy-machine-quality image. It expects the page to line up with its finite set of letter and font definitions. Think hall monitor with clipboard and a stopwatch - they get things done and keep it efficient, but you better fall in line with the exact expectations. With a clean print, it is better than 99% accurate - tough to beat for the job.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI-powered OCR.&lt;/strong&gt; Same target as classic OCR, but a neural network does the deciphering instead of a template. (think feature detection as opposed to rigid patterns) It handles the bad scans, complex layouts and unusual fonts that break a template matcher.  So this is probably that hall monitor you liked a bit better - can look the other way on small exceptions, but largely still keeps things running as expected.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ICR.&lt;/strong&gt; Intelligent Character Recognition learns character shapes from examples rather than matching templates. Lots and lots and lots of examples, which is what makes it good at educated guesses on handwriting.  In the halls, this monitor knows that no two kids are exactly the same, leaving room for self-expression in an otherwise rigid rule book.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What works best?  Only you, your examples and a bunch of testing can decide that.  So having all of these options in one place, like the Apryse Server SDK, is definitely an advantage.&lt;/p&gt;

&lt;p&gt;Apryse's default OCR module moved to deep learning neural networks in version 12.0, released July 2026. That brought roughly 16% better word recognition, much better tolerance of low-quality scans and complex layouts, no GPU requirement, and - my favorite part - a jump from 6 supported languages to more than 80.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Default OCR, AI-powered since v12.&lt;/strong&gt; For print with poor scans, complex layouts, or non-English pages. 80+ languages and the most versatile of the set. Licensed under the OCR add-on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alternative OCR.&lt;/strong&gt; For clean print, where speed and lean hardware matter most. Runs 2 to 3x faster. Also the OCR add-on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;IRIS OCR.&lt;/strong&gt; For disconnected text regions like magazine covers and CAD drawings, where it handles fragmented layouts. Separate IRIS OCR add-on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Handwriting ICR.&lt;/strong&gt; Anything written by hand, and the best handwriting results of the four. Separate ICR add-on.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Developer note: these are each separate modules you download from &lt;a href="https://docs.apryse.com/core/guides/info/modules" rel="noopener noreferrer"&gt;https://docs.apryse.com/core/guides/info/modules&lt;/a&gt;, and they map to specific add-ons for your Server SDK license. When you're ready for the call with sales, make sure to mention which items you're using.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where does recognition sit in an LLM pipeline?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Right at the front, before anything interesting happens. The pattern people now search for as OCR LLMs is really a pipeline, and recognition is step 1 of 5. Everything after it inherits whatever it produced before.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Recognition.&lt;/strong&gt; The page becomes text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chunking.&lt;/strong&gt; That text gets split into passages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embedding.&lt;/strong&gt; Passages become vectors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval.&lt;/strong&gt; A question pulls back the passages that look relevant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generation.&lt;/strong&gt; The model answers using what it retrieved.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Nothing downstream can repair a problem in step one. An embedding model has no way to know that "rnedication" was "medication" before the scan blurred, and a retrieval step cannot find a passage that was never extracted.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens when recognition output is noisy?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; It fails quietly rather than loudly, which is the worse of the two options. Five failure modes account for most of it, and not one of them throws an error.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Character-level corruption.&lt;/strong&gt; One misread letter creates a token the model has never seen, so the passage embeds strangely and stops matching the questions it should match.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phantom whitespace.&lt;/strong&gt; Words split or run together, breaking exact-match retrieval and keyword search.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Broken reading order.&lt;/strong&gt; A two-column page read straight across interleaves two unrelated arguments into nonsense.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lost tables.&lt;/strong&gt; Rows and columns flatten into a stream of digits, so a figure detaches from whatever it was measuring.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Missing content entirely.&lt;/strong&gt; Handwritten fields that classic OCR could not parse simply are not there, and nothing flags their absence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one is the quietest failure of the set. An empty field looks identical to a field the patient left blank.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you get OCR output into an LLM?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; A tool like Apryse OCR returns structured JSON or XML containing the recognized text and its position on the page. OCR text alone is a wall of words. Text plus coordinates is something a database, a rules engine or a retrieval pipeline can act on, because you can map any answer back to the exact spot it came from.&lt;/p&gt;

&lt;p&gt;There are always about as many options to accomplish the same task as stars in the sky (the reason I'd argue that coding is art), but let's take a look at the basic OCR code in Python with Apryse.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;apryse_sdk&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;

&lt;span class="n"&gt;PDFNet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Initialize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;LicenseKey&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;PDFNet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;AddResourceSearchPath&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;../../../OCRModuleWindows/Lib/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;OCRModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;IsModuleAvailable&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OCR module not available. Download it from https://dev.apryse.com/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;doc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;PDFDoc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;input_path&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scanned_invoice.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;opts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OCROptions&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;opts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;AddLang&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eng&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Get the recognition result as JSON before it goes anywhere
&lt;/span&gt;    &lt;span class="n"&gt;json&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;OCRModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;GetOCRJsonFromPDF&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;opts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Inspect, filter or correct here, then apply it back to the document
&lt;/span&gt;    &lt;span class="n"&gt;OCRModule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;ApplyOCRJsonToPDF&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Save&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output_path&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scanned_invoice_searchable.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;PDFNet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Terminate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For more detail and other languages, check out the formal &lt;a href="https://docs.apryse.com/core/samples/ocrtest" rel="noopener noreferrer"&gt;OCR Sample code&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Important not to gloss over the notes here, particularly: "Inspect, filter or correct here," &lt;code&gt;GetOCRJsonFromPDF&lt;/code&gt; gives you the recognition result &lt;em&gt;before&lt;/em&gt; it becomes a searchable PDF, which means this is the spot to add some checks. Filter low-confidence words, strip special characters or reject a page outright before any of it reaches your chunker. If you are feeding an LLM, that inspection point is the best way to ensure higher quality or know when something needs more human review.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why not just send the page to a multimodal model?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Sometimes you should. If you have neither security nor cost-saving concerns, they are great. I've seen some really impressive accuracy. Unfortunately, the great result is not always repeatable and can get expensive. When deployment control, cost and predictability can't be compromised, that's the reason to run recognition yourself.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your documents leave the building.&lt;/strong&gt; For patient records or anything with a data residency requirement, that is where the conversation ends. No cloud API runs inside an air-gapped network.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost scales with pages.&lt;/strong&gt; Fine for 50 documents. Budget line item at 50,000.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It may not give the same answer twice.&lt;/strong&gt; Recognition engines are deterministic. Generative models are not, which matters when an auditor asks why one document produced two different records.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency.&lt;/strong&gt; A round trip to someone else's server adds time you may not have.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No coordinates.&lt;/strong&gt; A model hands you text. A recognition engine hands you text plus exactly where on the page it came from.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So just to be clear, this can be a big lift to put together. ChatGPT, Gemini or Claude tools will return results with less effort. But when you need to control everything, it's time to start building.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where should you start?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Test your worst documents first and measure the error rate on your own files. Settle any data residency questions before you pick an engine, because it may decide for you.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Start with your worst examples.&lt;/strong&gt; Bad scans are the real test. The good ones are easy once worst-case is optimized.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure the error rate on your own documents.&lt;/strong&gt; Character Error Rate and Word Error Rate are the standard measures, and no vendor's published figure tells you what yours will be.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check whether your data is legally allowed to leave your network.&lt;/strong&gt; That usually settles the engine question on its own.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Match the module to the page.&lt;/strong&gt; Print, handwriting and collage layouts want different engines, and you may need several inside one project. Use the right tool for each job.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the recognition output before you read the model's answer.&lt;/strong&gt; Most pipeline bugs live upstream of the part everyone is looking at.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you've decided that you are building, definitely give Apryse a try. Get a trial key at &lt;a href="https://docs.apryse.com/guides/get-started" rel="noopener noreferrer"&gt;docs.apryse.com/guides/get-started&lt;/a&gt;. The 30-day trial includes every module mentioned here, with no page cap - so you can put your worst scan through all four before you commit.&lt;/p&gt;

&lt;p&gt;Happy Coding!&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Julie Love&lt;/strong&gt; is Director of Developer Experience at Apryse. She has spent an entire career in tech spanning everything from burning CD screen savers to mastering proprietary code, sales engineering and a variety of in betweens. Puns always intended.&lt;/p&gt;




&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/pdf-text-extraction-ocr-ai" rel="noopener noreferrer"&gt;PDF text extraction with OCR and AI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/ocr-extraction-json" rel="noopener noreferrer"&gt;OCR extraction to JSON&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/smart-data-extraction-beyond-ocr" rel="noopener noreferrer"&gt;Smart Data Extraction beyond OCR&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://apryse.com/blog/intelligent-document-processing-vs-traditional-ocr" rel="noopener noreferrer"&gt;Intelligent document processing vs traditional OCR&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.linkedin.com/pulse/icr-handwriting-recognition-python-when-do-you-need-self-hosted-love-dpadc/" rel="noopener noreferrer"&gt;ICR Handwriting Recognition in Python: when do you need self-hosted instead of a vision LLM?&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
    </item>
  </channel>
</rss>
