Quick Answer: Self-hosted ICR is a recognition engine trained on handwriting that runs inside your own network and can be set to return text plus coordinates. A vision LLM reads handwriting using an enormous training set and usually returns a block of text from someone else's server. Pay close attention to your needs around deployment, cost and audit-ability to make the best choices for you.
I've heard learning cursive in school is a thing of the past, much like my spiral notebook and my handwritten to-do lists.
Yes, there are usually doodles in my notebooks as well. It is great for those of us with visual and spatial memories who need to put ideas together on paper before they make sense anywhere else. (At least that is my justification and I am sticking to it.)
Luckily, kids these days do not need to learn cursive while we have so many cool digital toys now.
Let's break it down. (Due to the limitations of the written word, you will need to mentally cue a great late-80s hip hop track and a dance of your choice.)
Why is handwriting so much harder to read than printed text?
Quick Answer: Printed text follows solid, definable, repeatable rules. Handwriting breaks five of them at once, which is why handwriting support is not just a quick add-on to a printed-text engine. It has to be a separate model trained on separate data.
- Everyone writes differently. This may be one of the world's largest understatements. Fonts are finite. Handwriting variation is not. (in the uncountable infinity category, for my fellow math nerds.) Even the same person writes the same letter differently depending on the number of lattes consumed and how annoying the traffic was on their way to fill out the form.
- Cursive has no gaps. No need to mind the gap here, because there are none. No reliable boundary between characters means nowhere obvious to cut.
- Words wander. Print sits on the line. Handwriting drifts, slants, and takes a little trip to any nearby region that looks fun or available.
- Training data is scarce. Labeled handwriting collections are far smaller and rarer than printed text. Side quest, handwriting sample set: IAM Handwriting Database
- Every source looks different. A doctor's note has nothing in common with an 18th-century census record. And let's be real, you thought teaching a computer to beat a human at Go was hard? There is no quantity of doctors' notes large enough to solve that great mystery of life.
Then there are forms, which pile all of the challenges into one place. Typed labels that have to relate to handwritten answers, deciphered separately and reassembled correctly.
One practical note before you start building: If you have any influence over the forms themselves, use it. Grid-style boxes force people to space out their letters and define a usable area for each field. If that's too limiting, keep your field boxes black (non-scan colors like red, pink or cyan can drop out completely in the pre-process), and your field labels above or to the left of the fields (not inside or touching the boundaries). Blank lines invite the sprawl that breaks everything downstream. Design with scaffolding wherever you can.
What is ICR, and how is it different from OCR?
Quick Answer: Intelligent Character Recognition (ICR) is a neural network trained on handwriting examples rather than font definitions, so it can read writing it has never seen before. Optical Character Recognition (OCR) is based on matching the shape of a letter and expected characteristics - based on printed fonts.
Also of note, this is ever evolving. AI-powered OCR can sometimes pick up handwriting pretty well. As AI continues its expected exponential growth, will be interesting to see where this all ends up.
But as of 2026 as I write this, OCR and ICR each are the right tool for their different jobs, and there are many cases where you may need both. When you are building your self-hosted solution, make sure to review which engines are available. Tools like the Apryse Server SDK have lots of options (a few flavors of OCR and ICR), which make it easier to test multiple options out at one time.
Let's take a look at a simple Apryse SDK ICR implementation.
from apryse_sdk import *
PDFNet.Initialize(LicenseKey)
PDFNet.AddResourceSearchPath("../../../ICRModule/Lib/")
if not HandwritingICRModule.IsModuleAvailable():
print("Handwriting ICR module not available. Download it from https://docs.apryse.com/core/guides/info/modules#handwriting-icr-module")
else:
doc = PDFDoc(input_path + "intake_form.pdf")
opts = HandwritingICROptions()
opts.SetPages("2-3")
# Run ICR and apply the result as hidden text
HandwritingICRModule.ProcessPDF(doc, opts)
doc.Save(output_path + "intake_form_searchable.pdf", SDFDoc.e_linearized)
doc.Close()
PDFNet.Terminate()
Full and beautiful sample, complete with C#, C++, Java, PHP, and all the favorites: https://docs.apryse.com/core/samples/icrtest
SetPages is doing quiet work here. Running ICR only on the pages that contain handwriting is faster. Skip sections like the signature box and use OCR on typed pages. (Although, it might be an entertaining exercise to see if what signatures come out as. lol)
Developer note: Apryse ICR is a separate module you download from https://docs.apryse.com/core/guides/info/modules#handwriting-icr-module, and it maps to its own add-on to your Server SDK license. If you're using this, make sure to have it on your list of topics for sales.
When should you use a vision LLM instead?
Quick Answer: When volume is not huge and your documents are allowed to leave your network, then you can't beat the big players in the game. Vision LLMs are part of the AI ecosystem that is supposed to keep improving exponentially. Everything is subject to change quickly. Keep testing.
Because if the growth and change in the industry, nothing I type can be counted on for any real length of time. But here are a few high level considerations:
Reach for a vision LLM when:
- Legal discovery, a few hundred pages. Low volume contains the cost, and the documents are public.
- Historical archives and freeform cursive, where searching beats structuring. Broad training data copes with wildly unpredictable inputs, and you want a searchable layer rather than structured fields.
Reach for self-hosted recognition when:
- Medical intake forms under HIPAA. Data cannot leave the network. That is the whole decision.
- 10,000 forms a night. Per-page costs become significant enough to be worth the extra setup.
- Anything an auditor will review. Recognition is deterministic. Same document, same output, every time.
There is one comparison that needs no benchmark. A recognition engine is deterministic: same document, same output, every time. A generative model is not. If you have an auditor involved with your output, making sure you know the output is repeatable is important.
Healthcare keeps showing up for a reason. Handwritten intake forms, prescriptions and clinical notes are where patient data workflows pile up, and the time saving is measurable. A prospective multi-center study in Critical Care found OCR cut data entry time in intensive care units by a mean of 43.9%. Reported figures for ICR in claims and e-prescription workflows run as high as a 70% reduction.
In the end, there are a lot of numbers to throw around for these tools, and they are only as good as the time period they were printed. Technology is changing almost daily - usually improving these numbers for the better.
To sum up: manual data entry stinks. Make it easier, and both time needed and accuracy are better.
How do you get from a scanned form to structured JSON?
Quick Answer: ICR gets you a searchable PDF and a JSON file full of words and coordinates. It does not get you {"date_of_birth": "1985-03-14"}. Turning recognized words into named fields can take a few more steps.
That position data is the part people undervalue. Text alone is a wall of words. Working in the Apryse Server SDK, here's one approach to that work:
- Find the fields. The Data Extraction module, run with the form engine, returns each field's position, type and a confidence score.
- Convert those positions into zones. Field detection measures from the top-left. PDFs measure from the bottom-left. Subtract each Y value from the page height, or ICR will confidently read the wrong part of the page.
- Run ICR on just those zones. Faster than the whole page, and it keeps the engine away from the signature box.
- Apply the results to a clean template. Key-value extraction does not work on scanned documents, which is the gotcha that derails people. Apply the ICR output onto a blank copy of the form instead, at zero opacity.
- Extract the pairs, then check the doubtful ones. Route anything below your confidence threshold to a person.
Roger Dunham's full worked example walks all of it in JavaScript, including the coordinate flip and the template trick.
It is a lot of steps for one form. That is the honest trade. You get a pipeline that runs entirely inside your own network, on documents that are not allowed to leave it. You pay for that in set up, but you also get to dial into exactly what you need.
What confidence threshold should you set?
Quick Answer: Every recognition engine returns a per-field confidence score alongside the text. Your threshold is the line below which a field stops going straight into the database and goes to a person instead. Around 90% is a common starting point, tuned against your own documents.
This is the mechanism behind human-in-the-loop, and it is worth being concrete about, because "we'll add review later" is how review never gets added.
Nothing is going to be 100% accurate. Set the threshold too high and you drown your reviewers in fields that were fine. Set it too low and wrong data lands in the record silently. The only way to find your number is to run a few hundred of your own documents, compare against ground truth, and look at where the errors actually cluster.
Two things that make the number easier to live with.
- Set it before you build the pipeline, not after you have discovered the problem in production.
- Measure your reviewers too, because people are not 100% accurate either.
Remember the goal is always "better", not perfect.
Where should you start?
Quick Answer: Test your worst handwriting first, decide the data residency question early because it usually settles the engine choice, and pick your confidence threshold before you build.
- Start with your worst examples. The messy intake forms, not the tidy ones. Good scans are easy once the bad ones work.
- Measure Character Error Rate and Word Error Rate on your own documents. Doesn't matter what anyone else's tests showed, only what you see on what you're working with at the time you are starting.
- Check whether your data is legally allowed to leave your network. That usually decides between ICR and a vision model on its own.
- Decide your confidence threshold before you build the pipeline, and design the review queue at the same time. Everyone will thank you later for planning this upfront.
- Redesign the form if you can. Boxes beat blank lines.
- Get a trial key. Why not start with testing Apryse? Every module is included in the 30-day trial with no page cap: docs.apryse.com/guides/get-started. Sample code for ICR is at docs.apryse.com/core/samples/icrtest.
With good recognition in place, all of those pieces of paper become sources your systems can actually use. Every visit your oldest client made is exhaustively documented and, better still, searchable. Your filing cabinets turn into a nostalgic history of useful data you can reach in seconds.
Happy Coding!
Julie Love is Director of Developer Experience at Apryse. She has spent an entire career in tech spanning everything from burning CD screen savers to mastering proprietary code, sales engineering and a variety of in between. Puns always intended.
Further reading
- ICR vs OCR: what's the difference and when does it matter?
- A simple guide to handwriting ICR
- Intelligent document processing vs traditional OCR
- From paper to patient records
- Auto-recognize and process a form
- How does AI text recognition work? OCR, ICR and the Python pipeline that feeds your LLM
alt text for top image: handwritten to do list, complete with doodles. Unchecked - research the status of cursive education; checked - write about OCR vs ICR and stuff; unchecked - [empty lines]; unchecked - remember what else I was supposed to do; checked - recall my 2nd grade teacher used to call me "messy bessy" because of my handwriting; checked - realize I still remember the name Elliot Lim because he sat next to me and had OCR legible handwriting; checked - remember the gross injustice of our alphabetical seating.
Top comments (0)