Quick Answer: AI text recognition converts images of characters into machine-readable text, and it runs before anything an LLM does. Classic OCR matches shapes against known fonts. AI-powered OCR replaces that template matching with a neural network. ICR is trained on handwriting. This is step 1 of 5, and arguably your most important step. Downstream tools can't repair what step 1 gets wrong.
AI is currently center stage, in the spotlight, on every single space anyone can half-correctly call a stage right now. The most shiny of new toys even us geeks have seen in ages.
So as is quite fitting with the time, the AI step in a document workflow gets all the attention. The recognition step gets glossed over, that is, until the answers start coming back wrong and it's time to dig into the why.
Here is what catches people up: they blame it on hallucinations. But the recognition step is just doing its job - reporting what the characters look most like on an imperfect document.
How does AI text recognition actually work?
Quick Answer: There are three major players in the game. Classic OCR compares pixels against known letter shapes and picks the closest match. AI-powered OCR swaps that template matching for a neural network, still aimed at print but far more forgiving of bad scans. ICR is trained on real handwriting, so has more capacity to correctly analyze every new handwritten letter instance.
- Classic OCR. Optical Character Recognition assumes a standard font on a copy-machine-quality image. It expects the page to line up with its finite set of letter and font definitions. Think hall monitor with clipboard and a stopwatch - they get things done and keep it efficient, but you better fall in line with the exact expectations. With a clean print, it is better than 99% accurate - tough to beat for the job.
- AI-powered OCR. Same target as classic OCR, but a neural network does the deciphering instead of a template. (think feature detection as opposed to rigid patterns) It handles the bad scans, complex layouts and unusual fonts that break a template matcher. So this is probably that hall monitor you liked a bit better - can look the other way on small exceptions, but largely still keeps things running as expected.
- ICR. Intelligent Character Recognition learns character shapes from examples rather than matching templates. Lots and lots and lots of examples, which is what makes it good at educated guesses on handwriting. In the halls, this monitor knows that no two kids are exactly the same, leaving room for self-expression in an otherwise rigid rule book.
What works best? Only you, your examples and a bunch of testing can decide that. So having all of these options in one place, like the Apryse Server SDK, is definitely an advantage.
Apryse's default OCR module moved to deep learning neural networks in version 12.0, released July 2026. That brought roughly 16% better word recognition, much better tolerance of low-quality scans and complex layouts, no GPU requirement, and - my favorite part - a jump from 6 supported languages to more than 80.
- Default OCR, AI-powered since v12. For print with poor scans, complex layouts, or non-English pages. 80+ languages and the most versatile of the set. Licensed under the OCR add-on.
- Alternative OCR. For clean print, where speed and lean hardware matter most. Runs 2 to 3x faster. Also the OCR add-on.
- IRIS OCR. For disconnected text regions like magazine covers and CAD drawings, where it handles fragmented layouts. Separate IRIS OCR add-on.
- Handwriting ICR. Anything written by hand, and the best handwriting results of the four. Separate ICR add-on.
Developer note: these are each separate modules you download from https://docs.apryse.com/core/guides/info/modules, and they map to specific add-ons for your Server SDK license. When you're ready for the call with sales, make sure to mention which items you're using.
Where does recognition sit in an LLM pipeline?
Quick Answer: Right at the front, before anything interesting happens. The pattern people now search for as OCR LLMs is really a pipeline, and recognition is step 1 of 5. Everything after it inherits whatever it produced before.
- Recognition. The page becomes text.
- Chunking. That text gets split into passages.
- Embedding. Passages become vectors.
- Retrieval. A question pulls back the passages that look relevant.
- Generation. The model answers using what it retrieved.
Nothing downstream can repair a problem in step one. An embedding model has no way to know that "rnedication" was "medication" before the scan blurred, and a retrieval step cannot find a passage that was never extracted.
What happens when recognition output is noisy?
Quick Answer: It fails quietly rather than loudly, which is the worse of the two options. Five failure modes account for most of it, and not one of them throws an error.
- Character-level corruption. One misread letter creates a token the model has never seen, so the passage embeds strangely and stops matching the questions it should match.
- Phantom whitespace. Words split or run together, breaking exact-match retrieval and keyword search.
- Broken reading order. A two-column page read straight across interleaves two unrelated arguments into nonsense.
- Lost tables. Rows and columns flatten into a stream of digits, so a figure detaches from whatever it was measuring.
- Missing content entirely. Handwritten fields that classic OCR could not parse simply are not there, and nothing flags their absence.
That last one is the quietest failure of the set. An empty field looks identical to a field the patient left blank.
How do you get OCR output into an LLM?
Quick Answer: A tool like Apryse OCR returns structured JSON or XML containing the recognized text and its position on the page. OCR text alone is a wall of words. Text plus coordinates is something a database, a rules engine or a retrieval pipeline can act on, because you can map any answer back to the exact spot it came from.
There are always about as many options to accomplish the same task as stars in the sky (the reason I'd argue that coding is art), but let's take a look at the basic OCR code in Python with Apryse.
from apryse_sdk import *
PDFNet.Initialize(LicenseKey)
PDFNet.AddResourceSearchPath("../../../OCRModuleWindows/Lib/")
if not OCRModule.IsModuleAvailable():
print("OCR module not available. Download it from https://dev.apryse.com/")
else:
doc = PDFDoc(input_path + "scanned_invoice.pdf")
opts = OCROptions()
opts.AddLang("eng")
# Get the recognition result as JSON before it goes anywhere
json = OCRModule.GetOCRJsonFromPDF(doc, opts)
# Inspect, filter or correct here, then apply it back to the document
OCRModule.ApplyOCRJsonToPDF(doc, json)
doc.Save(output_path + "scanned_invoice_searchable.pdf", 0)
PDFNet.Terminate()
For more detail and other languages, check out the formal OCR Sample code.
Important not to gloss over the notes here, particularly: "Inspect, filter or correct here," GetOCRJsonFromPDF gives you the recognition result before it becomes a searchable PDF, which means this is the spot to add some checks. Filter low-confidence words, strip special characters or reject a page outright before any of it reaches your chunker. If you are feeding an LLM, that inspection point is the best way to ensure higher quality or know when something needs more human review.
Why not just send the page to a multimodal model?
Quick Answer: Sometimes you should. If you have neither security nor cost-saving concerns, they are great. I've seen some really impressive accuracy. Unfortunately, the great result is not always repeatable and can get expensive. When deployment control, cost and predictability can't be compromised, that's the reason to run recognition yourself.
- Your documents leave the building. For patient records or anything with a data residency requirement, that is where the conversation ends. No cloud API runs inside an air-gapped network.
- Cost scales with pages. Fine for 50 documents. Budget line item at 50,000.
- It may not give the same answer twice. Recognition engines are deterministic. Generative models are not, which matters when an auditor asks why one document produced two different records.
- Latency. A round trip to someone else's server adds time you may not have.
- No coordinates. A model hands you text. A recognition engine hands you text plus exactly where on the page it came from.
So just to be clear, this can be a big lift to put together. ChatGPT, Gemini or Claude tools will return results with less effort. But when you need to control everything, it's time to start building.
Where should you start?
Quick Answer: Test your worst documents first and measure the error rate on your own files. Settle any data residency questions before you pick an engine, because it may decide for you.
- Start with your worst examples. Bad scans are the real test. The good ones are easy once worst-case is optimized.
- Measure the error rate on your own documents. Character Error Rate and Word Error Rate are the standard measures, and no vendor's published figure tells you what yours will be.
- Check whether your data is legally allowed to leave your network. That usually settles the engine question on its own.
- Match the module to the page. Print, handwriting and collage layouts want different engines, and you may need several inside one project. Use the right tool for each job.
- Read the recognition output before you read the model's answer. Most pipeline bugs live upstream of the part everyone is looking at.
If you've decided that you are building, definitely give Apryse a try. Get a trial key at docs.apryse.com/guides/get-started. The 30-day trial includes every module mentioned here, with no page cap - so you can put your worst scan through all four before you commit.
Happy Coding!
Julie Love is Director of Developer Experience at Apryse. She has spent an entire career in tech spanning everything from burning CD screen savers to mastering proprietary code, sales engineering and a variety of in betweens. Puns always intended.
Top comments (0)