DEV Community

Cover image for My OCR Pipeline Asked for 44 GB of RAM to Read Three Invoice Columns. The Fix Wasn't a Better Model.
Evgeny Khramov
Evgeny Khramov

Posted on Originally published at hram.github.io AI-assisted

My OCR Pipeline Asked for 44 GB of RAM to Read Three Invoice Columns. The Fix Wasn't a Better Model.

The task was narrow. Take a phone photo of an UPD (a standard Russian invoice form), read the VAT rate from column 7, group the rows by rate, sum columns 8 and 9, and check the sums against the document total.

photo → find three columns → read the numbers → group by VAT rate → check the sums
Enter fullscreen mode Exit fullscreen mode

My wife is an accountant and does this by hand, so I started building a tool for it with a coding agent. I assumed OCR would be the hard part. It turned out to be one of the easiest.

The first version: give the whole page to PaddleOCR

My own spec asked for the obvious architecture: OpenCV preprocessing, then PaddleOCR's table recognition as the main path, with classical geometry only as a fallback.

The agent built it: page detection, perspective warp, deskew, CLAHE, then two full-page passes — TableRecognitionPipelineV2 for structure and a separate PaddleOCR(lang="ru") for text.

The first real phone photo, about 5 MB, made PaddleOCR try to allocate roughly 44 GB and the request died with ResourceExhaustedError.

The obvious hypothesis was "the photo is too big", so the agent capped the long side at 2600 px. A large synthetic image now completed — in about 44 seconds, peaking near 10 GB RSS. It looked like a fix.

Then the same class of failure came back on another real document.

The image size was a symptom

From that point the agent ran every experiment under systemd-run with a hard MemoryMax and swap disabled, so a bad hypothesis would kill one process instead of freezing the laptop.

The first reproducible run under a 14 GB limit:

value
input image 1462×2600
model loading ~85 s
RSS after loading ~1.97 GB
first full-page inference over 14 GB, process killed

Even the downscaled image didn't fit. The cause was the composition of the pipeline: PaddleOCR(ru) ran its own text detector over the whole page, and TableRecognitionPipelineV2 carried its own models for layout, OCR, and table structure. The page went through several heavy models, with part of the work duplicated.

PaddleOCR wasn't the problem. Feeding it a whole page to get three columns was.

The new scheme:

ORIGINAL → small working copy → classical CV: geometry → coordinates
ORIGINAL → small crops at those coordinates → text recognition model only
Enter fullscreen mode Exit fullscreen mode

Working image for geometry, original for reading. With recognition running only on crops, peak memory in a control run was about 627 MB.

But that only works if the geometry is right. And the standard way of getting it right made things worse.

Aligning by the sheet edge was worse than not aligning at all

The first geometric version did what a document scanner does: find the sheet contour, warp it to a rectangle. But the sheet often didn't fit in the frame, and the paper edge isn't necessarily parallel to the printed table.

Morphology then found zero or one long horizontal line instead of about thirty. Rotating by the angle of the horizontals fixed them and tilted the verticals by about 0.31° — so the distortion was shear or perspective, not rotation. A shear correction helped.

At that point I asked the agent why step two was fixing a tilt that step one should have removed. It proposed moving the tilt measurement into step one — tidier, but the reference object was still the sheet of paper.

Years ago I wrote a Sudoku recognizer, and I told the agent how it worked: look for near-vertical and near-horizontal lines on the uncropped frame within a range of angles, find their intersections, and derive the transform from those.

The useful part wasn't the code, it was the choice of reference. If I need the table's geometry, the table is what should be aligned.

source frame → LSD line segments → horizontal/vertical lines
→ intersections = table nodes → homography (RANSAC) → one warp of the original
Enter fullscreen mode Exit fullscreen mode

A Sudoku grid is regular, so every node's ideal position is known. An invoice has columns and rows of different sizes, so the agent did a rough transform from the outer geometry and refined it over many nodes with outlier rejection. Debugging took several iterations: LSD found both edges of a thick line, and digits in a narrow column produced false segments.

Then all variants were measured with one metric:

Variant Tilt, horizontals / verticals Mean node residual
No alignment 0.032° / 0.053° 0.76 px
By sheet contour 0.374° / 0.010° 1.97 px
Contour + shear 0.105° / 0.004° 0.98 px
By table grid 0.010° / 0.002° 0.26 px

Look at the second row. Aligning by the sheet contour was quantitatively worse than doing nothing: the tilt we kept correcting was partly introduced by our own preprocessing. The grid-based method cost about 0.8 s more per page at that stage.

A one-pixel error on the working copy clips a digit on the original

With the grid found on the reduced copy, I scaled the coordinates back and cut cells from the original. Some money values lost their last digit: one pixel on the working copy is several on the original, and morphology widens the table lines further.

The fix was a second coarse-to-fine level. The working-image grid gives an approximate boundary; then, in the original, the algorithm searches a ±10 px window for the actual printed line and crops from there.

The control metric was edge_ink — crops with ink touching the edge. It went from 67 of 84 to 3 of 84, and those three were specks and pen marks, not clipped digits.

The biggest accuracy gain was a crop, not a model

With detection and table recognition gone, one small text recognition model remained. Two findings from comparing inputs:

  • Binarization helped geometry and hurt recognition. The model read grayscale better.
  • A strip of several columns scored slightly higher than separate cells at first, but I kept separate cells: one lost comma in a strip breaks the parsing of its neighbors, while a cell keeps its own bbox, crop, and result as evidence.

The remaining errors clustered in double-height rows. There the value occupies a small part of a tall cell, and the model resizes the whole crop to a fixed 48 px height, so the digits shrink.

Adding a tight crop around the connected ink components, on a hand-labeled golden document:

separate cells: 76 → 82 out of 83
strip:          79 → 83 out of 83
Enter fullscreen mode Exit fullscreen mode

Enabling MKL-DNN for the recognition model, on the same tight crops, cut a single call from about 155 to 44 ms for cells and from 366 to 93 ms for the strip.

The tight crop brought its own errors — dots, pen marks, and grid-line remnants stretching the box — and the component rules had to be tightened.

Constrain the decoder by what the field can contain

One error remained: a smudged 10% read as 1Q%. A Q → 0 rule would encode one photo's defect. But I know the field's semantics: a VAT rate is digits, a percent sign, or the text «без НДС» ("no VAT"); a money amount is digits and separators.

The restriction went into CTC decoding as a per-field character whitelist. That gave 83 of 83 on the golden document.

It can't be global. The column-number row contains labels like 1а and 10а, and a numeric whitelist corrupts them.

Perfect characters, a row that never existed

A second form, TORG-12 (another Russian delivery note), arrived rotated and on curved paper.

Orientation: PaddleOCR's classifier on the central crop got 1 of 24 artificial rotations wrong, because the page center held signatures instead of text. Voting over a 3×3 grid of regions, skipping empty ones, gave 24 of 24. Document type came from the already-read column-number row, with UNKNOWN returned when features conflict.

The dangerous problem was structural. Global row detection found 16 rows instead of 18: the curvature merged two pairs of neighbors. Since columns are read independently, the pipeline could take an amount from one physical row and the VAT from the next and emit a record that was never on the paper — with every character read correctly.

The fix was to stop solving the whole page again: search for row boundaries only inside the narrow strip of needed columns, where curvature barely shows. That found all 18 rows, with no mixed values on this document.

Speed: a dictionary beat a heavier model

By now a document took 4.4–4.8 s, 75–80% of it in the recognition network. The agent compared seven recognizers (several PaddleOCR generations and Tesseract) in two modes on two documents.

  • PP-OCRv6_tiny was fastest, but confused 6 and 8 and broke the column-number row.
  • en_PP-OCRv3_mobile_rec was about 1.8–2.6× faster than the original model, with one systematic error: the smudged 10% became 19%.

A heavier model was one option. But the set of legal rates is finite: 0%, 5%, 7%, 10%, 20%, 22%, «без НДС». So for this field the decoder computes the CTC probability of each legal entry and picks the most probable. That is more than a character whitelist: the decoder knows the full dictionary of values. On the smudge, 10% wins.

Result with en_PP-OCRv3_mobile_rec as the default: 135 of 135 labeled values on the two-document collection, and 4.84 → 2.81 s on the TORG-12, 6.06 → 3.19 s on the UPD.

That is two documents, not an accuracy claim. And the fastest hybrid, with the tiny model, was rejected because it lost one real value.

What the agent did and what I did

The first full-page pipeline matched my own spec — I asked for it. After the crashes we switched to an experiment log: hypothesis before a change, measurements and artifacts after. The agent ran variants under the memory limit, measured time and RSS, saved crops and overlays, and moved the winners into the main pipeline.

The agent found the real cause of the OOM, after the resize fix failed. I changed the direction where the model of the task had to change: the Sudoku episode replaced the paper edge with the table grid as the reference.

Where this stands

It started as image → OCR → data. It is now:

image → geometry → physical rows and cells → OCR
→ domain constraints → structured data → validation
Enter fullscreen mode Exit fullscreen mode

The neural network sits in the narrowest spot, reading a small fragment that is already located and cropped. Almost none of the big gains came from a bigger model: memory dropped because the network stopped seeing the page, geometry improved because the reference became the table, accuracy improved because the crop got tighter and the decoder learned the field's semantics.

PaperLedger is not a finished product. My wife doesn't use it in her daily work; I run experiments on her real scenarios and show her the results, and real use would need her company's approval because of the data in the documents.

The full case study is on my site: https://hram.github.io/en/articles/paper-ledger/

Top comments (0)