DEV Community

Feroz Sheikh
Feroz Sheikh

Posted on

Parsing bank-statement PDFs in the browser: text layer, tables, and a balance check

Generic "PDF to Excel" converters are great until you feed them a bank statement. Columns drift, page headers repeat mid-table, wrapped descriptions become extra rows, and the closing balance no longer matches the opening balance plus the sum of amounts. That is not a minor formatting bug — it is a wrong ledger.

I spent a while on this while building a free, browser-only converter for text-layer statements. Here is the practical checklist I wish I had on day one.

1. Confirm there is a real text layer

Open the PDF and try to select a single transaction line. If you cannot select text, or you get mojibake, you have a scan (or broken font encoding). Ordinary converters will invent numbers. You need OCR in the statement language first, and you should budget time to fix rows by hand.

If selection works, you have a text PDF. Prefer a table-aware extractor over a full-page "PDF to Word" pass:

  • Excel (Microsoft 365 on Windows): Data → Get Data → From File → From PDF, then pick the table on each page
  • Tabula (open source): draw a box around the transaction table and export CSV

2. Prefer the bank export when it exists

Most online banking apps can export CSV or Excel for a date range. That path beats every PDF pipeline. Use the PDF route only when the bank will not give you a spreadsheet (or you only have the PDF the branch emailed).

3. Reconcile row by row, not "looks about right"

After extraction, start from the opening balance and walk each row:

previous_balance ± amount  ==  new_balance
Enter fullscreen mode Exit fullscreen mode

Any failure is a split cell, a merged cell, a repeated page header, or a description that wrapped onto a second line. Those four failure modes cover almost every "converted mess" I have seen.

A small implementation tip in JavaScript: keep amounts as integers of the minor unit (pence/cents) while you walk the ledger, and only format for display at the end. Floating-point drift will otherwise create false mismatches on long statements.

4. Keep the file on the device

Bank statements are high-sensitivity documents. Uploading them to a random converter is a privacy decision, not just a convenience one. A client-side parser (pdf.js + your own column heuristics) never sends the bytes to a server. Trade-off: large scanned PDFs OCR slowly in the browser, and you will not magically fix a true scan without an OCR engine that supports the language.

Disclosure

I built a free browser converter that runs this balance check locally on text PDFs (not scans): opentoolsuite.com/guides/convert-bank-statement-pdf-to-excel. Use it, or rebuild the same pattern with Tabula + a spreadsheet — the reconciliation step is the part that actually matters.

Top comments (0)