I Built a Browser-Based PDF Bank Statement Parser
I recently built and launched SheetParse, a browser-based tool that converts bank statement PDFs into spreadsheet-ready data.
The idea started with a simple problem: PDF is a great format for reading a bank statement, but it's a pretty terrible format when you actually want to work with the transaction data.
If you need to analyze, filter, reconcile, or move the transactions into Excel, you quickly end up dealing with manual copy-pasting or unreliable PDF converters.
So I decided to build my own solution.
The interesting part: bank statements aren't standardized
At first, I thought the problem would basically be:
PDF → Extract text → Parse transactions → Excel
It turned out to be much messier.
Different banks use completely different layouts.
Some have separate debit and credit columns. Others combine them. Transaction descriptions can span multiple lines. Headers and footers can appear between transactions. Some PDFs contain actual text, while others are essentially scanned images.
So the actual pipeline looks more like:
PDF
↓
Text extraction / OCR
↓
Statement detection
↓
Bank-specific parsing
↓
Transaction normalization
↓
Validation
↓
CSV / XLSX
I decided to process everything locally
One of the main decisions I made was to keep document processing in the browser whenever possible.
Bank statements contain sensitive financial information, so I didn't particularly like the idea of requiring users to upload their documents to a random server just to convert them.
With SheetParse, the document can be processed locally in the browser.
That also means I don't need a large backend infrastructure just to process someone's PDF.
Bank-specific parsers
Another thing I learned is that trying to create one universal parser isn't necessarily the best approach.
Instead, SheetParse has parsers designed around different statement formats, along with a generic parser for unsupported formats.
The goal is to normalize everything into a consistent structure such as:
Date
Description
Debit
Credit
Balance
Once the data is normalized, generating CSV or XLSX becomes much easier.
OCR is a different problem
Text-based PDFs and scanned PDFs are very different.
If a PDF contains actual text, extraction is relatively straightforward.
But if every page is essentially an image, there isn't any transaction text to extract.
That's where OCR becomes useful.
SheetParse uses OCR as a fallback for these types of documents.
The tradeoff is that OCR is slower and can introduce additional ambiguity, especially with things like numbers, dates and transaction descriptions.
Extraction isn't enough
One of the biggest things I learned while building this is that extracting rows isn't the same as extracting correct data.
For financial documents, an incorrect transaction can be worse than a missing transaction because it can silently affect someone's calculations.
So validation becomes important.
Things like:
- duplicate transactions
- broken dates
- incorrect debit/credit values
- transactions split across lines
- repeated headers
- page boundaries
- running balances
all need to be considered.
The parser needs to understand the structure of the statement rather than simply grabbing everything that looks like a row.
Column Studio
I also wanted SheetParse to do more than simply extract the transactions and immediately export them.
That's why I built Column Studio in SheetParse.
After the data has been extracted, Column Studio lets you work with the resulting columns before exporting the final spreadsheet.
Instead of having to take the raw output into another tool just to make small changes, you can prepare the dataset within SheetParse itself.
The idea is to make the workflow closer to:
PDF
↓
Extract
↓
Clean / organize columns
↓
Review
↓
Export
This became particularly useful because different people want different spreadsheet structures.
Someone might only need:
Date | Description | Amount
while another workflow might need:
Date | Description | Debit | Credit | Balance
Rather than forcing one output format on everyone, Column Studio gives the user more control over the extracted dataset.
Why build it in the browser?
There were a few reasons.
Privacy
The document doesn't have to be uploaded to a backend just for parsing.
Infrastructure
There isn't a document-processing server running for every conversion.
Simplicity
The application can remain relatively lightweight.
Cost
Client-side processing reduces the infrastructure required to operate the service.
There are obviously tradeoffs too. Large PDFs and OCR can consume significant CPU and memory in the browser, so performance is something I'm continuing to work on.
What I learned
The biggest lesson wasn't really about PDF parsing.
It was that real-world documents are messy.
A parser can work perfectly against ten sample PDFs and then encounter a completely different layout in the eleventh.
That's why supporting more banks isn't simply a matter of adding another parser. Each format needs representative documents, edge cases and regression tests.
I'm now much more interested in the reliability and validation side of document processing than I was when I started the project.
SheetParse is live
The first public version is now available:
It currently supports multiple bank statement formats, a generic parser, and OCR fallback for scanned PDFs.
It's still an early project, but it's something I've wanted to build for a while, and it's finally out in the wild.
If you've worked with PDF parsing, OCR, financial data extraction, or document processing, I'd be interested to hear how you've approached similar problems.
What are some of the weirdest PDF formats you've had to deal with?
Top comments (2)
Please feel free to let me know your feedbacks or improvement ideas about sheetparse so I can make it more better!
The validation stage is the part I'd preserve through Column Studio and export. One useful fixture is two legitimate transactions with the same date, description and amount, plus a repeated page-boundary row. A value-based dedupe can wrongly remove a real payment, while keeping every row can double-count the overlap. Keeping source page/row evidence lets the review distinguish those cases. I'd also carry an unresolved reconciliation flag into the exported data instead of letting the final CSV look fully checked after someone hides the Balance column. Does Column Studio retain those source and validation fields when columns are reorganized?