DEV Community

Muhammad Sharif
Muhammad Sharif

Posted on

I Built a Bank Statement Converter That Never Uploads Your Files

The problem

A bank statement shows where you live, where you shop, and how much you earn. Bookkeepers and accountants handle this kind of document for many clients, every month.

When I looked for a tool to convert a PDF statement into a spreadsheet, almost every option asked me to do the same thing: upload the file to their server, trust their privacy policy, and hope it gets deleted later.

That felt backwards. So I built something different.

The idea: no server, no upload

StatementHarbor is a bank statement converter that runs entirely in your browser. There is no upload button because there is no server to upload to.

The benefits turned out to be bigger than I expected:

  • Privacy: the file never leaves your computer, so there is nothing to leak or breach.
  • No signup: there is no account system because there is nothing to store.
  • No limits: conversions cost me nothing, so there are no page limits or paywalls.
  • Offline: once the page has loaded, it keeps working without internet.

How it works

At a high level, the pipeline looks like this:

  1. Read the PDF locally. The browser reads the file with the File API, and a JavaScript PDF library extracts the text.
  2. Get text with positions. A PDF has no real "table" concept. It is a set of text fragments with x/y coordinates, so I extract each fragment together with its position.
  3. Rebuild the rows. I group fragments with similar y-values into rows, then sort each row by x-value.
  4. Detect columns. I identify the date, description, amount, and balance columns from the header row and the alignment of values.
  5. Export. The result becomes a CSV for Excel or Google Sheets, or a QBO file that QuickBooks Online can import directly.

A simplified example of the row-grouping idea:

// Group text items into rows by their vertical position
function groupIntoRows(items, tolerance = 3) {
  const rows = [];
  for (const item of items) {
    const y = item.transform[5];
    let row = rows.find(r => Math.abs(r.y - y) <= tolerance);
    if (!row) {
      row = { y, cells: [] };
      rows.push(row);
    }
    row.cells.push(item);
  }
  return rows
    .sort((a, b) => b.y - a.y)
    .map(r => r.cells.sort((a, b) => a.transform[4] - b.transform[4]));
}
Enter fullscreen mode Exit fullscreen mode

The hard parts

Parsing bank statements is messier than it looks. A few things that gave me trouble:

  • Every bank has its own layout. Column positions, header names, and spacing all differ.
  • Multi-line descriptions. One transaction can span two or three lines, and you have to decide which lines belong to which row.
  • Date formats. A statement from a US bank and one from an Indian bank write dates differently, and some omit the year on each row.
  • Debits and credits. Some statements use separate columns, others use minus signs or "CR/DR" markers.
  • Scanned statements. These are images inside a PDF, so there is no text to extract. OCR is the next thing on my roadmap.

What I learned

  1. Client-side processing is more capable than people think. Modern browsers handle PDF parsing without any trouble.
  2. Privacy can be a feature and a cost saver. No server means no hosting bill that scales with usage, and no data to protect.
  3. Real-world data is the real test. The parser only got better once I tried it on many different statements.

Try it and tell me what breaks

It's free and needs no signup: statementharbor.com

It currently works with text-based PDF statements from banks like Chase, Wells Fargo, Bank of America, HDFC, ICICI, and SBI, with step-by-step pages for each.

If a statement doesn't parse cleanly, I would love to know which bank so I can fix it. Leave a comment below or use the contact page on the site.

What's next

  • OCR support for scanned statements
  • More bank formats

Thanks for reading! I'd love to hear how you would approach client-side PDF parsing, or whether you've run into the same privacy concerns with upload-based tools.

Top comments (2)

Collapse
 
launchgatecheck profile image
Launch Gate •

For the position-based row grouping, I'd add a synthetic two-page statement where both pages have transactions at the same y-coordinate, with a description continuing onto the next page. Keep page identity separate from y-position, then reconcile opening balance + credits - debits against closing balance before export. An unmatched continuation or amount should produce a visible review state rather than a tidy but wrong CSV/QBO. Do you surface unassigned text fragments and a reconciliation difference?

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to