DEV Community

Cover image for Why PDF Tables Break When You Extract Them — and How to Handle It
BLOODY GAMING
BLOODY GAMING

Posted on

Why PDF Tables Break When You Extract Them — and How to Handle It

Working with PDFs sounds simple until you actually need the data inside them.

A PDF might contain a perfectly clean table with hundreds of rows. You can open it, read everything, and even search through the document.

Then you try to extract that table into Excel.

Suddenly the columns don't line up.

Some rows are split into two.

Numbers end up in the wrong cells.

And sometimes the output is completely empty.

I've run into this problem enough times to realize that the issue usually isn't just the extraction tool. The bigger problem is how PDFs store information in the first place.

If you're building a document-processing workflow, understanding this makes PDF extraction much easier to reason about.

A PDF Isn't Really a Spreadsheet

This is the first thing that confused me when I started working with PDF data.

When we see this:

Product        Quantity        Price
Keyboard       20              $500
Mouse          35              $700
Monitor        10              $2,000
Enter fullscreen mode Exit fullscreen mode

we naturally think:

row → column → cell
Enter fullscreen mode Exit fullscreen mode

That's how a spreadsheet works.

A PDF doesn't necessarily work that way.

A PDF can essentially say:

Put this piece of text at this position on the page.

Then put another piece of text somewhere else.

Visually, those pieces form a table.

Structurally, they might have very little relationship with each other.

That's why extracting a table from a PDF can be much harder than extracting data from CSV or XLSX.

The Difference Between Text and Position

Imagine the following PDF page:

Date          Description             Amount
01/10/2026    Hosting                  $20
03/10/2026    Domain                   $12
05/10/2026    Software                 $49
Enter fullscreen mode Exit fullscreen mode

A human immediately understands the three columns.

A PDF extraction process may instead receive something closer to:

01/10/2026
Hosting
$20
03/10/2026
Domain
$12
05/10/2026
Software
$49
Enter fullscreen mode Exit fullscreen mode

The information hasn't necessarily disappeared.

The relationship between the pieces has.

The extraction logic now has to reconstruct that relationship.

And that's where things get interesting.

Digital PDFs vs Scanned PDFs

Before building any extraction pipeline, determine what kind of PDF you're dealing with.

There are two common cases.

Digital PDF

The document contains actual text objects.

You can usually:

  • Select text
  • Search for words
  • Copy text
  • Zoom without the text becoming part of an image

These PDFs are generally easier to process programmatically.

Scanned PDF

The page is essentially an image.

You can see the text, but the computer may not actually have text to extract.

For example:

[Image of a bank statement]
Enter fullscreen mode Exit fullscreen mode

There is no actual "Salary" string sitting inside the PDF waiting for your code to read.

You need OCR.

OCR Changes the Problem

OCR stands for Optical Character Recognition.

Instead of extracting existing text, OCR looks at an image and tries to recognize characters.

The workflow becomes something like:

PDF
 ↓
Page image
 ↓
OCR
 ↓
Recognized text
 ↓
Table detection
 ↓
Structured data
 ↓
Excel / CSV / JSON
Enter fullscreen mode Exit fullscreen mode

Every additional step introduces another opportunity for errors.

For example, OCR may confuse:

0 → O
1 → I
5 → S
Enter fullscreen mode Exit fullscreen mode

That's not a huge problem when extracting ordinary paragraphs.

It's a serious problem when extracting financial data.

Imagine:

₹5,000
Enter fullscreen mode Exit fullscreen mode

becoming:

₹50O0
Enter fullscreen mode Exit fullscreen mode

A human can often spot the mistake.

A script may not.

Why Columns Get Mixed Up

Let's say the PDF visually contains:

Name          Amount       Status
John          1200         Paid
Sarah         850          Pending
David         2400         Paid
Enter fullscreen mode Exit fullscreen mode

But the underlying text has different coordinates.

A parser may extract:

John
1200
Paid
Sarah
850
Pending
David
2400
Paid
Enter fullscreen mode Exit fullscreen mode

Now you need some logic to decide:

  • Which text belongs to the same row?
  • Where does one row end?
  • Where does the next row start?
  • Which X-coordinate represents each column?

This is why table extraction often relies on both text content and spatial information.

Coordinates Matter More Than You Think

When dealing with PDFs programmatically, the position of text can be just as important as the text itself.

A simplified representation might look like:

Text        X       Y
John        100     500
1200        300     500
Paid        450     500
Sarah       100     470
850         300     470
Pending     450     470
Enter fullscreen mode Exit fullscreen mode

Now the structure becomes easier to reconstruct.

Items with similar Y coordinates probably belong to the same row.

Items with similar X coordinates probably belong to the same column.

Real documents are more complicated, of course, but the basic idea is useful.

Tables Without Borders Are Even More Interesting

Developers sometimes assume that table extraction depends on visible borders.

Not necessarily.

A PDF can contain a table that has no lines at all.

Something like:

Product              Quantity       Price

Keyboard             20             500
Mouse                35             700
Monitor              10             2000
Enter fullscreen mode Exit fullscreen mode

A human sees a table because of the alignment.

A parser needs to infer that alignment.

That's why two visually similar documents can produce very different extraction results.

Multi-Page Tables Create Another Problem

Consider a report where the table starts on page 1 and continues for 20 pages.

Every page might contain:

Date | Description | Debit | Credit
Enter fullscreen mode Exit fullscreen mode

The extraction output could become:

Date | Description | Debit | Credit
...

Date | Description | Debit | Credit
...

Date | Description | Debit | Credit
...
Enter fullscreen mode Exit fullscreen mode

Those repeated headers aren't transaction records.

If you're building an automated pipeline, you need to recognize and remove them.

This is one of those problems that seems trivial until you process thousands of documents.

Don't Trust the First Successful Conversion

A file downloading successfully doesn't mean the extraction succeeded.

This distinction is important.

There are at least three different outcomes:

1. Conversion failed

No usable output was produced.

2. Conversion succeeded technically

An Excel file was generated, but the structure is wrong.

3. Conversion succeeded semantically

The resulting spreadsheet actually represents the information correctly.

The third one is what you really care about.

A generated .xlsx file by itself doesn't prove anything.

A Better Way to Validate Extracted Data

If you're building a document-processing system, add validation.

For example, suppose you're extracting financial transactions.

You could check:

Number of extracted rows
+
Number of valid dates
+
Number of numeric amounts
+
Expected column count
+
Opening/closing balance
Enter fullscreen mode Exit fullscreen mode

If the original statement says there are 250 transactions and your parser produces 37 rows, something clearly went wrong.

Likewise, if an amount column suddenly contains mostly text, that's a warning sign.

Validation is often more useful than simply improving the extraction algorithm forever.

What Should You Do With the Extracted Data?

Once the table is reliable, you can convert it into whatever format your application needs.

For example:

{
  "date": "2026-10-05",
  "description": "Software subscription",
  "amount": 49,
  "type": "debit"
}
Enter fullscreen mode Exit fullscreen mode

From there, you could store it in:

  • PostgreSQL
  • MySQL
  • MongoDB
  • CSV
  • Excel
  • An internal API
  • An accounting system

This is where PDF extraction becomes more than a file-conversion problem.

It becomes a data-ingestion problem.

A Practical Workflow

For a general-purpose PDF table extraction system, I'd think about the pipeline like this:

Upload PDF
    ↓
Identify PDF type
    ↓
Text available?
   / \
 Yes  No
  ↓    ↓
Parse  OCR
  \    /
   ↓  ↓
Detect table
    ↓
Reconstruct rows/columns
    ↓
Validate data
    ↓
Clean formatting
    ↓
Export
Enter fullscreen mode Exit fullscreen mode

The important part is the validation step.

Skipping it can make an automated system look successful while quietly producing bad data.

When Manual Cleanup Is Actually Fine

There's sometimes a temptation to automate every last part of the process.

That's not always necessary.

If someone converts a five-page PDF once a month, spending two minutes cleaning the spreadsheet may be perfectly reasonable.

Automation becomes more valuable when:

  • The same process happens repeatedly
  • Documents contain hundreds of rows
  • Multiple users need the workflow
  • Data needs to enter another system
  • Manual errors are expensive
  • Documents arrive continuously

The right level of automation depends on the workload.

If You Just Need a Spreadsheet

Not everyone needs to build an extraction pipeline.

If your goal is simply to take a table from a PDF and work with it in Excel, an online PDF to Excel converter can be much quicker than setting up a complete document-processing stack.

The same rules still apply, though.

Check whether the PDF is digital or scanned, review the extracted table, and verify important values.

The Main Lesson

PDF extraction becomes much easier once you stop thinking of a PDF as a document full of tables.

Think of it as a visual layout that you have to reconstruct into structured data.

Sometimes the structure is already available through text objects.

Sometimes you have to recover it from coordinates.

Sometimes you first need OCR.

And sometimes the document is simply messy enough that a little manual cleanup is the sensible option.

Once you understand what's happening underneath, those mysterious “why did my columns move?” problems become much easier to diagnose.

The goal isn't just to produce an Excel file.

The goal is to produce correct, usable data.

Top comments (0)