DEV Community

Cover image for Extract tables from a PDF
Felipe Cardona
Felipe Cardona

Posted on Originally published at anyformat.ai

Extract tables from a PDF

Extracting a table from a PDF means getting its rows and columns back as data, a list of rows or a pandas DataFrame, instead of a pile of text. It is harder than it looks, because a PDF stores characters at positions on a page and has no concept of a row or a column. A library has to infer the table from the lines drawn around the cells or from how the words line up.

You usually want this to load figures into a spreadsheet or database, to feed a table into an analysis, or to give a language model a table it can read. This page runs the three most used Python libraries on the same five fictional PDFs, shows the code that worked, and reports where each one broke.

What you will need

  • Python 3.10 or newer.
  • One PDF with a table in it, ideally a fictional or sample document, because the examples print what they find and text from real client documents can end up in terminal and job logs. Use the hardest one you have: no ruling lines, merged header cells or a table that runs across pages. A tidy ruled table makes every library look good.
  • One virtual environment per library, so their dependencies do not collide:
pip install pdfplumber       # environment 1
pip install camelot-py       # environment 2
pip install pymupdf          # environment 3
Enter fullscreen mode Exit fullscreen mode

Step 1. Check that the PDF has a text layer

All three libraries read the text stored in the PDF. A scan is an image with no text layer, and each of them returned an empty list on our scanned test file without raising an error. Check every page first, because a text PDF can have a short cover page or a blank one:

import pdfplumber

with pdfplumber.open("table.pdf") as pdf:
    pages_without_text = [
        number
        for number, page in enumerate(pdf.pages, start=1)
        if len((page.extract_text() or "").strip()) <= 50
    ]

if pages_without_text:
    print("pages with little or no text (possible scans):", pages_without_text)
else:
    print("every page has a text layer")
Enter fullscreen mode Exit fullscreen mode

A page listed there is worth a look rather than a verdict, since a cover or a blank page also has little text. If a page with a table has no text layer, none of the code below will work and you need OCR or a layout-aware parser first; see the best Python OCR libraries for the OCR options.

Step 2. Extract the table with each library

pdfplumber, the lightest of the three, reads ruled tables by default:

import pdfplumber

with pdfplumber.open("table.pdf") as pdf:
    tables = pdf.pages[0].extract_tables()

if tables:
    print(tables[0])  # a list of rows
else:
    print("no table found")
Enter fullscreen mode Exit fullscreen mode

Camelot returns a pandas DataFrame, with lattice for ruled tables and stream for borderless ones:

import camelot

tables = camelot.read_pdf("table.pdf", pages="all", flavor="lattice")
print(len(tables))
if len(tables):
    print(tables[0].df)
Enter fullscreen mode Exit fullscreen mode

PyMuPDF finds tables with find_tables, and the result converts to a DataFrame too:

import pymupdf

doc = pymupdf.open("table.pdf")
tabs = doc[0].find_tables()
if tabs.tables:
    print(tabs[0].extract())      # a list of rows
    df = tabs[0].to_pandas()
Enter fullscreen mode Exit fullscreen mode

What we saw

We generated five fictional PDFs with a known table: one with ruling lines, one without any, one with merged header cells and a multi-line cell, one that runs across two pages with a repeated header, and one that is a scan. We compared every output cell by cell against the table we generated. The versions were pdfplumber 0.11.10, Camelot 2.0.0 and PyMuPDF 1.28.2, and each library was run in its default mode and in its text-based mode.

pdfplumber Camelot PyMuPDF
Ruled table Correct Correct Correct
No ruling lines, default mode Empty Empty Empty
No ruling lines, text mode Right cells plus blank rows Right cells plus a title row Right cells plus blank rows
Merged header, multi-line cell Correct in default mode; broken in text mode Correct in lattice; broken in stream Correct in default mode; broken in text mode
Two pages Two tables, header repeated Two tables, header repeated Two tables, header repeated
Scan, no text layer Empty list, no warning Empty list, no warning Empty list, no warning

This is five small, clean, generated files, so it shows how each library fails and says nothing about how often it fails on your documents.

Ruling lines decide everything in the default modes

The default mode of each library is the same idea with a different name: find the lines drawn around the cells and read the text inside them. On the ruled table all three were correct. On the table with no ruling lines all three returned nothing, and on that one you have to switch to the text-based strategy, which guesses cells from word positions.

The text-based mode works but needs cleaning

With vertical_strategy and horizontal_strategy set to "text", pdfplumber found the right cells but added a blank row between every row of the table (11 rows where 6 were expected):

import pdfplumber

with pdfplumber.open("table.pdf") as pdf:
    found = pdf.pages[0].extract_tables(
        {"vertical_strategy": "text", "horizontal_strategy": "text"}
    )

rows = [r for r in found[0] if any(r)] if found else []  # drop the blank rows
Enter fullscreen mode Exit fullscreen mode

After dropping empty rows the table was correct. PyMuPDF's text strategy gave the same blank rows, and Camelot's stream mode added the page title as a first row. On the table with merged header cells and a multi-line cell, the text modes broke: the cell text was split across rows and only one of five rows came out intact in pdfplumber and PyMuPDF. PyMuPDF also cut "reorder monthly" to "reorder mont" in that run.

Merged cells and page breaks are not handled for you

A merged header such as "Sales" over two sub-columns comes out as "Sales" followed by an empty cell, with the sub-headers on a second row. A table that runs across two pages comes back as two tables, each with its own copy of the header, and none of the three libraries joins them. You concatenate the pieces and drop the repeated header yourself.

A scan returns an empty list, not an error

On the scanned file every library returned an empty result without a warning, which is easy to miss in a pipeline. That is why Step 1 checks for a text layer first.

The same test on a real financial PDF

The video runs the three libraries on a real financial document from the public ParseBench dataset: two tables with 45 companies, multi-line names and merged cells. Every row was checked against the ground truth. The results shown in the video were:

Approach Rows correct What went wrong
pdfplumber 51.1% Drops names in merged cells
Camelot, lattice 0% No ruling lines to detect
Camelot, stream 88.9% Merges the two tables
PyMuPDF 82.2% Scrambles text in complex cells
anyformat 95.6% 43 of 45 companies

It is one document and one run each, so read it as a pattern and not as a benchmark. The pattern matches the synthetic test above: ruling lines decide the default modes, and the text-based modes keep more rows but damage them.

None of the three libraries tells you that a row is wrong. They return a table either way, so a damaged cell looks exactly like a correct one, and that is why the check below is worth writing. anyformat returns the same table with a confidence score per field and the source text each value was read from, which is what lets a person check only the doubtful rows.

Check on your own PDFs

Write down one row you know is in your table and test for it. The function reports only which check failed, never the cell contents, so its output is safe to log.

def check_table(table: list[list[str]], expected_row: list[str]) -> list[str]:
    problems = []
    if not table:
        problems.append("no table found")
    elif expected_row not in table:
        problems.append("expected row is not intact")
    return problems

print(check_table(tables[0] if tables else [], ["North", "1200", "1350", "1480"]))
Enter fullscreen mode Exit fullscreen mode

Run it for each library on the same file and you have a comparison for your documents, which is worth more than ours.

Reproduce the test

The five PDFs come from a short script, so you can rerun everything. It needs pip install reportlab:

import json
from reportlab.lib.pagesizes import A4
from reportlab.platypus import Table, TableStyle, SimpleDocTemplate, Paragraph, Spacer
from reportlab.lib import colors
from reportlab.lib.styles import getSampleStyleSheet

hdr = ["Region", "Q1", "Q2", "Q3"]
rows = [["North", "1200", "1350", "1480"], ["South", "980", "1010", "1100"],
        ["East", "1500", "1620", "1710"], ["West", "760", "810", "905"],
        ["Total", "4440", "4790", "5195"]]

def simple(filename, data, grid):
    style = [("FONTSIZE", (0, 0), (-1, -1), 11), ("ALIGN", (1, 0), (-1, -1), "RIGHT"),
             ("BOTTOMPADDING", (0, 0), (-1, -1), 6), ("TOPPADDING", (0, 0), (-1, -1), 6)]
    if grid:
        style += [("GRID", (0, 0), (-1, -1), 0.8, colors.black),
                  ("BACKGROUND", (0, 0), (-1, 0), colors.lightgrey)]
    table = Table(data, colWidths=[110, 90, 90, 90])
    table.setStyle(TableStyle(style))
    SimpleDocTemplate(filename, pagesize=A4).build(
        [Paragraph("Sales by region", getSampleStyleSheet()["Heading2"]), Spacer(1, 12), table])

simple("a_ruled.pdf", [hdr] + rows, grid=True)
simple("b_borderless.pdf", [hdr] + rows, grid=False)

merged = [["Product", "Sales", "", " Notes"], ["", "H1", "H2", ""],
          ["Widget", "100", "120", "Best seller\nreorder monthly"],
          ["Gadget", "80", "95", "Discontinued\nin Q4"], ["Gizmo", "60", "70", "New"]]
table = Table(merged, colWidths=[90, 70, 70, 150])
table.setStyle(TableStyle([("GRID", (0, 0), (-1, -1), 0.8, colors.black),
                           ("SPAN", (1, 0), (2, 0)), ("SPAN", (0, 0), (0, 1)), ("SPAN", (3, 0), (3, 1)),
                           ("ALIGN", (0, 0), (-1, 1), "CENTER"), ("VALIGN", (0, 0), (-1, -1), "MIDDLE")]))
SimpleDocTemplate("c_merged.pdf", pagesize=A4).build([table])

long = [hdr] + [[f"Item{i:02d}", str(100 + i * 7), str(200 + i * 3), str(300 + i * 11)] for i in range(1, 61)]
table = Table(long, colWidths=[110, 90, 90, 90], repeatRows=1)
table.setStyle(TableStyle([("GRID", (0, 0), (-1, -1), 0.8, colors.black),
                           ("BACKGROUND", (0, 0), (-1, 0), colors.lightgrey), ("FONTSIZE", (0, 0), (-1, -1), 11)]))
SimpleDocTemplate("d_twopage.pdf", pagesize=A4).build([table])
Enter fullscreen mode Exit fullscreen mode

For the scan, render the ruled PDF to an image and save it back as a PDF with no text layer:

import pymupdf

src = pymupdf.open("a_ruled.pdf")
pix = src[0].get_pixmap(dpi=150)
out = pymupdf.open()
page = out.new_page(width=pix.width, height=pix.height)
page.insert_image(page.rect, pixmap=pix)
out.save("e_scan.pdf")
Enter fullscreen mode Exit fullscreen mode

Install and licence notes

Camelot is the heaviest install of the three, around 157 MB with OpenCV, against about 42 MB for pdfplumber and 53 MB for PyMuPDF in our environments. Camelot 2.0.0 no longer needs Ghostscript by default and converts pages through pypdfium2. Older 0.x versions, which most tutorials still cover, failed on every default call with "Ghostscript is not installed" in our run, even with the gs binary on the path, so install the current version.

pdfplumber is MIT licensed and Camelot is MIT licensed. PyMuPDF is dual licensed under AGPL 3.0 or a commercial license from Artifex, which matters if you ship closed-source software; read its licensing page for your case. All three had a release within the last four months on 2026-10-01.

When these libraries are not enough: anyformat

These libraries fit PDFs with a text layer and tables that are ruled or cleanly aligned. They stop fitting when tables have no lines, merged cells, multi-line cells or page breaks, or when the files are scans. For those cases people turn to a layout-aware parser, a vision language model or an OCR step that finds table structure.

anyformat's parsing is one option. Its parse result lists a document's blocks, and table blocks carry their rows, according to the SDK reference. We did not run it on the five synthetic files; in the video it ran on the real financial PDF above and got 43 of 45 companies right, which is one document and one run. For long tables specifically, see long-document extraction.

Watch the test

The video version of this test runs about five minutes and covers the three libraries and the real-PDF results above.

https://www.youtube.com/watch?v=yf3IP0Bx1oI

Frequently asked questions

What is the best Python library to extract tables from a PDF?
It depends on the table. For ruled tables in text-based PDFs all three worked in our test, and pdfplumber was the lightest to install. For tables without ruling lines, none of the defaults worked and every library needed its text-based mode plus cleanup.

Can pdfplumber extract tables from a scanned PDF?
No. It reads the text layer, and a scan has none, so it returns an empty list. You need OCR or a layout-aware parser first.

How do I get a PDF table into a pandas DataFrame?
Camelot returns one directly (tables[0].df), PyMuPDF converts with to_pandas(), and with pdfplumber you pass the list of rows to pandas.DataFrame, using the first row as the header.

Why does my table come out with blank rows?
The text-based strategies in pdfplumber and PyMuPDF find the cell edges from word positions and can add empty rows between real ones. Drop the rows where every cell is empty.

How do I extract a table that spans several pages?
Each library returns one table per page. Concatenate the tables and drop the repeated header rows yourself.


Originally published at anyformat.ai/blog/extract-tables-from-pdf.

Top comments (0)