PDF is one of the most common document formats in daily work, but its "tables" are a different thing from tables in Excel: a PDF file itself does not store structured table data. A table on a page is merely a visual effect presented by text and lines according to their positional relationships. Therefore, reading PDF tables programmatically is essentially about letting a library analyze the page layout and identify the structure of rows and columns.
This article uses Spire.PDF for Python's PdfTableExtractor to demonstrate how to detect tables in a PDF page by page and read each cell as text. The overall process has only three steps: load the document → extract tables page by page → iterate through rows and columns to read cells.
1. Environment Setup
Install the library via pip:
pip install Spire.PDF
If you are just learning or handling small documents, you can also install the free version:
pip install Spire.Pdf.Free
Note that the free version has a page limit when loading and processing PDFs (at most the first 10 pages). Documents longer than 10 pages need to be split before processing, or you need to use the full PyPI version.
2. Core Code
from spire.pdf import PdfDocument, PdfTableExtractor
# Load PDF document
pdf = PdfDocument()
pdf.LoadFromFile("input.pdf")
# Create a PdfTableExtractor object
table_extractor = PdfTableExtractor(pdf)
# Extract tables from each page
for i in range(pdf.Pages.Count):
tables = table_extractor.ExtractTable(i)
for table_index, table in enumerate(tables):
print(f"Table {table_index + 1} on page {i + 1}:")
for row in range(table.GetRowCount()):
row_data = []
for col in range(table.GetColumnCount()):
text = table.GetText(row, col).replace("\n", " ")
row_data.append(text.strip())
print("\t".join(row_data))
3. Step-by-Step Code Explanation
1. Load the PDF document
pdf = PdfDocument()
pdf.LoadFromFile("input.pdf")
PdfDocument is the document entry object. LoadFromFile() loads the PDF by passing in the file path. The file path can be either relative or absolute.
2. Create the table extractor
table_extractor = PdfTableExtractor(pdf)
PdfTableExtractor is responsible for analyzing the layout page by page and detecting table structures. When constructing it, you need to pass in the already loaded PdfDocument object.
3. Extract tables page by page
for i in range(pdf.Pages.Count):
tables = table_extractor.ExtractTable(i)
pdf.Pages.Count is the total number of pages in the document, and ExtractTable(i) returns all tables detected on page i. Note that it returns a list — a single page may contain multiple tables at the same time; if there are no tables on that page, it returns an empty list.
4. Iterate through the table's rows and columns
for row in range(table.GetRowCount()):
row_data = []
for col in range(table.GetColumnCount()):
text = table.GetText(row, col).replace("\n", " ")
row_data.append(text.strip())
print("\t".join(row_data))
-
GetRowCount()/GetColumnCount()get the number of rows and columns in the table; -
GetText(row, col)reads the text of the specified cell. A cell may contain line breaks; here\nis replaced with a space, and thenstrip()removes leading and trailing whitespace to avoid extra blank lines in the output; - Finally,
\t(tab) is used to join the cells of a row for output, making it easy to copy into Excel or use as TSV.
4. Running Effect
Running the above code on a PDF containing a simple table produces output roughly as follows:
Table 1 on page 1:
Name Department Salary
Zhang San Technical Department 12000
Li Si Marketing Department 9800
Wang Wu Finance Department 10500
Each line corresponds to a row in the table, and cells are separated by tabs, so it can be pasted directly into spreadsheet software.
5. Advanced: Save the Results as CSV
Printing to the console is only suitable for quick viewing. If you want to save the results, you can use Python's built-in csv module, with no need to install an additional library:
import csv
from spire.pdf import PdfDocument, PdfTableExtractor
pdf = PdfDocument()
pdf.LoadFromFile("input.pdf")
extractor = PdfTableExtractor(pdf)
with open("output.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
for i in range(pdf.Pages.Count):
for table in extractor.ExtractTable(i):
for row in range(table.GetRowCount()):
row_data = [
table.GetText(row, col).replace("\n", " ").strip()
for col in range(table.GetColumnCount())
]
writer.writerow(row_data)
pdf.Close()
When writing with csv.writer, the program automatically handles special characters such as commas and quotation marks contained in cells, which is more reliable than manually concatenating strings.
6. Notes and Limitations
1. Suitable for tables with clear borders. PdfTableExtractor relies on analyzing the page layout. It recognizes tables with clear borders and regular structures relatively well; for tables without visible borders, cells spanning multiple lines, or complex layouts where headers are not clearly identified, it may fail to detect them or produce incomplete structures.
2. Scanned documents require OCR first. If the PDF is a scanned image (without a text layer), no text-based tool can directly read the content. You need to use OCR first to convert the image into text. This library does not provide OCR functionality.
3. Page limit of the free version. The free version can process at most the first 10 pages. When handling long documents, it is recommended to add a min(pdf.Pages.Count, 10) limit in the loop, or split the document by page.
4. Release resources after use. When batch processing a large number of files, remember to call pdf.Close() at the end to release resources.
7. Other Optional Libraries
If you encounter complex tables that Spire.PDF cannot handle, you can try the following open-source libraries as supplements:
- pdfplumber : Based on pdfminer.six, it provides finer-grained access to page objects and is suitable for handling complex layouts;
- camelot : Identifies structures by detecting table lines and works well for bordered tables;
- tabula-py : Wraps Java's tabula, and its output is directly compatible with pandas DataFrame.
In actual projects, you can combine them according to the table type: regular tables can be handled by any library, while for complex ones, it is better to test on a small sample first before choosing.
8. Summary
The process of extracting PDF tables with Spire.PDF for Python is very straightforward: load the document → create PdfTableExtractor → call ExtractTable() page by page → iterate through rows and columns to read cells. It is suitable for handling bordered, clearly structured tables, and with the csv module, the results can be saved to disk. For special cases such as borderless tables and scanned documents, OCR or other libraries are needed as supplements.

Top comments (1)
tr.ee/dev-to