DEV Community

Allen Yang
Allen Yang

Posted on

How to Split PDF Documents with Python

How to Split PDF Documents with Python

Large PDFs are everywhere in day-to-day work — consolidated quarterly reports, scanned contracts, batches of exported invoices. The recipient usually needs only a few pages: individual pages split out as separate files, or a specific page range handed to a different team. Doing this by hand in a PDF editor is slow and impossible to scale. With Python you can split a document page by page, extract selected pages, or even divide one oversized page into several standard pages — all of which slot naturally into automated workflows. This article demonstrates these three common splitting scenarios using Spire.PDF for Python.

Why Split PDFs with Python

Handling the job in a script has a few concrete advantages:

  • Batch processing: one script can process hundreds of files without manual clicks.
  • Flexible granularity: split a whole document page by page, extract only chosen pages, or subdivide a single long page.
  • Easy integration: splitting is just one step in a pipeline that may also download, archive, or email the results.

Setting Up the Environment

Install the Spire.PDF library:

pip install Spire.PDF
Enter fullscreen mode Exit fullscreen mode

Then import the required modules in your script:

from spire.pdf import *
from spire.pdf.common import *
Enter fullscreen mode Exit fullscreen mode

Splitting a PDF into Single-Page Documents

The most direct approach is turning every page into its own PDF. PdfDocument.Split() does this with a single filename pattern — no per-page handling required:

from spire.pdf import *
from spire.pdf.common import *

# Load the PDF document
doc = PdfDocument()
doc.LoadFromFile("Sample.pdf")

# Split by filename pattern; each page becomes an individual PDF
doc.Split("SplitDocument-{0}.pdf")

doc.Close()
Enter fullscreen mode Exit fullscreen mode

Split() takes a filename template containing the {0} placeholder, which is replaced with the page index starting from 0. A 10-page document produces ten single-page files, from SplitDocument-0.pdf through SplitDocument-9.pdf. Call Close() afterwards to release the document's resources.

Extracting Specific Pages into a New PDF

More often you do not need every page — just a selection combined into a new document. The approach: create a new PdfDocument, add a same-sized page for each source page you want, then draw the source page's content into it:

from spire.pdf import *
from spire.pdf.common import *

# Load the source document
oldPdf = PdfDocument()
oldPdf.LoadFromFile("Sample.pdf")

# Create the target document
newPdf = PdfDocument()
page = None

# Extract pages at index 1 and 2 (0-based)
for i in range(1, 3):
    # Add a page with the same size as the source page
    page = newPdf.Pages.Add(oldPdf.Pages[i].Size, PdfMargins(0.0))
    # Draw the source page content onto the new page
    oldPdf.Pages[i].CreateTemplate().Draw(page, PointF(0.0, 0.0))

newPdf.SaveToFile("ExtractPages.pdf")
newPdf.Close()
oldPdf.Close()
Enter fullscreen mode Exit fullscreen mode

The first argument to Pages.Add() receives the source page's Size, so the new page matches the original dimensions, and PdfMargins(0.0) zeroes out the margins to keep the content aligned. CreateTemplate() produces a drawing template of the source page, and Draw() renders it onto the new page with layout fully preserved. The loop range range(1, 3) extracts pages 2 and 3 — adjust the range to pull any combination of pages.

Splitting One Long Page into Multiple Pages

A special case: the source PDF contains a very tall page (a long screenshot, a continuous table) and you want it cut into standard-height pages. The trick is to build a document whose page height is half the source page's height, then draw the source content with automatic pagination enabled:

from spire.pdf import *
from spire.pdf.common import *

# Load the source document and get the first page
doc = PdfDocument()
doc.LoadFromFile("LongPage.pdf")
page = doc.Pages[0]

# New document: zero margins, same width, half the source height
newPdf = PdfDocument()
newPdf.PageSettings.Margins.All = 0.0
newPdf.PageSettings.Width = page.Size.Width
newPdf.PageSettings.Height = page.Size.Height / float(2)

# Add a new page and configure automatic pagination
newPage = newPdf.Pages.Add()
format = PdfTextLayout()
format.Break = PdfLayoutBreakType.FitPage
format.Layout = PdfLayoutType.Paginate

# Draw the source content; it flows onto extra pages automatically
page.CreateTemplate().Draw(newPage, PointF(0.0, 0.0), format)

newPdf.SaveToFile("SplitOnePage.pdf")
newPdf.Close()
doc.Close()
Enter fullscreen mode Exit fullscreen mode

Two properties of PdfTextLayout do the work: Layout = PdfLayoutType.Paginate makes the drawing flow onto additional pages when it exceeds the page, and Break = PdfLayoutBreakType.FitPage cuts at the page boundary. Since the new page is half the height of the source, one long page spreads across two standard pages; decrease the height divisor to cut it into more pieces.

Practical Tips

  • Output location: the filename pattern in Split() can include a path prefix (for example "output/SplitDocument-{0}.pdf") to write the results into a specific directory.
  • 0-based indexing: check doc.Pages.Count before extracting pages and choose the loop range accordingly to avoid out-of-range errors.
  • Granularity: for long-page splitting, a smaller divisor for Height yields more pages — tune it to the actual content height.
  • Resource cleanup: call Close() on both the source and target documents; this matters even more in batch processing.
  • Free edition: the free edition of Spire.PDF limits the number of pages processed, which is worth checking for large documents.

Conclusion

This article covered three common ways to split PDF documents in Python: using Split() to break an entire document into single-page files, using CreateTemplate().Draw() to extract selected pages into a new document, and leveraging PdfLayoutType.Paginate to divide one long page across multiple standard pages. Combined with file iteration and simple page math, these operations fit cleanly into workflows such as batch splitting and per-recipient document distribution.

Top comments (0)