DEV Community

Jack9012
Jack9012

Posted on

Downloading PDF Documents from URLs with Python

In scenarios such as automated office work, document collection, and batch resource fetching, it is often necessary to download PDF files from the network programmatically. Directly writing the binary stream returned by an interface to a local file can easily lead to corrupted files or format anomalies. This article uses requests to handle network requests, combined with Spire.PDF for .NET to complete PDF stream validation and persistent saving, providing a ready-to-run download solution with built-in file validity checking.

Download PDF from URL

1. Installing Environment Dependencies

1.1 Network Request Library: requests

Used to make HTTP requests and retrieve remote PDF binary data:

pip install requests
Enter fullscreen mode Exit fullscreen mode

1.2 PDF Processing Library

This example relies on this library for in-memory stream loading, PDF validation, and local file export. The installation command is as follows:

pip install spire.pdf
Enter fullscreen mode Exit fullscreen mode

2. Complete Runnable Code

import requests
from spire.pdf import *

def download_pdf_from_url():
    # Specify the remote PDF resource URL
    url = "resource/sample.pdf"

    # Send a GET request to fetch the file's binary data
    response = requests.get(url)
    # Automatically raise 4xx/5xx HTTP errors to catch dead links and server issues in advance
    response.raise_for_status()
    # Wrap the downloaded byte data as an in-memory stream
    stream = Stream(response.content)
    # Load the PDF document from the memory stream, automatically validating whether the file is a legitimate PDF
    document = PdfDocument(stream)
    # Save the validated PDF to a local file
    document.SaveToFile("Downloaded.pdf")
    # Close the document to release memory resources
    document.Close()
    print("PDF downloaded and saved successfully!")

if __name__ == "__main__":
    download_pdf_from_url()
Enter fullscreen mode Exit fullscreen mode

3. Line-by-Line Code Breakdown

3.1 Module Imports

import requests
from spire.pdf import *
Enter fullscreen mode Exit fullscreen mode
  • requests: a general-purpose Python HTTP library, responsible for remote file downloads;
  • The PDF-related namespace: from Spire.PDF for .NET, providing full capabilities for stream reading and PDF document operations.

3.2 Remote Request and Exception Handling

response = requests.get(url)
response.raise_for_status()
Enter fullscreen mode Exit fullscreen mode

raise_for_status() is crucial fault-tolerant logic: when a link returns 404, the server returns 500, or access is denied, the program throws an exception directly instead of generating a corrupted blank file.

3.3 Loading the PDF from a Memory Stream (Core Advantage)

The conventional download logic directly uses open to write the byte stream to a file, without being able to determine whether the returned data is a valid PDF.

This solution first converts the binary data returned by the interface into a Stream, then hands it to PdfDocument for loading:

  1. Automatically validates whether the data stream conforms to the PDF standard;
  2. Filters out anomalies such as interrupted downloads and HTML error pages returned in place of the file;
  3. Processes everything in memory, without generating temporary cache files.

3.4 Saving the File and Releasing Resources

SaveToFile supports custom output locations via relative or absolute paths. After processing, call Close() to release the memory occupied by the document, which effectively prevents memory overflow during long batch download loops.

4. Practical Usage Notes

  1. URL requirements In the example, resource/sample.pdf is a relative resource path; in production, replace it with a full http/https public PDF link. For private links requiring login authentication or Cookie validation, add the headers and cookies parameters to requests.get.
  2. Extended exception handling The basic code only intercepts HTTP errors. In a production environment, it is recommended to wrap the logic in try except to catch exceptions such as PDF parsing failures and insufficient file-write permissions:
   try:
       # Complete download logic
   except requests.exceptions.RequestException as http_err:
       print(f"Network request failed: {http_err}")
   except Exception as pdf_err:
       print(f"PDF processing failed, the file may be corrupted: {pdf_err}")
Enter fullscreen mode Exit fullscreen mode
  1. Batch download adaptation Store multiple PDF URLs in a list, call download_pdf_from_url in a loop, and modify the output file names to avoid overwriting, enabling batch archiving of remote PDFs.

5. Summary

This download solution combines the network capabilities of requests with Spire.PDF for .NET's PDF parsing capabilities. Its biggest highlight is validating file legitimacy at download time, solving the pain point of ordinary download methods easily producing corrupted PDFs. The code is compact and highly extensible, making it suitable for various development scenarios such as script automation, backend document synchronization, and crawler resource collection.

Top comments (0)