DEV Community

Cover image for What's Hiding Inside a PDF? How to Extract Embedded File Attachments via API
PDF4me
PDF4me

Posted on

What's Hiding Inside a PDF? How to Extract Embedded File Attachments via API

A PDF is not always just the pages you can see. Open the file properties on a supplier invoice and you might find a UBL or ZUGFeRD-style XML file riding along inside it, the same numbers the human-readable page shows, but in a shape a machine can actually parse. A signed contract exported from a document management system can carry the original Word file, a redline history, or a spreadsheet of exhibits, all tucked inside the same PDF a reviewer only ever opens to read page one. Government and legal filings do this constantly, because the PDF/A-3 archive format and PDF Portfolios were built specifically to let one file carry other files inside it without breaking the "it's just a PDF" promise to whoever opens it.

The problem is that almost nothing in a typical document pipeline goes looking for those files. OCR reads the visible page. Text extraction reads the visible page. A human reviewer, understandably, reads the visible page. The attachment sits there, correctly embedded, technically present, and functionally invisible, until someone remembers to check, usually after the data it contained was needed somewhere and wasn't found.

What Extract Attachment from PDF Actually Recovers

PDF4me's Extract Attachment from PDF endpoint exists for exactly that gap. You send it a PDF, base64-encoded as docContent with its docName, and it hands back every file embedded inside that PDF as a separate object in an outputDocuments array, each one with its own fileName and its own base64 streamFile. Nothing about the attachment is interpreted or reformatted. If a spreadsheet went in, a spreadsheet comes out, byte-identical, ready to be decoded and saved.

The endpoint covers more than one packaging style: standard PDF attachments (the paperclip kind you would add manually in Acrobat), PDF Portfolios (a PDF built as a container for several documents at once), and PDF/A-3 files, the archival format that intentionally allows one embedded machine-readable file alongside the human-readable original. For anything larger than a one-off PDF, the endpoint also accepts an optional async boolean: set it, and instead of waiting on the response, you get a 202 Accepted back immediately with a Location header to poll.

Here is the request/response shape in Python, field names confirmed live against the endpoint's own docs page:

import requests
import base64

API_KEY = "your-api-key"
BASE_URL = "https://api.pdf4me.com/api/v2/"

with open("invoice.pdf", "rb") as f:
    doc_content = base64.b64encode(f.read()).decode("utf-8")

payload = {
    "docContent": doc_content,
    "docName": "invoice.pdf",
    "async": False
}

headers = {
    "Authorization": f"Basic {API_KEY}",
    "Content-Type": "application/json"
}

response = requests.post(BASE_URL + "ExtractAttachmentFromPdf", json=payload, headers=headers)

if response.status_code == 200:
    result = response.json()
    for doc in result["outputDocuments"]:
        file_bytes = base64.b64decode(doc["streamFile"])
        with open(doc["fileName"], "wb") as out:
            out.write(file_bytes)
        print(f"Recovered: {doc['fileName']}")
elif response.status_code == 202:
    location = response.headers["Location"]
    print(f"Processing async, poll: {location}")
else:
    print(f"Error {response.status_code}: {response.text}")
Enter fullscreen mode Exit fullscreen mode

The async branch above (202 plus a Location header to poll) matches the pattern this cluster has already live-verified on sibling file-conversion endpoints; the GitHub sample repo was not independently re-checked this run, so treat the exact retry and polling cadence as something to confirm against your own account before relying on it in production.

Where Do the Recovered Files Actually Go?

Extraction on its own is only half a workflow. The other half is what happens the moment those files land, and that is where PDF4me's automation platform integrations do the actual routing.

On n8n, the Extract Attachment From PDF node takes the same three inputs (how the PDF is supplied, the PDF itself, and its filename) and drops each recovered file back into the workflow as its own item, ready to feed into whatever node comes next.

On Zapier, the same action is framed around the workflows people actually build with it: pulling originals back out of a compliance archive, recovering the source file from a legal discovery export, or unpacking an email-to-PDF conversion that quietly buried the original attachment inside the PDF wrapper. The output includes the file's content, name, extension, a retrievable URL, and its size.

Make documents the same module with a detail worth planning around: because a single PDF can contain more than one embedded file, the recommended pattern is to chain a Make Iterator immediately after the module. Feed that Iterator into a Google Drive or Dropbox module, and you have a working "unpack and file away" pipeline without writing a parser for any of it.

If you're building on Power Automate, it's worth knowing up front that, unlike its n8n, Zapier, and Make counterparts, there is currently no dedicated Power Automate connector action published for this specific endpoint. The REST endpoint itself is still reachable from a Power Automate HTTP action if you need the capability there today.

What It Will Not Do for You

Extract Attachment from PDF recovers files. It does not read them. If the embedded object is a spreadsheet, you get the spreadsheet back exactly as it went in, not a summary of its rows. If it's an XML invoice payload, you get the XML, not a parsed set of line items.

And if a PDF has no embedded attachments at all, a perfectly normal case, the outputDocuments array simply comes back empty rather than erroring. It's also worth being honest that this endpoint solves a recovery problem, not a discovery problem: it will pull out every attachment a PDF actually has, but it will not tell you which PDFs in an archive are worth checking to begin with.

A Concrete Example Worth Picturing

Picture an accounts payable inbox that receives supplier invoices as PDFs generated by a dozen different ERPs. Some are ordinary scans. Others are PDF/A-3 e-invoices with a structured XML payload embedded, specifically so an accounting system does not have to OCR a total that is already sitting there in machine-readable form. Run every incoming PDF through Extract Attachment from PDF first, and the outputDocuments array tells you, instantly, which invoices came with a ready-to-parse XML file and which ones genuinely need OCR or manual entry.

Getting Set Up

Every call needs an API key from the PDF4me dashboard, attached per the connect to the API guide. Every request goes to the same base URL whether you're calling this endpoint or any other one in the catalog. Before wiring it into a real pipeline, PDF4me's API Tester is the fastest way to confirm what a specific PDF actually returns: upload a file, run the call directly in the browser, and see the real outputDocuments shape come back before you write a single line of integration code around an assumption about it.

Somewhere in your document pipeline right now, there is probably a PDF carrying more than the page you're reading. The only real question is whether you go looking for it on your own terms, or wait until someone needs what's inside it and discovers, too late, that it was there all along.

Website: pdf4me.com
Documentation: docs.pdf4me.com

Top comments (0)