DEV Community

Cover image for How to Extract Line Items, Totals, and Customer Info from Purchase and Sales Orders via API
PDF4me
PDF4me

Posted on

How to Extract Line Items, Totals, and Customer Info from Purchase and Sales Orders via API

Every procurement, fulfillment, or ERP-adjacent team ends up building the same throwaway tool eventually: something that opens a purchase order PDF, finds the line items, finds the total, and matches it against whatever system of record actually runs the business. Then a new supplier sends a PO in a different layout, the regex breaks, and someone spends an afternoon patching a parser nobody wanted to own in the first place.

Purchase orders and sales orders are two sides of the same transaction (one side issues it, the other side fulfills it) but they get treated as a document-handling problem twice, once per direction, because most parsing tools make you describe the layout before they'll read it. That's the part worth fixing, not the OCR.

Why one parser has to cover both documents

A purchase order and a sales order carry almost the same information: an order number, a date, who it's for, what's being ordered, how many, at what price, and what the total comes to. The difference is which side of the transaction wrote it. A buyer's procurement system generates POs; a seller's order-entry system generates SOs. Downstream, a reconciliation or fulfillment system usually needs to read both, because one company's outgoing PO is often another company's incoming SO in the same workflow.

That shared shape is exactly why a schema-driven parser works here in a way that template matching doesn't. A capture-area tool anchors to pixel coordinates, so a vendor that moves their logo six millimeters to the left breaks the template. A regex-based tool anchors to text patterns, so "Total:" vs "Total Due:" vs "Amount Due" each needs its own rule. Neither approach generalizes across a supplier list you don't control the formatting of.

What the AI Order Parser actually returns

PDF4me's AI Order Parser skips capture areas and regex entirely: you send it a PDF or image, it returns a fixed JSON shape, and per PDF4me's own Make.com integration documentation that shape is identical whether the source document is a purchase order or a sales order ("works without template setup," "the two document types return the same schema").

The documented output fields:

  • orderNumber (string): order or reference number, whatever format the document uses (PO-2024-001, SO-78423, etc.)
  • orderDate (string): normalized to ISO 8601 (YYYY-MM-DD)
  • customerName (string): the customer or vendor name on the document
  • lineItems (array): each row carries description, quantity, unitPrice, total
  • subTotal (number): amount before taxes, shipping, discounts
  • total (number): final amount including all charges
  • currency (string): ISO 4217 code (USD, EUR, INR, ...)
  • shippingAddress (string): shipping or delivery address
  • jobId (string): PDF4me's own job ID, for audit/support
  • success (boolean) / message (string): completion status

A realistic response:

{
  "success": true,
  "jobId": "a1b2c3d4-e5f6-4789-9abc-def012345678",
  "orderNumber": "PO-2024-00187",
  "orderDate": "2026-09-14",
  "customerName": "Meridian Fabrication Ltd.",
  "currency": "USD",
  "lineItems": [
    { "description": "M8 Stainless Hex Bolt, 40mm", "quantity": 500, "unitPrice": 0.42, "total": 210.00 },
    { "description": "M8 Stainless Washer", "quantity": 500, "unitPrice": 0.06, "total": 30.00 }
  ],
  "subTotal": 240.00,
  "total": 259.20,
  "shippingAddress": "4410 Industrial Pkwy, Unit 12, Dayton, OH 45417",
  "message": "Extraction completed."
}
Enter fullscreen mode Exit fullscreen mode

That's enough to populate an ERP line, run a reconciliation threshold check, or trigger a fulfillment workflow without a human opening the PDF first.

A documentation gap worth being upfront about: there's no published REST-only walkthrough for this specific action on docs.pdf4me.com (confirmed 404, not assumed), and the official pdf4me-api-samples repo on GitHub doesn't have an Order Parser folder either, unlike several of the other AI parsers in this cluster. Make.com is currently the one platform with a dedicated, field-by-field written guide for this action. If you're calling this over the raw REST API, the schema above (verified against Make's docs) is what to expect back, but test against your own documents before you build a strict validator around it.

A second gap, this one a discrepancy rather than a missing page: PDF4me's Power Automate connector (listed on Microsoft's own connector reference as AI - Order Parser / operation ID ProcessOrderAi) documents a noticeably different output shape for what is nominally the same action, built around invoiceToName/deliverToName pairs, a products array, and a few fields that read like leftovers from one specific customer's template rather than a general-purpose schema. That's either a different model version, a stale page on Microsoft's side, or a connector-specific response wrapper, and there's no way to tell which from the outside. If you're building on Power Automate, run a live test order through the action and check the actual field names it returns before wiring a Flow around either schema.

Where this fits in a no-code or custom stack

  • Make: the one platform with a dedicated, written walkthrough for the AI-Process Order module, including the field table above. Map Order Name and the document binary, run an Iterator over lineItems if you need one ERP row per line item, done.
  • Power Automate: the same capability exists as a native connector action (ProcessOrderAi), available across Power Automate, Power Apps, and Copilot Studio (Premium) per Microsoft's own connector listing, with a 100-calls-per-60-seconds throttle per connection. Verify the live output fields yourself, per the discrepancy noted above.
  • REST API: PDF4me's broader AI parser family is reachable directly for teams building their own backend instead of a no-code flow, following the same API key and connection pattern as PDF4me's other endpoints.
  • When the fixed schema isn't enough: orders carrying fields the standard parser doesn't name (a custom SKU format, a project code, a buyer-specific approval field) are a better fit for PDF4me's Universal Document Parser, which lets you specify the exact fields to extract, or the Document Parser / custom analyzer path, which lets you define a reusable extraction profile in the PDF4me dashboard. Reach for either one before writing a post-processing script to bolt extra fields onto the Order Parser's output.

The actual use case this solves

Picture a small distributor that takes purchase orders by email, as PDFs, from forty different retail accounts, each using whatever order-entry system that retailer happens to run. Nobody on the distributor's side wants to build forty templates. Dropping every incoming PO through the Order Parser, regardless of sender, normalizes them all into the same orderNumber / lineItems / total shape before anything touches the warehouse management system. The matching logic downstream (does this line item exist in inventory, does the total reconcile, does this need a human to look at it) only has to be written once, against one schema, instead of once per retailer.

The same logic runs in reverse for a seller reconciling their own outgoing sales order confirmations against what a buyer's procurement system actually received, which is the other half of why one shared schema across both document types is worth more than it sounds.

What this doesn't solve

This parser reads documents, it doesn't validate business logic. It won't tell you whether a total is actually correct given the lineItems listed (that's a reconciliation check you still write), and it won't catch a line item that's missing entirely if the source document itself never listed it. Treat the extraction as trustworthy input to your own validation step, not a replacement for one.

Website: pdf4me.com
Documentation: docs.pdf4me.com

Top comments (0)