What Does an AI Mortgage Document Parser Actually Return? Loan Terms, Borrower Info, and Closing Details Explained
Hand a mortgage file to a generic OCR tool and you get a transcript: every character on every page, roughly in reading order, with no idea which number is the loan amount and which one is the appraised value. That distinction is the whole job in a loan-processing or compliance pipeline, where the next system in line needs loanAmount in one field and appraisedValue in another, not forty pages of text it has to re-read itself. A mortgage package makes this worse than almost any other document type, because it is not one document. It is a stack: a promissory note, a deed of trust, an appraisal, a closing disclosure, sometimes a title report, each with its own layout and its own vocabulary for the same underlying facts. So what does a purpose-built mortgage document parser actually hand back, and why does its field list look the way it does?
Why a Transcript Is Not a Loan File
Lenders, servicers, and the software vendors who build for them do not want text. They want a loan amount they can run a calculation against, a closing date they can schedule a disbursement around, and a borrower name they can match to an existing record, every time, regardless of which title company or loan origination system produced the page. A generic OCR engine reads a mortgage document correctly as characters and still leaves someone to figure out which string means what, document by document, lender by lender. That is the gap PDF4me's AI Mortgage Document Parser is built to close. According to its own product description, it uses "machine learning to extract structured data from mortgage documents," which in practice means it understands mortgage paperwork well enough to return loan terms, borrower details, and closing information as named fields, not as a page of recovered text.
The Response Shape, Field by Field
Send the parser a mortgage document as a PDF, PNG, JPG, or JPEG, and the response groups into several clear categories. I live-checked the field list directly against PDF4me's own Power Automate and n8n integration docs for this action rather than trusting a single summary table, and it is worth doing the same before you wire anything to it, because the two pages don't phrase the grouping identically even though the field names match.
Loan information: documentType, lenderName, loanAmount, interestRate, loanTerm, loanType, effectiveDate.
Borrower details: borrowerName, coBorrowerName, employerName, position, income, employmentStatus, employmentStartDate.
Property and appraisal: propertyAddress, legalDescription, propertyType, conditionRating, appraisedValue, appraiserName, appraiserLicenseNumber, plus a comparableSales array holding each comparable's address, sale price, sale date, and adjustments.
Closing and title: closingCosts, cashToClose, settlementAgent, closingDate, disbursementDate, currentOwner, liensAndEncumbrances, easements, titleInsuranceAmount, propertyTaxesStatus.
Payments and compliance: a projectedPayments array (one entry per payment type, each carrying principal and interest, mortgage insurance, estimated escrow, and the estimated total monthly payment), plus notarySignatureWitness, controlNumber, parish, currency, witnessName, witnessType, preparerCompany, homesteadActCompliance, farmLandActCompliance, and a handful of lien-related flags: securesLiabilities, guaranteeMortgage, collateralMortgage.
Processing metadata on every response: success, message, warnings, fallbackUsed, jobId, jobIdExt. Check success and warnings before you trust anything else in the payload; a response with fallbackUsed: true is telling you the extraction took a lower-confidence path on at least one field.
A roughly shaped (not literal) version of what comes back looks like this:
{
"documentType": "Deed of Trust",
"lenderName": "Example Lending Co.",
"loanAmount": 285000.00,
"interestRate": 6.125,
"loanTerm": 30,
"borrowerName": "Jane A. Doe",
"propertyAddress": "412 Elm Street, Springfield",
"appraisedValue": 310000.00,
"closingCosts": 6420.15,
"cashToClose": 14200.00,
"settlementAgent": "Example Title & Escrow",
"closingDate": "2026-11-14",
"comparableSales": [
{ "address": "398 Elm Street", "salePrice": 305000.00, "saleDate": "2026-08-02" }
],
"projectedPayments": [
{ "principalAndInterest": 1732.21, "mortgageInsurance": 0.00, "estimatedEscrow": 410.00, "estimatedTotalMonthlyPayment": 2142.21 }
],
"success": true,
"warnings": [],
"fallbackUsed": false,
"jobId": "a1b2c3d4-...",
"jobIdExt": "ext-..."
}
That is well over forty named fields spanning seven real-world categories a loan file actually has, returned from one call instead of built field by field in a mapping layer. One flag worth knowing before you quote a number in your own docs or a blog post: the n8n documentation page labels its own field list "30 total" while enumerating closer to fifty individual fields across those same categories. The discrepancy looks like a stale summary count on PDF4me's side rather than a real difference in what the parser returns, so pull one live response and count the fields yourself before repeating an exact number.
Two Inputs That Change What You Get Back
Two optional parameters are worth knowing before the first run. documentType lets you tell the parser what kind of mortgage paperwork it is looking at, such as "Deed of Trust," "Promissory Note," "Mortgage Agreement," or "Loan Application," which the documentation says improves extraction accuracy. That makes sense: a promissory note and a title report share almost none of the same layout, so telling the parser which one it is looking at upfront removes a layer of guessing.
customFieldKeys goes further. It is an array of additional field names, beyond the standard schema above, that this specific pipeline needs. A servicer tracking an internal risk score or a regional disclosure code the standard schema does not cover can request it through this parameter rather than waiting on a schema update. That is a meaningfully different model from a rigid, fixed-output parser, and it is one of the more useful details buried in the request body rather than the headline feature list.
The Same Action, Four Ways In, Three Ways to Hand It a File
The AI Mortgage Document Parser is not REST-only. It ships as a native action in Power Automate, a native module in Make, a native node in n8n, and a native action in Zapier, with the same required inputs and the same output schema across all four.
In Power Automate, the file comes in as binary content through Mortgage Document File Content, mapped straight from a connector like SharePoint, OneDrive, Dropbox, or an email attachment, alongside the required Mortgage Document Name, which matters for a loan-ops team that already lives in Microsoft 365 and wants the extracted fields to land directly in a SharePoint list or a Dynamics 365 record without a custom connector in between.
n8n's node, named "AI-Process Mortgage Document," is the most explicit about this choice: an Input Data Type parameter switches between Binary Data (reads from a previous node's binary output via Input Binary Field), Base64 String (via Base64 Mortgage Document Content), and URL (via Mortgage Document URL) and the filename always goes in separately through Mortgage Document Name. Make exposes the same three input modes under its own field names. In Zapier, a mortgage PDF dropped into a watched folder can turn into a fully structured record in a CRM or spreadsheet with no code written at all.
That consistency matters more than it sounds. A pipeline built once against this schema in Make can be rebuilt in n8n, or handed to a no-code teammate in Zapier, without anyone re-learning what closingCosts or settlementAgent means, because the field names do not change between platforms. The parser itself reportedly handles both scanned and digitally generated documents without altering the source file, which is the detail that actually matters if the original has to go into a compliance archive untouched. That handling claim comes from PDF4me's own integration documentation rather than from independently observed test output in this piece, so treat it as the vendor's stated behavior pending your own verification against a real scanned file.
When a Fixed Schema Is Not Enough
The Mortgage Document Parser works because its schema is already built for one document family. But not every document a developer needs to parse is a mortgage, and PDF4me's broader AI lineup reflects that: it sits alongside dedicated parsers for pay stubs, tax documents, receipts, and purchase or sales orders, each with its own fixed schema tuned to that document type. For anything outside those categories, there is a second mechanism worth knowing about: the generic AI Document Parser, which runs against an Analyzer you define yourself in the PDF4me dashboard, with your own field names, types, and descriptions, called by an AnalyzerId instead of a fixed schema. Edit that schema later and every workflow pointed at the same Analyzer ID picks up the change automatically, with no Zap or flow to rebuild. Choosing between the two is really a question of fit: reach for the Mortgage Document Parser when the document family is already mortgage paperwork and the standard fields cover the job, and reach for a self-defined Analyzer when the document type is something the standard parsers were never built for.
Worth noting for anyone looking for a plain REST endpoint: there is no dedicated pdf4me-api page for this specific parser outside the four integration platforms above (confirmed as a 404 on the expected path while researching this piece). If that changes, it will show up alongside the other pdf4me-api endpoints in the documentation.
Where to Start
If a lending, servicing, or proptech pipeline currently has a human opening PDFs to copy numbers into a loan system, this is the kind of task that disappears first. One call, one of four platforms, and a loan amount, a closing date, and a borrower name that were never typed by hand.
Website: pdf4me.com
Documentation: docs.pdf4me.com
Top comments (0)