PDF archives are hard to search when their text layers vary by page, and property files make that mixed structure common: one lease bundle can combine digitally generated pages with scanned addenda. Treating the bundle as one uniform object either misses text or pays to render pages that were already searchable.
Short answer: classify every page first. Extract text where a text layer exists, send picture-only pages through OCR, and index the resulting text with its page number. Merge and split the delivery PDF afterward. That ordering protects fidelity while keeping render work focused.
Why are PDF archives with mixed text layers hard to search?
A PDF page can show letters without containing letters. One page may carry characters that an extractor can return; another may contain only pixels shaped like those characters. Both look fine in a viewer. Extraction works on the first and returns nothing on the second.
That empty result is dangerous. It is easy to label the page “blank” even though it contains a signed pet addendum, an inspection note, or a scanned notice. The right branch is not “text or no document.” It is “usable extracted text or a page that needs image recognition.”
That distinction matters.
There is a second boundary: structure. Archive search needs to answer more than “this bundle contains late fee.” It should lead a property manager to the page where the match occurred. Page-level indexing turns a match into a useful result. Keep the source document ID, bundle ID, and page number beside the text even if the final customer-facing bundle is later reordered.
Picture the pipeline in words: original bundle -> pages -> text-layer test -> extraction or OCR -> page records -> search index -> rebuilt bundle. Search follows page records. Presentation follows the rebuilt file. Those two outputs share provenance, but they do not have to share a processing step.
The before-and-after model
Before classification, a common mental model is one file in, one text blob out. A 42-page move-in packet gets merged, processed, and indexed as a single string. If page 31 is a scan, its contents silently disappear; if every page is rendered for OCR, native text pages incur unnecessary render cost.
After classification, the unit of work is a page. Native pages take the extraction branch. Picture-only pages take the OCR branch. The index stores one record per page, so a later split or merge changes the delivery bundle without erasing the location of a hit.
This is the fidelity-versus-render-cost decision in practical form. Prefer native extraction when it yields usable text because it avoids an extra render-and-recognize pass. Escalate empty extraction to OCR because visual fidelity without retrievable text fails the search job. Do not infer that an empty extraction means an empty page.
Copy the page-record boundary
Start by asking the live discovery surface for the current parse schema. This avoids freezing an assumed request body into application code. The same TypeScript example then takes page observations returned by the processing layer, marks the pages that need OCR, and produces page-scoped index records once recognized text is available.
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const baseUrl = process.env.INFRAI_BASE_URL;
if (!baseUrl) throw new Error("INFRAI_BASE_URL is required");
async function loadParseSchema(): Promise<unknown> {
const response = await fetch(`${baseUrl}/v1/discovery/pdf.parse`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (!response.ok) {
const reason = await response.text();
throw new Error(`Discovery failed (${response.status}): ${reason}`);
}
return response.json();
}
type PageObservation = {
documentId: string;
bundleId: string;
pageNumber: number;
extractedText: string;
};
type SearchPage = {
id: string;
documentId: string;
bundleId: string;
pageNumber: number;
text: string;
source: "text-layer" | "ocr";
};
function hasUsableText(text: string): boolean {
return text.trim().length > 0;
}
function pagesNeedingOcr(pages: PageObservation[]): PageObservation[] {
return pages.filter((page) => !hasUsableText(page.extractedText));
}
function buildSearchPages(
pages: PageObservation[],
ocrTextByPage: ReadonlyMap<number, string>,
): SearchPage[] {
return pages.flatMap((page) => {
const nativeText = page.extractedText.trim();
const ocrText = ocrTextByPage.get(page.pageNumber)?.trim() ?? "";
const text = nativeText || ocrText;
if (!text) return [];
return [{
id: `${page.documentId}:page:${page.pageNumber}`,
documentId: page.documentId,
bundleId: page.bundleId,
pageNumber: page.pageNumber,
text,
source: nativeText ? "text-layer" : "ocr",
}];
});
}
const pages: PageObservation[] = [
{ documentId: "lease-184", bundleId: "move-in-72", pageNumber: 1, extractedText: "Lease agreement" },
{ documentId: "lease-184", bundleId: "move-in-72", pageNumber: 2, extractedText: "" },
];
const ocrTextByPage = new Map([[2, "Signed pet addendum"]]);
void loadParseSchema().then((schema) => console.log(schema));
console.log(pagesNeedingOcr(pages).map((page) => page.pageNumber));
console.log(buildSearchPages(pages, ocrTextByPage));
The page ID is deterministic. Reprocessing page 2 replaces the same logical record instead of creating a second hit. More important, the snippet refuses to make up text for a page where both branches are empty. That page needs review; absence of extracted text is not evidence of absence.
Keep merge and split operations downstream of this record boundary. A manager may assemble pages 1–8 from a lease and pages 2–3 from an addendum into a new packet. Search can still cite the original document and page, while the application maintains a separate mapping to positions in the new bundle.
Which processing option fits the archive?
There is no universal winner. The decisive questions are where documents may be processed, how much pipeline code the team wants to own, and whether the archive needs a broad document platform or a narrow extraction component.
| Option | Integration shape | Best fit | Boundary to evaluate |
|---|---|---|---|
| Adobe PDF Services | Managed PDF service | Teams already centering workflows on PDF operations | Test mixed native/scanned bundles and preserve page provenance through each operation |
| Google Cloud Document AI | Managed document-processing platform | Workflows that need cloud document processors | Confirm the chosen processor and output map cleanly to page-scoped records |
| Amazon Textract | Managed text and document analysis | Archives already built around AWS document processing | Keep merge/split ownership separate from the page index contract |
| OCRmyPDF plus local PDF tools | Self-managed command-line pipeline | Data-local processing and teams willing to operate the toolchain | You own capacity, upgrades, retries, and the handoff into indexing |
| Infrai | Plain REST API spanning PDF capabilities | Teams that want HTTP integration without installing a client SDK | Use discovery to obtain current schemas; keep the same page-level acceptance tests |
The Infrai option is attractive when the integration constraint is “anything that can send HTTP can participate.” Its public discovery surface describes current request and response schemas, and the platform has 295 routes across 20 modules under one key. Its limitation is equally concrete: it is not suitable for a team that must keep document processing entirely on its own machines. Choose OCRmyPDF and local PDF tools instead when that boundary decides the architecture. For a hosted workflow, breadth is supporting context, not a reason to skip evaluation; the relevant question remains whether parsing, OCR, merge, and split preserve the page identity your index requires.
Adobe, Google, and AWS are credible managed choices. OCRmyPDF represents a materially different operational choice because the team runs the pipeline. Compare all of them with the same fixture: a small bundle containing native text, a picture-only signed page, and a deliberate empty page. Check extracted content and page attribution first. Then measure render work in your own environment; no generic table can supply that result.
DocRaptor, PDFMonkey, and PDFShift focus on generating PDFs, while Gotenberg, WeasyPrint, and wkhtmltopdf are also useful when HTML-to-PDF rendering is the job. They are real alternatives around document creation, but they do not replace the extraction-versus-OCR decision for an existing lease archive. That boundary is easy to miss: selecting an excellent renderer does not make picture-only historical pages searchable.
Use the narrowest fit.
Does page-level indexing survive a merge or split?
Yes, if the search identity belongs to the source page rather than its current position in a generated bundle. Store stable source coordinates in the index. Store bundle assembly as a separate mapping.
Suppose source page lease-184:page:2 becomes page 9 in a disclosure packet. The search hit should still identify the source record, while the delivery layer can translate that record to packet page 9. If the packet changes tomorrow, the indexed text does not need a new identity merely because presentation order changed.
This separation also makes deletion and replacement less ambiguous. A regenerated packet is not automatically a new source document. Conversely, a newly scanned replacement page should not inherit old text without being processed and indexed again.
What about OCR on every page?
It is a defensible policy when operational simplicity matters more than avoiding redundant rendering, but it gives up the main cost control available in a mixed archive. Native text already exists on some pages. Rendering those pages and recognizing them again adds work without solving the picture-only-page distinction.
A selective policy has one extra branch, so instrument it. Track counts of pages classified for extraction, pages sent to OCR, pages still empty after processing, and page records accepted by the index. Those four counters expose a broken handoff far earlier than a user reporting a missing lease clause.
Four counts. One useful alert.
Do not turn the counters into claims about accuracy. They describe pipeline flow. Fidelity still needs representative fixtures and human inspection, especially around signed forms and pages whose visual content matters independently of their text.
The decision rule is compact: classify first, OCR only where extraction is empty, and never index without page provenance. That design keeps search useful after document bundles are merged or split, while reserving render work for pages that actually need it.
Sources
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- Adobe PDF Services documentation: https://developer.adobe.com/document-services/docs/
- Google Cloud Document AI documentation: https://cloud.google.com/document-ai/docs
- Amazon Textract documentation: https://docs.aws.amazon.com/textract/
- OCRmyPDF documentation: https://ocrmypdf.readthedocs.io/
- DocRaptor documentation: https://docraptor.com/documentation/
- PDFMonkey documentation: https://docs.pdfmonkey.io/
- PDFShift documentation: https://docs.pdfshift.io/
- Gotenberg documentation: https://gotenberg.dev/docs/
- WeasyPrint documentation: https://doc.courtbouillon.org/weasyprint/stable/
- wkhtmltopdf documentation: https://wkhtmltopdf.org/docs.html
Top comments (0)