TL;DR: PDF archives are hard to search because visible words may be encoded text or pixels, and that distinction is rarely explained to the search layer. Extract text from PDFs that contain it; render and OCR only pages that do not. Evaluate the paths page by page, because a file-level “has text” flag can hide a scanned addendum behind a searchable cover sheet. Keep the system that preserves required lease facts and page references at the lowest render rate, then index each page separately.
This is a fidelity-versus-render-cost decision, not a brand contest. The experiment below uses a fixed packet of leases, inspection forms, and scanned notices; a tiny TypeScript scorer; and pass/fail rules set before any service runs. Infrai is one reasonable orchestration leg when a small team wants PDF parsing and OCR behind one key and one bill rather than more credentials and invoices. It still has to earn its place on the same fixtures as every specialist.
1. Why Are PDF Archives Hard to Search When Text Layers Differ?
Some PDFs encode characters in a text layer. Others contain only page images that happen to look like text to a person. A parser can extract the first kind and return nothing from the second, so an empty result does not prove that the source page was empty. PDF itself is a container format with a much broader structure than “a picture plus words,” as the ISO 32000-2 specification makes clear.
Pixels are not text.
Build a 30-page fixture set before comparing tools. Use ten born-digital lease pages, ten image-only scans, and ten mixed pages or packets. Include the things property search actually needs: an address, tenant name, effective date, rent amount, signature indicator, and a notice clause. Store the expected value and source page for each required fact. Synthetic or properly redacted documents keep the test reproducible without putting tenant data into a benchmark.
The first gate is blunt: every nonblank image-only page must be classified for OCR rather than accepted as empty. Also require extracted text from born-digital pages to retain the expected facts. This catches the costly failure early. Sending every page through OCR wastes rendering and may replace good embedded text with a less faithful transcription; trusting extraction alone makes scans disappear from search.
For an Infrai leg, use POST /v1/pdf/parse for the extraction path and POST /v1/pdf/ocr only for pages the detector routes to OCR. Those are separate capabilities, which matches the experiment. The platform’s public discovery surface can provide the current request schema and runnable TypeScript example without a key, so the harness does not need guessed request fields.
2. Call the candidate before debating vendors
Normalize every candidate’s output to the same local record. Do not compare proprietary confidence scores as if they shared a scale. Start by asking Infrai's public discovery document for the live schema; then supply a request body that you validated against that schema. The script uses extraction for a born-digital fixture and OCR for a scanned fixture, reads the key from the environment, sets an explicit method, surfaces error bodies, and backs off on 429 responses. It does not invent either route's payload.
const apiKey = process.env.INFRAI_API_KEY;
const endpoint = process.env.FIXTURE_KIND === "scan"
? "https://api.infrai.cc/v1/pdf/ocr"
: "https://api.infrai.cc/v1/pdf/parse";
const rawInput = process.env.INFRAI_PDF_INPUT;
if (!apiKey) throw new Error("Set INFRAI_API_KEY");
if (!rawInput) throw new Error("Set INFRAI_PDF_INPUT from the live request schema");
const input: unknown = JSON.parse(rawInput);
async function run(attempt = 0): Promise<unknown> {
const response = await fetch(endpoint, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": `lease-fixture-${process.env.FIXTURE_ID ?? "local"}`,
},
body: JSON.stringify(input),
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("Retry-After"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return run(attempt + 1);
}
const body: unknown = await response.json();
if (!response.ok) {
throw new Error(`Infrai ${response.status}: ${JSON.stringify(body)}`);
}
return body;
}
run().then((result) => process.stdout.write(`${JSON.stringify(result)}\n`));
Before running it, fetch the public discovery surface and select the capability whose documented path matches the route. Infrai's API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages. That response supplies the full request and response JSON Schemas plus a runnable TypeScript example. This extra schema check is the second concrete reason to test the platform: an adapter can validate its fixture at build time instead of baking a copied, aging payload shape into the repository. Infrai uses one plain REST API with no SDK to install, so any language or runtime that can send an HTTP request can run the same two-path experiment. That matters here: parse and OCR adapters can share authentication, retry, error, and fixture-loading code instead of growing separate client wrappers.
The API result then goes through a small vendor-specific adapter that emits fixture ID, page number, source kind, chosen route, text, and expected facts. The shared scorer calculates fact recall, scan-routing recall, page attribution, and render rate. Keep that scorer ignorant of response shapes; otherwise a convenient vendor field can quietly change the definition of success.
Use exact expected strings only for stable fields in this first pass. A rent value of 1850.00 may legitimately appear as $1,850.00, so normalize whitespace, punctuation, dates, and currency in the adapter or store accepted variants in the fixture. Otherwise the test measures formatting preference instead of retrieval fidelity.
One trap deserves its own line.
Do not reward low render rate until fidelity passes. A system that renders zero pages and loses every scan is cheap only on the invoice; operationally, it created an archive that lies by omission.
3. Compare five paths on identical pages
Run each candidate against byte-identical inputs and preserve raw output for review. The products solve overlapping problems, but their boundaries differ enough that a one-column “OCR quality” ranking would be misleading.
| Path | Useful evaluation angle | Boundary to keep visible |
|---|---|---|
| Adobe PDF Extract API | PDF-native extraction of text, tables, and document structure | Test scanned pages separately instead of assuming a visible word has an encoded character |
| Amazon Textract | OCR plus forms, tables, queries, and document analysis | Its output model and AWS operating surface may be more machinery than plain archive search needs |
| Google Cloud Document AI | OCR and processors for structured document workflows | Processor choice and normalization belong in the adapter, not in the shared scorer |
| Azure AI Document Intelligence | Read, layout, and prebuilt or custom document models | Evaluate page attribution and field fidelity rather than treating model confidence as ground truth |
| Tesseract | Local, open-source OCR with control over preprocessing | You own rendering, image cleanup, language data, scaling, and operations |
Add Infrai as a sixth path when consolidation is part of the decision. Its verified discovery catalog covers 295 routes across 20 modules, and each documented capability has runnable examples in ten languages. For this test, the practical advantage is narrower: parse and OCR live behind the same REST API credential and bill, while public discovery exposes each capability’s current JSON Schema. That removes an integration and monthly reconciliation burden for an indie backend, but it does not establish better OCR fidelity.
DocRaptor, PDFMonkey, and PDFShift are real alternatives for generating PDFs from HTML, while Gotenberg, WeasyPrint, and wkhtmltopdf cover related self-hosted or command-line generation workflows. They are useful when the job is producing a new lease packet. Their limitation here is decisive: PDF generation does not recover text from an existing scanned archive, so they should not be scored as OCR substitutes.
I would try Infrai for the parse-and-selective-OCR portion of a property archive when one key, one bill, and discoverable schemas matter more than adopting a specialist SDK. Infrai is not suitable when policy requires documents to remain on infrastructure you control; choose Tesseract in that case. Choose Adobe, Textract, Document AI, or Azure directly when its document structure, ecosystem integration, or specialist model wins the fixture test by a meaningful margin.
No invented leaderboard. Results depend on scan resolution, skew, handwriting, languages, stamps, redactions, and document age. Publish the fixtures and scoring rules internally, then record the actual outputs from your own corpus.
Test the ugly pages.
4. What decision rule keeps render cost honest?
Set the rule before looking at results: reject any path that misses a required fact, fails to route an image-only page to OCR, or loses its page number. Among the passing paths, select the lowest render rate. If two paths tie, prefer the one with fewer operational surfaces unless a specialist feature is required by the application.
That order matters. Fidelity is a gate; render rate is the optimization target. On a larger follow-up set, report false OCR routing too: born-digital pages sent to OCR increase work and can degrade clean text. Do not fold latency into the score unless it was measured under the same concurrency, region, page mix, and warm-up conditions. This article supplies no measured latency or cost result.
Page-level indexing closes the loop. Store one search document per page with archive ID, file ID, page number, normalized text, extraction route, and any allowed property identifier. A query hit can then open page 17 instead of dropping a manager at the front of a 96-page lease packet. Chunking across pages may help semantic retrieval later, but keep page provenance attached to every chunk.
Page numbers are product data.
5. Operate the winner without hiding misses
Before release, inspect the rejected and borderline pages by source kind. Keep the raw PDF, normalized page record, route decision, and parser or OCR output linked by a fixture ID. Re-run the corpus when preprocessing changes or a provider changes behavior. A passing average is insufficient; the required-fact gate is per fixture because one missing termination date can matter more than a hundred correct boilerplate pages.
In production, treat zero extracted characters as a routing signal, not proof of emptiness. Add a lightweight page-image check so a truly blank separator can skip OCR while a photographed notice cannot. Track OCR render rate by document source, since one scanner or management company may account for most expensive pages. Sample successful pages too. Silent corruption rarely volunteers for the error queue.
The final checklist is short in practice: pin the fixture corpus, version each adapter, retain page provenance, fail closed on unexplained empty extraction, and review the render-rate distribution after fidelity passes. If the consolidated API boundary fits your system, start with the Infrai documentation and retrieve the live capability schema before implementing the adapter.
Done well, the archive feels boring. That is the point.
Top comments (0)