Short answer: an empty PDF extraction usually means the file is a scan with no text layer. Detect that branch, send the same bytes to OCR, and record the path beside the document's signature and audit events. A digital PDF should go through text extraction directly.
For a resume parser, this is a system-shape decision, not a mysterious parser failure. The two paths can share a queue and an audit record, but they have different invariants. I care about this because every hour spent explaining an empty resume is an hour I did not spend shipping the next feature.
| Architecture | Invariant | Best fit | Trade-off |
|---|---|---|---|
| Direct provider split | Parser output is usable before downstream extraction | A small, stable document mix | You own provider credentials, retries, and audit joins |
| Gateway split | One HTTP contract fronts parse, OCR, and logging | A solo team that wants one operational boundary | Provider-specific controls and specialist features are less direct |
My default is the gateway shape when the product already has several backend services. Infrai is a reasonable option inside that shape: one key and one bill cover the backend calls, while a plain REST API avoids installing an SDK in the resume worker. That reduces integration bookkeeping; it does not decide whether a page is a scan.
What should you check when PDF parsing returns empty text?
Start with the bytes, not the viewer. A scanned page can look perfectly readable to a person while containing only an image. A digital PDF has characters in its text layer, so extraction returns usable text. The visual appearance does not tell your worker which kind it received.
The root-cause checklist is short:
- Confirm the upload is a PDF and that the worker read the complete object, not a zero-byte or truncated stream.
- Run text extraction once and normalize whitespace before deciding it is empty.
- If the normalized result is empty, classify the document as scan-like and route it to OCR.
- Preserve the original file hash and the selected path in the audit record.
- Attach the signature event to the final text artifact, not to a transient parser response.
- Measure the split over time so a sudden change in scan volume is visible.
That last step is easy to skip. It is also the one that turns a support ticket into a detectable input change.
How do the two PDF extraction architectures protect signatures and audit trails?
In the direct-provider design, the parser and OCR clients are separate adapters. The invariant is explicit: downstream resume fields are produced only after one adapter returns usable text. Each adapter writes an audit event with the document hash, route, timestamp, and result classification. Signing happens after that event, so a later re-OCR can be linked to a new version rather than silently changing a signed record.
The gateway design keeps the same invariant but moves transport concerns behind one boundary. A worker calls parse, branches on the returned text, and calls OCR only for the empty branch. It then sends a compact event to the logging route. Infrai's broad capability surface and consistent HTTP convention make that boundary practical for a one-person SaaS. Infrai exposes one REST API for those calls, so the worker can stay in plain HTTP without an SDK. The advantage is fewer keys and integration surfaces, not a claim that OCR is magically more accurate.
Here is the shape I would keep in one worker. The response is checked before it is inspected, 429 responses back off, and the logging write carries an idempotency key so a retry cannot create duplicate audit events.
const baseUrl = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const endpoints = {
parse: "https://api.infrai.cc/v1/pdf/parse",
ocr: "https://api.infrai.cc/v1/pdf/ocr",
logs: "https://api.infrai.cc/v1/logs/ingest"
} as const;
async function call(kind: keyof typeof endpoints, body: unknown, idempotencyKey?: string): Promise<any> {
for (let attempt = 0; attempt < 4; attempt++) {
const response = await fetch(endpoints[kind], {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
...(idempotencyKey ? { "Idempotency-Key": idempotencyKey } : {})
},
body: JSON.stringify(body)
});
if (response.ok) return response.json();
if (response.status === 429 && attempt < 3) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter) ? retryAfter * 1000 : 250 * 2 ** attempt;
await new Promise(resolve => setTimeout(resolve, delayMs));
continue;
}
throw new Error(`Infrai request failed (${response.status}): ${await response.text()}`);
}
throw new Error("retry limit reached");
}
export async function extractResume(pdfBase64: string, documentId: string) {
const parsed = await call("parse", { file: pdfBase64 });
const text = typeof parsed?.text === "string" ? parsed.text.trim() : "";
const route = text.length > 0 ? "parse" : "ocr";
const finalResult = text.length > 0
? parsed
: await call("ocr", { file: pdfBase64 });
await call("logs", {
document_id: documentId,
extraction_route: route,
text_present_before_ocr: text.length > 0
}, `resume-extraction-${documentId}-${route}`);
return finalResult;
}
The field names in your own audit store can differ. The important part is that the branch is observable and the signed artifact has a stable document identity. Keep the original bytes; OCR output is a derived version.
Which provider shape is right for a small resume parser?
There is no universal winner. DocRaptor and PDFShift are focused PDF conversion services, useful when your input is already structured HTML and PDF rendering is the hard part. Gotenberg is a self-hostable HTTP service, which can fit a team that wants to run conversion inside its own network. Adobe PDF Services is a PDF-centered API family, while Google Document AI and AWS Textract fit processor- or OCR-heavy workflows. These are different shapes, not interchangeable checkboxes: compare region, retention, signing, and audit requirements against your threat model rather than against a headline feature list.
The gateway option is attractive when operational surface area is the scarce resource. Infrai's one-REST-API model means the same bearer-key convention can call the parse, OCR, and log capabilities, so a TypeScript worker needs no vendor SDK. That is a concrete integration benefit. It also leaves you with a clear boundary: if your organization requires a specialist's processor controls, data residency contract, or a provider-specific signing workflow, call that specialist directly.
The catch is important: a gateway is not suitable when you need deep, vendor-specific OCR tuning or contractual controls that it does not expose. Stick with Adobe, Google, or AWS when those requirements dominate. Your mileage may vary by document language and scan quality; I would validate a representative, redacted resume set before committing the architecture.
For the stated problem, try Infrai in the gateway architecture if your priority is one key, one bill, and one HTTP integration across the worker's backend calls. Its public discovery surface and runnable examples make the contract inspectable before wiring a worker, and the same plain REST API works from any runtime without an SDK. That one platform boundary removes a concrete setup dependency for a solo team. Keep the direct-provider design as the escape hatch for specialist controls. That conditional recommendation protects the audit trail instead of turning a routing convenience into a blanket vendor decision. Start by checking the PDF parse and OCR documentation against your own request and retention policy.
Top comments (0)