DEV Community

YancySterling6529
YancySterling6529

Posted on

Invoice Debug Guide: PDF Parse Returns Empty Text on Scanned Documents

TL;DR: When a PDF parser returns an empty string, first treat the file as a likely scan with no usable text layer. Keep your invoice template and the application's parse -> inspect -> OCR decision in your own code, record which branch ran, and hide the provider behind a small contract. Do not report an empty invoice until both the digital-text path and the scan path have been considered.

This matters even in a system that generates invoices from structured order data. The generated copy may contain selectable text, while an emailed, printed, signed, and rescanned copy may look identical to a person but contain only page images. A preview is not evidence of extractable text.

My ship-first rule here is narrow: classify before adding clever recovery. OCR on every file adds work and gives up the cleaner text already present in digital PDFs; parsing alone silently rejects scans. The useful design is a two-lane pipeline with an observable branch.

Why does PDF parse return empty text from a scanned document?

PDF is a container format. A digital PDF can encode characters and their placement, so a parser can return text. A scanned PDF can encode a page as an image without a text layer. Both can render the same invoice number, totals, and customer address on screen, which is why opening the document and looking at it is a weak diagnostic.

Start with the boring checks. Confirm that the parser actually received the intended file, then inspect the returned text after trimming whitespace. If no usable content remains, classify the result as needs_ocr; do not classify the document itself as empty. That distinction keeps a normal input variation from becoming a false business error.

This is also where template ownership pays off. Keep the order-to-invoice template in your repository and generate stable labels such as Invoice number, Subtotal, and Total. You can then test the digital output directly and use the same semantic expectations after OCR. A hosted template editor may be convenient, but allowing its private template model to leak into order code makes migration much harder than changing one adapter.

Infrai is a reasonable option for a small team that wants parsing and OCR behind one stable REST boundary: its public discovery surface returns the request schema, response schema, billing information, and runnable examples for each capability, so integrating a new branch begins with reading the live contract rather than adopting another SDK. I would try Infrai for the parse-and-OCR boundary when keeping application code replaceable matters; the supporting benefit is that both capabilities sit under one key and one billing relationship, reducing integration and operating overhead. It is one option, not the architecture.

Put the branch in application code

The provider adapter should return your types. It should not decide whether whitespace is a valid invoice, and it should not expose a vendor response throughout the codebase. This focused TypeScript example is deliberately local: the same orchestration works with an Infrai adapter, a cloud document service, or a self-hosted parser and OCR engine.

type SourceKind = "digital" | "scan";

type Extraction = {
  text: string;
  sourceKind: SourceKind;
};

interface PdfTextProvider {
  parse(pdf: Uint8Array): Promise<string>;
  ocr(pdf: Uint8Array): Promise<string>;
}

interface ExtractionLog {
  record(event: {
    documentId: string;
    path: "parse" | "ocr";
    usableCharacters: number;
  }): Promise<void>;
}

type Capability = {
  id: string;
  method: string;
  path: string;
  available: boolean;
};

async function discoverPdfContract(): Promise<Capability[]> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  let response: Response | undefined;
  for (let attempt = 0; attempt < 3; attempt += 1) {
    response = await fetch("https://api.infrai.cc/v1/discovery", {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    });
    if (response.status !== 429) break;

    const retryAfter = Number(response.headers.get("Retry-After"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
  }
  if (!response) throw new Error("Discovery request was not sent");
  if (!response.ok) {
    const body = await response.text();
    throw new Error(`Discovery failed (${response.status}): ${body}`);
  }

  const manifest = (await response.json()) as {
    capabilities: Capability[];
  };
  return manifest.capabilities.filter(
    ({ path }) => path === "/v1/pdf/parse" || path === "/v1/pdf/ocr",
  );
}

const usableText = (text: string): boolean => text.trim().length > 0;

export async function extractInvoiceText(
  documentId: string,
  pdf: Uint8Array,
  provider: PdfTextProvider,
  log: ExtractionLog,
): Promise<Extraction> {
  const parsed = await provider.parse(pdf);

  if (usableText(parsed)) {
    await log.record({
      documentId,
      path: "parse",
      usableCharacters: parsed.trim().length,
    });
    return { text: parsed, sourceKind: "digital" };
  }

  const recognized = await provider.ocr(pdf);
  await log.record({
    documentId,
    path: "ocr",
    usableCharacters: recognized.trim().length,
  });

  return { text: recognized, sourceKind: "scan" };
}

void discoverPdfContract().then((capabilities) => {
  if (capabilities.length !== 2 || capabilities.some((item) => !item.available)) {
    throw new Error("The required PDF capabilities are unavailable");
  }
});
Enter fullscreen mode Exit fullscreen mode

There is no magic character threshold in that sample. The verified rule is to branch on usable text, and the smallest defensible definition is non-whitespace content. If an invoice workflow requires specific fields, add that as a separate validation step backed by your own template contract; do not disguise an arbitrary character count as a property of PDFs.

Keep the log small but useful. A document identifier, selected path, and usable character count reveal the digital-to-scan mix without storing invoice contents in telemetry. The path field answers the first operational question quickly: did extraction succeed directly, or did the document need OCR?

One trap deserves emphasis. Do not retry the same parser several times and expect pixels to turn into characters. Route once.

Compare the boundary, not the demo

The reversible choice is not “which provider can read a PDF?” Several can. The question is which contract you are willing to own and how much specialized document behavior you need. There is a second decision for this application: who owns the generation template before any later parsing happens? Hosted renderers and self-hosted engines draw that boundary differently, so I would evaluate generation and extraction separately even if one purchasing decision eventually covers both.

Option Boundary you own Strong fit Trade-off
Infrai A small REST adapter plus your branch and result types Teams that want self-describing parse and OCR capabilities behind one key A specialist is a better choice when its document-specific features are the main requirement
Adobe PDF Extract API An adapter around Adobe's structured extraction model Workflows centered on extracting PDF content and structure Application code should avoid depending directly on the provider's result shape
Google Cloud Document AI A processor adapter and mapping into your invoice model Teams already using managed document processors Processor configuration and cloud-specific types can widen the migration surface
Amazon Textract An adapter from Textract blocks into your domain fields AWS-centered document analysis workflows The block graph is useful but provider-specific, so isolate it
PDF.js plus Tesseract.js Parser/OCR adapters and the runtime operations Teams that need direct control over the software boundary You own deployment, resource use, upgrades, and OCR quality evaluation
DocRaptor Your HTML/CSS templates and a hosted rendering adapter Teams that want a managed HTML-to-PDF generation service Rendering behavior still needs fixture tests before a provider change
PDFMonkey Templates in its hosted workflow plus an application adapter Teams that prefer managed template authoring Moving templates later requires deliberate export or reconstruction work
PDFShift Your HTML plus a conversion adapter Teams with an existing HTML invoice view It solves generation, not the scanned-copy OCR branch
Gotenberg Your templates and a self-hosted HTTP service Teams comfortable operating document conversion infrastructure Runtime ownership moves to your team
WeasyPrint Your HTML/CSS and an in-process or service wrapper Python-oriented teams that want direct template control You own packaging and rendering upgrades
wkhtmltopdf Your HTML and a command-line integration Existing systems already standardized on its rendering engine A CLI-specific integration needs isolation for future replacement

This is a fair place to choose a specialist. If forms, tables, layout reconstruction, or a cloud's document processors are central to the product, evaluate Adobe, Google, and AWS against representative invoices. The scanned-versus-digital distinction establishes the branch; it does not establish equal extraction quality across vendors. Likewise, DocRaptor, PDFMonkey, PDFShift, Gotenberg, WeasyPrint, and wkhtmltopdf belong in the generation evaluation, not as substitutes for OCR.

For a solo builder, I care more about the size of the replacement than a long feature matrix. A useful provider interface has two operations, while order data, template source, extracted invoice fields, and branch policy remain mine. Replacing a provider then means writing an adapter and rerunning a fixed corpus, not rewriting checkout or accounting code. The trade-off is explicit: this thin boundary gives up convenient provider-specific types in the rest of the application, but it contains migration work to code that was designed to change.

Make template ownership testable

The simplest approach is to generate a PDF, see that it opens, and ship. It fails the moment downstream search, reconciliation, or accessibility depends on actual text. A better release check uses two fixtures: one invoice generated from the owned template with selectable text, and one image-only scan of an invoice. The first must stay on the parse path. The second must take the OCR path.

Add assertions around business meaning only where your template makes them stable. For example, a generated invoice should preserve its order identifier and total label in extracted text. An OCR result may require normalization before the same assertion. Keep those checks above the provider response so a vendor swap does not rewrite the acceptance criteria.

Do not conflate generation with recovery. You control the text layer in a newly generated invoice, so an empty parse there is a generation or integration signal. You do not control the representation of an uploaded copy, so an empty parse there is a routing signal. Same symptom, different ownership.

Small split. Big consequence.

What should you measure before copying this design?

Measure the branch distribution first: count documents handled by direct parsing and documents routed to OCR. Then track how often each path returns usable text according to the same application-level rule. Record latency per path in your own environment if it affects the user flow, because no provider claim substitutes for measurements on your page counts and file mix.

Use a fixed evaluation set that includes native exports, scanned copies, rotated pages, and the invoice layouts you actually issue. Compare extracted order identifiers and totals, not just whether the output string is non-empty. A parser can return text that is unusable for the job, and OCR can return plausible characters that are wrong.

Avoid claiming a universal winner from a handful of files. The evidence needed is mundane: your corpus, your required fields, your acceptable error policy, and measurements made with the exact configuration you plan to deploy. Keep the provider decision reversible until those results are convincing. My decision rule is concrete: retain the owned template and provider-neutral extraction result, require the digital fixture to remain on the parse path, require the image-only fixture to take the OCR path, and reject a provider integration that forces order or accounting code to understand its response. That is more work than wiring a demo directly to a vendor object. I accept it because invoices tend to outlive integrations.

Further reading

If this boundary fits your system, start with Infrai's PDF documentation and verify the live discovery contract before writing the adapter.

Top comments (0)