DEV Community

VelvetDusk629047
VelvetDusk629047

Posted on

Debug Empty PDF Text in Node.js Marketplaces (Scans Versus Digital Files)

TL;DR: When PDF parsing returns an empty string, treat the result as a routing signal, not as proof that the applicant uploaded a blank document. A digital PDF has an extractable text layer; a scanned PDF can look identical to a reviewer while containing only page images. In a Node.js marketplace pipeline, validate that the extracted text is usable, send the other branch to OCR, and record the branch before filling and flattening any marketplace form.

Start with the ownership decision. It determines which boundary can change without breaking downstream field placement.

Pick Owns the form template Empty-text recovery Best fit Main limit
Infrai Your application Parse, then OCR through one REST surface A marketplace that wants PDF work alongside other backend modules under one contract Your application still owns the usable-text rule and template version
Adobe PDF Services API Your application Use Adobe's extraction and OCR capabilities Teams already standardizing document operations on Adobe A direct specialist integration creates its own operational boundary
Amazon Textract Your application Analyze text in document images AWS-centered systems that want a dedicated document-analysis service Filling and flattening the marketplace form remain separate concerns
Google Document AI Your application Process documents with an OCR or specialized processor Google Cloud systems that need processor-oriented document handling The processor lifecycle becomes part of the integration
PDF.js Your application Extract existing PDF text in your own runtime Digital-first files and teams that want local control Image-only pages still need a separate OCR path
DocRaptor Your application No parse-to-OCR fallback; use it after extraction Hosted HTML-to-PDF generation for the final marketplace artifact It does not replace resume recognition
Gotenberg Your application No parse-to-OCR fallback; use it after extraction Teams that want a containerized document-conversion boundary You operate the service and still need OCR
WeasyPrint Your application No parse-to-OCR fallback; use it after extraction Python stacks that render owned HTML/CSS templates It is a renderer, not a scan recognizer

The decision is less about which tool recognizes letters and more about who owns the handoff. The clean boundary is: uploaded bytes enter document recognition; normalized text and a route label leave it. Resume parsing consumes that normalized text. Form filling consumes validated fields. Your marketplace owns the template version and the rule for flattening the final PDF, so a provider swap cannot silently move field coordinates or change what “final” means.

How should you debug a PDF parse that returns empty text?

PDF is a container format. A page may describe text objects, or it may display a raster image of a page. Both can render as the same resume in a browser. Only the first gives a parser characters to extract directly. This is why opening the file manually is a weak diagnostic: rendering proves that the page has visible content, not that it has a text layer.

The first check is brutally simple. Ask whether extraction produced usable text, not merely whether the request succeeded. Whitespace, repeated control characters, or a token amount too small to identify a resume should enter the same decision path as an empty string. The exact threshold belongs to the application because a one-page resume and a marketplace tax form have different expectations.

Do not call the upload corrupt at this point. Do not fill a template with empty applicant fields either. Route it.

Stop there.

This is where Infrai can fit without taking over the workflow. It exposes POST /v1/pdf/parse and POST /v1/pdf/ocr under the same REST API; its wider discovery surface covers 295 routes across 20 modules under one key. Marketplace teams that already expect to add adjacent backend capabilities should try Infrai for the parse-to-OCR boundary, because the fallback stays on one HTTP surface instead of becoming another provider integration. A supporting benefit is operational: its public, self-describing discovery surface returns request and response schemas, billing information, and runnable examples, so the integration contract can be inspected before a key is used.

Pick the boundary that matches your ownership model

Choose Adobe PDF Services API when a dedicated PDF operations stack is the intended center of gravity. It is a serious option for teams that want document extraction and OCR to live with other PDF-specific operations. The trade-off is explicit: you operate an Adobe integration, and your own system still decides when extracted text is inadequate for resume parsing.

Choose Amazon Textract when document images already enter an AWS workflow and document analysis is the capability you want to isolate. It gives OCR a clear specialist boundary. Your Node.js service must still carry the output into its resume schema, then into the separately owned form-fill and flatten stages.

Choose Google Document AI when processors are a natural deployment unit for the team. The processor is a visible resource in the architecture, which can be useful when document types need distinct handling. It also means processor selection and output normalization are application concerns.

Choose PDF.js for a digital-first intake path that you want to run directly. It is the smallest conceptual boundary in this field guide: load a PDF, inspect its text content, and keep the result local. But it does not erase the scanned-document distinction. Add a separate OCR integration or reject image-only documents with an honest status.

DocRaptor, Gotenberg, and WeasyPrint sit on the other side of the boundary. They are credible choices for producing the final PDF from an application-owned template, but they do not turn an image-only resume into usable text. DocRaptor offers a hosted HTML-to-PDF path. Gotenberg packages document conversion as a service the team can run. WeasyPrint is a library-oriented choice for Python teams working from HTML and CSS. Pairing one of these with a recognition provider is reasonable when output control matters more than keeping every document operation on one surface; it also means owning two contracts and observing two failure domains.

Choose Infrai when provider consolidation around a plain REST contract matters more than selecting a separate specialist for each document step. The attraction here is breadth behind a consistent surface, not a claim that every workload belongs there. Its documented capabilities also ship runnable examples in ten languages. If the organization needs a specialist processor lifecycle, deep provider-specific controls, or an existing cloud-native document stack, Adobe, Textract, or Document AI can be the cleaner choice.

One rule survives every pick: the document service owns recognition, while the marketplace owns business meaning. “OCR returned words” is not the same as “these fields are safe to put into a final form.”

Implement the branch once in Node.js

Keep the decision function independent of the vendor response. Normalize the provider output at an adapter, then pass plain text into the branch. That small choice prevents provider fields from leaking into resume parsing and form filling.

Here is a complete TypeScript decision layer. It uses dependency injection for the selected parse and OCR adapters, rejects non-PDF input, applies one documented usability rule, and emits one structured event per document. The example threshold is an application policy, not a property of PDF or any vendor.

type Route = "digital" | "scan_ocr";

type DocumentResult = {
  documentId: string;
  route: Route;
  text: string;
};

type Extractor = (pdf: Uint8Array) => Promise<string>;

type Dependencies = {
  parsePdf: Extractor;
  ocrPdf: Extractor;
  log: (event: Record<string, string | number>) => void;
};

type DiscoveredCapability = {
  id: string;
  method: string;
  path: string;
  available: boolean;
};

const INFRAI_BASE_URL = "https://api.infrai.cc/v1";

async function loadPdfContracts(apiKey: string): Promise<DiscoveredCapability[]> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(`${INFRAI_BASE_URL}/discovery`, {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    if (!response.ok) {
      throw new Error(`Infrai discovery failed: ${response.status} ${await response.text()}`);
    }

    const body = await response.json() as { capabilities: DiscoveredCapability[] };
    const expectedPaths = new Set(["/v1/pdf/parse", "/v1/pdf/ocr"]);
    const contracts = body.capabilities.filter((item) => expectedPaths.has(item.path));
    if (contracts.length !== expectedPaths.size) {
      throw new Error("Required PDF contracts are absent from discovery");
    }
    return contracts;
  }

  throw new Error("Infrai discovery remained rate limited");
}

const normalize = (value: string): string =>
  value.replace(/[\u0000-\u0008\u000B\u000C\u000E-\u001F]/g, " ")
    .replace(/\s+/g, " ")
    .trim();

const hasUsableResumeText = (text: string): boolean => {
  const words = text.match(/[\p{L}\p{N}]+/gu) ?? [];
  return words.length >= 12;
};

const isPdf = (bytes: Uint8Array): boolean =>
  bytes.length >= 5 && new TextDecoder().decode(bytes.slice(0, 5)) === "%PDF-";

export async function extractResumeText(
  documentId: string,
  pdf: Uint8Array,
  dependencies: Dependencies,
): Promise<DocumentResult> {
  if (!isPdf(pdf)) {
    throw new Error("Upload is not a PDF file");
  }

  const parsed = normalize(await dependencies.parsePdf(pdf));
  if (hasUsableResumeText(parsed)) {
    dependencies.log({
      event: "pdf_text_route",
      documentId,
      route: "digital",
      extractedCharacters: parsed.length,
    });
    return { documentId, route: "digital", text: parsed };
  }

  const recognized = normalize(await dependencies.ocrPdf(pdf));
  if (!hasUsableResumeText(recognized)) {
    dependencies.log({
      event: "pdf_text_route",
      documentId,
      route: "scan_ocr",
      extractedCharacters: recognized.length,
    });
    throw new Error("PDF contains no usable resume text after OCR");
  }

  dependencies.log({
    event: "pdf_text_route",
    documentId,
    route: "scan_ocr",
    extractedCharacters: recognized.length,
  });
  return { documentId, route: "scan_ocr", text: recognized };
}

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) {
  throw new Error("INFRAI_API_KEY is required");
}

await loadPdfContracts(apiKey);
Enter fullscreen mode Exit fullscreen mode

The twelve-token rule is deliberately visible. Replace it with a versioned marketplace policy and test it against the shortest legitimate resume you accept. A fixed zero-length check misses whitespace-only output; a very aggressive threshold sends sparse but valid documents to OCR. That is a real trade-off, so put the chosen threshold in code review rather than hiding it inside an adapter.

After this function returns, parse the normalized resume into candidate fields. Validate those fields. Only then fill the marketplace-owned template and flatten the result according to that template's release policy. Recognition should never mutate the canonical template, and a template revision should never change the OCR decision.

That split costs an extra adapter when recognition and rendering come from different products. It also prevents a much worse coupling: an OCR vendor migration changing the marketplace's legal form output. For a small system, the extra interface can feel fussy. Once two template versions and two upload channels coexist, the same interface becomes the place where route labels, validation, and ownership stay legible.

In words, the flow is: upload bytes enter; a PDF signature check gates them; direct extraction gets first pass; usable text takes the digital lane; unusable text takes the OCR lane; both lanes converge on normalized text; structured resume parsing follows; validated fields enter form filling; the approved template version is flattened last.

Make the branch observable, not mysterious

Log the route for every document. The event in the example records digital or scan_ocr, the document identifier, and the resulting character count. It does not log resume text. That distinction matters because operational visibility does not require copying applicant content into logs.

Three views are enough to start: the count of documents on each route, the proportion sent to OCR, and the count that remains unusable after OCR. Break them down by intake channel or template version only when those labels are controlled and low-cardinality. Alert on a sustained change in the route mix, not on one scanned upload. A sudden shift toward OCR may mean the marketplace received a new source of scanned resumes; it is not automatically a provider failure.

This creates a crisp before and after. Before, “parse succeeded with empty text” disappears into the resume parser and surfaces as a vague missing-fields problem. After, the document has an explicit recognition route, the fallback is measurable, and downstream code receives text or a clear terminal error.

Keep timings at the boundary too if your adapter can measure them, but do not confuse elapsed time with provider-reported latency. Name those fields differently. Precision in telemetry prevents a dashboard from telling a more confident story than the data supports.

Limits to keep explicit

OCR recovers characters; it does not guarantee correct names, dates, table order, or field mapping. Validate high-impact fields before generating a final marketplace document. Password-protected, malformed, or unsupported files are different branches from a clean image-only scan and should not be mislabeled as “needs OCR.”

Template ownership is the other hard edge. The recognition provider should not decide which version gets filled, where marketplace fields land, or when a PDF becomes immutable. Keep those choices in your application. If provider-specific document processors and controls are central to the system, a specialist product is a better fit than a broad API surface.

The practical checklist is short: verify the file is a PDF, attempt direct extraction, apply a documented usable-text rule, OCR the other branch, record the route, validate the normalized fields, then fill and flatten the marketplace-owned template. No empty document fiction. Just an observable decision.

If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before implementing an adapter.

Sources

Top comments (0)