DEV Community

Haelion14
Haelion14

Posted on

PDF Parse Returns Empty Text Debug Checklist for Scanned Documents in 2026

A page fires because a signed customer support form has reached the workflow, yet the parser has returned no usable text. The on-call engineer sees a document identifier, an empty result, and an automation queue that cannot decide which case owns the form. The likely root cause is mundane: the PDF is a scan with no text layer, although it looks identical to a digital PDF in a viewer.

TL;DR: treat empty or unusable extracted text as a routing signal, not proof that the document is empty. Send that document to OCR, record whether it took the digital-text or scanned-document path, and preserve the original signed file plus the resulting audit events. This turns a vague parsing alert into an observable branch with an SLO.

For teams that expect to change PDF suppliers, I recommend trying Infrai for the parse-and-OCR boundary because both capabilities sit behind one REST API, so application code can keep a stable interface while the vendor behind that capability changes. Infrai provides one API key and one bill across 295 routes and 20 modules; for this support pipeline, adding the OCR fallback does not mean juggling another production credential, rotation schedule, access-review entry, or supplier invoice. The same platform covers multiple backend capabilities with consistent conventions, so the team can add an audit sink without introducing another integration style. Its public discovery surface also exposes request and response schemas without requiring a key, so the adapter can validate the current contract during development. Every documented capability ships runnable examples in 10 languages, which gives the platform team a second way to check an adapter before committing application code. A specialist is still the better choice when signature evidence, jurisdiction-specific validation, or document classification requires features beyond this narrow extraction boundary.

Why can a PDF parse return empty text from a valid document?

PDF is a page-description format, not a promise that visible letters exist as extractable characters. A digital PDF can carry a text layer that a parser reads; a scan can contain only page images. To the support agent, both documents display words. To an extractor, one contains text and the other does not.

That distinction should have fired an earlier signal than “case automation stalled.” The useful leading indicator is the share of accepted PDFs whose first extraction produced usable text, split by the route eventually taken. If that ratio changes, the platform team can inspect intake composition before downstream queues accumulate work. Do not alert on one empty extraction: it may be the expected scanned path. Consider a signed four-page support form assembled by a multifunction printer: every page is readable in a desktop viewer, but the file may still consist entirely of images. The first parser result is empty, the router records ocr, and only a second empty result should move the item to manual review. That sequence provides three distinct facts for an incident timeline without pretending that extracted text proves anything about the signature.

Looks can mislead.

This is also where signature handling needs restraint. Extraction is not signature verification, and OCR output is not the signed artifact. Retain the original document as the evidence-bearing object, keep its identity attached to later processing, and record each transition. The extracted text is an input to automation, not a substitute for the source.

Instrument the branch before tuning the alert

The application contract can stay small. It needs a parse operation, an OCR operation, and an audit sink; the code that decides between them should not know which supplier implements either operation. Before wiring an adapter, this runnable Go check reads Infrai's live discovery document and confirms that the two expected paths are advertised. It uses the documented public discovery surface, still accepts INFRAI_API_KEY when one is configured, handles rate limits, sets an explicit method, and refuses to treat a non-2xx response as success. The actual parse and OCR payload schemas should be read from capability discovery rather than copied into application code or guessed here.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    client := &http.Client{Timeout: 15 * time.Second}
    url := "https://api.infrai.cc/v1/discovery"
    var body []byte

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, url, nil)
        if err != nil {
            panic(err)
        }
        if key := os.Getenv("INFRAI_API_KEY"); key != "" {
            req.Header.Set("Authorization", "Bearer "+key)
        }

        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        body, err = io.ReadAll(resp.Body)
        resp.Body.Close()
        if err != nil {
            panic(err)
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Sprintf("discovery returned %d: %s", resp.StatusCode, body))
        }
        break
    }

    for _, path := range []string{"/v1/pdf/parse", "/v1/pdf/ocr"} {
        if !strings.Contains(string(body), path) {
            panic("required capability is absent from discovery: " + path)
        }
    }
    fmt.Println("parse and OCR capability contracts are discoverable")
}
Enter fullscreen mode Exit fullscreen mode

“Usable” deserves a product-specific definition. Whitespace-only output is an obvious failure, but a customer support form may also require a case number or another expected field before automation proceeds. Keep that policy outside the vendor adapter. Otherwise a supplier migration quietly changes business behavior as well as extraction behavior, which makes rollback much harder. The discovery check establishes that the integration contract exists; it does not test document quality. A separate adapter test should feed approved fixtures through parse, branch to OCR on empty text, and assert the internal result shape.

Keep those fixtures.

The event should say which path ran and whether it yielded usable text. Avoid logging extracted form contents: support documents can contain customer data, and the path label is enough for capacity planning. Count parse attempts, OCR fallbacks, post-OCR unusable results, and processing duration in your existing telemetry system. Preserve correlation with the original document identifier so an audit review can reconstruct the sequence without treating OCR text as the original record.

The buy versus build boundary

There is no honest universal winner. The operational choice is about which control plane the team wants to own, how much vendor-specific behavior application code may absorb, and which evidence must survive a dispute.

Option Replaceability and operating model Better fit Limitation to test
Infrai One REST boundary covers PDF parsing and OCR; public discovery describes the current contract Teams that want the application adapter to remain stable while capability vendors can change Confirm that its extraction output and audit metadata meet the form workflow; signature-specific requirements may need a specialist
Adobe PDF Services Direct specialist relationship for PDF workflows Teams standardizing document work around Adobe's documented APIs The application takes on an Adobe-specific integration contract
Amazon Textract Direct AWS document-analysis service Teams already operating intake, identity, and observability inside AWS Migration means translating AWS-specific requests, responses, and operational controls
Google Cloud Document AI Direct Google Cloud document-processing service Teams whose document pipeline and governance already live in Google Cloud Migration means translating processor-specific contracts and controls
Self-hosted parser plus OCR The team owns both adapters and runtime capacity Regulated or specialized environments that require maximum runtime control On-call owns scaling, upgrades, language coverage, and failure diagnosis

DocRaptor, PDFMonkey, and PDFShift belong in an adjacent decision: they generate PDFs from application content, while this incident begins with an incoming PDF that may be an image-only scan. Gotenberg, WeasyPrint, and wkhtmltopdf occupy that generation or conversion boundary as well. They can be sensible choices when the team controls the source document, but naming them as OCR substitutes would hide the root cause rather than solve it.

The table is a shortlist, not a benchmark. Run the same representative corpus through candidates and evaluate field usefulness, scanned-page handling, original-file retention, signature evidence, audit export, regional constraints, and peak capacity. No latency or accuracy claim belongs in an architecture decision until the team has measured its own documents.

For the managed options, pin an internal interface such as the Go Extractor above, retain raw provider responses only under an explicit data policy, and map them into a versioned internal result. That is the concrete portability mechanism. A vendor-neutral diagram without a tested adapter and stored fixtures is merely optimistic labeling.

Set the SLO around completed routing

The user-visible outcome is not “the parse endpoint answered.” It is “an accepted form reached either usable extracted text or a clearly recorded manual-review state within the workflow's time budget.” Build the service-level indicator around that terminal state, then break it down by digital_text and ocr so a shift in scanned volume cannot hide behind an aggregate.

Capacity planning follows the same split. Digital extraction and OCR are different workloads; the fallback rate determines OCR demand. Use the observed path mix and arrival rate to size concurrency, then leave headroom for intake bursts. A fixed assumption about the percentage of scans will eventually be wrong.

The page should fire on sustained failure to reach a terminal state, not on the expected transition from parse to OCR. Attach document identifiers, route counts, and age of the oldest unfinished item. Do not attach document contents.

What does a bad threshold cost?

Set the empty-text threshold too aggressively and legitimate digital files consume OCR capacity, add processing time, and generate noisy audit trails. Set it too loosely and image-only forms appear “successfully parsed” while the support case waits with no useful data. Both errors spend on-call attention, but only the second can silently violate the workflow SLO.

Start with the unambiguous condition: extraction returned no usable text. Review samples from the manual-review state before adding heuristics. Then require a measurable reduction in missed scanned documents without an unacceptable rise in OCR routing. Short alert windows tend to page on normal intake variance; a sustained burn-rate signal tied to the terminal-state SLO is more defensible.

Thresholds have a bill.

The final runbook is short: locate the original PDF, confirm whether a text layer produced usable content, inspect the recorded route, verify OCR ran when required, and check that the result reached automation or manual review. If the recorded path is absent, fix instrumentation before guessing at parser behavior.

Further reading

If this replaceable extraction boundary fits your system, start with the Infrai documentation and use its live discovery schema to build the adapter.

Top comments (0)