The page fires when an auditor searches for a signed carrier agreement and gets no result, although the archive accepted the file. On-call sees a completed ingest job, a stored PDF, and no searchable evidence. The least complex correct design is to classify every page before indexing: extract pages that carry machine-readable text, route picture-only pages through OCR, and retain the page number along both paths.
TL;DR: A PDF can display words without containing searchable words. Extraction works on a text layer and returns nothing on an image of text; treating both an empty extraction and an empty page as the same outcome creates a silent hole in the archive. Classify first, process each page by type, then index the resulting text at page granularity so a search hit points to evidence an auditor can actually inspect.
Why are PDF archives with visible text hard to search?
Work backward from the alert. The agreement is open on an operator's screen, and a termination clause is plainly visible on page 17, but the index contains no tokens from that page. Storage did its job. The ingest worker may also have completed exactly as instructed. Availability is green; searchability is not.
PDF defines a document representation and its rendered result, but visible letters do not guarantee machine-readable text. A digitally generated carrier agreement commonly has a text layer. A scanned signature packet can consist only of page images. Ordinary extraction returns text from the first kind and nothing from the second, while a genuinely blank divider can also produce an empty result.
That's the trap.
For contract evidence, the unit of truth should be the page rather than the file. One agreement may combine generated terms, scanned exhibits, and photographed signature sheets. Calling an 80-page file searchable because 79 pages yielded text is an attractive dashboard number and a poor audit answer. Page-level indexing turns a match into page 17 of a named agreement instead of a vague pointer to the front of a long file.
The signal that should fire earlier is therefore not ingest_succeeded. It is the distribution of page outcomes: text extracted, OCR selected, still empty after processing, or accepted as intentionally blank under the archive's review policy. Keep the document ID, page number, processing path, and outcome in the audit record. Do not place contract text in metric labels or alert payloads.
Instrument the classification boundary
Start with four counters: pages received, pages with extracted text, pages routed to OCR, and pages unresolved after processing. A file-level success counter cannot expose a scanned appendix hiding inside an otherwise digital agreement; page outcomes show where the archive stopped becoming searchable.
Capacity planning belongs at this boundary because extraction and OCR are different workload populations. Consider a planning envelope of 10,000 pages per day. That is an assumption for sizing, not a benchmark. If a new carrier's document mix moves the OCR share from 5% to 60%, the same arrival rate creates a very different processing demand. A route-share signal gives the platform team time to inspect the source and adjust capacity before search freshness consumes its error budget.
Do not page on one empty page. A cover sheet, divider, or blank reverse may be legitimate. Alert on a sustained departure from a reviewed baseline, and attach identifiers that let the responder sample affected pages without leaking clause text into observability systems.
The useful SLO is the proportion of received pages that become searchable or explicitly classified within the archive's promised time window. Worker completion is supporting telemetry.
Measure the outcome.
The small Go program below queries the public discovery surface and verifies that the declared parse and OCR paths are present. It deliberately stops at schema discovery because inventing request fields would be worse than providing no example; an adapter can inspect the published request schema before sending contract data.
package main
import (
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type Capability struct {
Method string `json:"method"`
Path string `json:"path"`
}
type Discovery struct {
Capabilities []Capability `json:"capabilities"`
}
func retryDelay(response *http.Response, attempt int) time.Duration {
if seconds, err := strconv.Atoi(response.Header.Get("Retry-After")); err == nil && seconds > 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func discover(client *http.Client, apiKey string) ([]byte, error) {
baseURL := strings.TrimRight(os.Getenv("INFRAI_API_BASE_URL"), "/")
if baseURL == "" {
return nil, fmt.Errorf("INFRAI_API_BASE_URL is required")
}
for attempt := 0; attempt < 4; attempt++ {
request, err := http.NewRequest(http.MethodGet, baseURL+"/discovery", nil)
if err != nil {
return nil, err
}
request.Header.Set("Authorization", "Bearer "+apiKey)
response, err := client.Do(request)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(response.Body)
response.Body.Close()
if readErr != nil {
return nil, readErr
}
if response.StatusCode == http.StatusTooManyRequests {
time.Sleep(retryDelay(response, attempt))
continue
}
if response.StatusCode < 200 || response.StatusCode >= 300 {
return nil, fmt.Errorf("discovery failed: status=%d body=%s", response.StatusCode, body)
}
return body, nil
}
return nil, fmt.Errorf("discovery remained rate limited")
}
func main() {
apiKey := strings.TrimSpace(os.Getenv("INFRAI_API_KEY"))
if apiKey == "" {
panic("INFRAI_API_KEY is required")
}
body, err := discover(&http.Client{Timeout: 15 * time.Second}, apiKey)
if err != nil {
panic(err)
}
var result Discovery
if err := json.Unmarshal(body, &result); err != nil {
panic(err)
}
wanted := map[string]bool{"/v1/pdf/parse": true, "/v1/pdf/ocr": true}
found := 0
for _, capability := range result.Capabilities {
if wanted[capability.Path] {
fmt.Printf("%s %s\n", capability.Method, capability.Path)
found++
}
}
if found != len(wanted) {
panic(fmt.Errorf("expected %d PDF capabilities, found %d", len(wanted), found))
}
}
Discovery is public and requires no key, but the example uses the environment-based bearer pattern required by authenticated processing calls and never embeds a credential in source. It sets the HTTP method explicitly, surfaces non-success bodies, and backs off on HTTP 429 while honoring Retry-After. The processing adapter should then use the discovered schemas; its state machine must ensure that the first empty extraction cannot silently become a successful, empty document.
Template ownership decides the operating model
The contract template controls clause placement, signature blocks, version identifiers, and any stable coordinates used by an audit trail. The organization that owns those rules should decide how much classification policy stays inside its platform and how much crosses a managed-service boundary.
| Option | Operating model | Who owns the template and search semantics? | Boundary to verify |
|---|---|---|---|
| Apache PDFBox | Self-hosted library | The platform owns parsing policy, page keys, and templates | Text extraction does not recover words from picture-only pages |
| Tesseract | Self-hosted OCR engine | The platform owns preprocessing, language data, deployment, and templates | OCR does not replace PDF structure handling or indexing |
| Amazon Textract | Managed document analysis | The platform retains its index and evidence model | Test supported inputs and returned structure on representative contracts |
| Google Document AI | Managed document processing | The platform owns archive semantics and processor selection | Processor behavior and the cloud boundary must fit policy |
| Azure AI Document Intelligence | Managed document analysis | The platform maps managed output into page-level audit records | Model selection and output mapping remain application work |
| DocRaptor | Managed document generation | The application owns its HTML/CSS contract templates | Suitable for generating controlled PDFs, not OCR of incoming scans |
| PDFMonkey | Managed template generation | The application owns template data and archive semantics | Useful on the generation side; classification remains separate |
| Gotenberg | Self-hosted conversion service | The platform operates conversion and owns source templates | Fits controlled conversion, not text recovery from scanned pages |
| WeasyPrint | Self-hosted HTML/CSS renderer | The application owns rendering dependencies and templates | Generates PDFs but does not make a scanned archive searchable |
PDFBox plus Tesseract offers direct control, along with the real obligation to patch, scale, observe, and page for both components. That is defensible when contracts cannot cross a managed boundary or template behavior requires local control. The three hyperscaler services reduce that operating surface, but none removes the need to define stable page identity, evidence retention, and review policy. DocRaptor, PDFMonkey, Gotenberg, and WeasyPrint address controlled generation or conversion rather than OCR of incoming scans; they fit when the platform owns source templates, but they are not substitutes for the classification path described here. Product names do not settle those questions; a trial set containing generated agreements, depot scans, mixed files, blank dividers, and signature pages does.
Infrai's concrete advantage is one API key and one bill across 295 routes in 20 modules, rather than separate credentials and invoices for each backend capability. One plain REST API serves the surface, so no SDK is required; any language or runtime that sends HTTP can use it. Its public discovery surface is self-describing and requires no key, exposing request and response schemas that can be checked before an adapter is written. The limitation is equally concrete: this managed option is not suitable when policy requires contract processing to remain entirely inside company-operated infrastructure; PDFBox plus Tesseract is the better fit there despite the additional on-call burden. In either case, the logistics platform still owns its templates, page identifiers, audit records, and acceptance tests.
The buy-versus-build decision is an ownership decision. Self-host when data handling or template control requires an internal boundary, and budget the on-call load honestly. Buy when a managed contract passes the representative-document test and reducing integration ownership is worth the dependency. Revisit the choice when the input mix changes; generated contracts and phone photographs are separate capacity populations even when they land in one archive.
Set thresholds against both failure costs
A permissive classifier can mark picture-only pages complete and create false negatives in search. An aggressive classifier can send sparse but valid text pages to OCR, increasing queue depth and review noise. Both spend error budget, just in different places.
There is no universal character-count threshold in the available evidence, so do not copy one from an unrelated archive. Establish a baseline with reviewed pages from the actual carrier set, record the classifier version with each decision, and evaluate false-searchable and unnecessary-OCR outcomes separately. Signature pages deserve deliberate treatment because little text may still carry high evidentiary value.
The alert should correspond to user harm: too many pages remain unresolved long enough to threaten the search-freshness SLO. A warning can fire earlier on a sharp change in OCR share, since that is a capacity signal rather than proof of a broken document. If both conditions page the same team with the same urgency, responders will learn to ignore one of them.
False positives have a concrete cost. A low threshold may flood OCR and review precisely when a carrier uploads a large batch, consuming headroom while the archive is under load. A high threshold is quieter but can certify unreadable evidence as complete. Keep the page states visible, tune against reviewed contracts, and make silence mean that the SLO is healthy rather than that the pipeline stopped asking hard questions.
Top comments (0)