OCR output is replaceable; the page-to-citation contract is not. For scanned documents that people must verify, choose per-page indexing, retain the page number as durable metadata, and keep the OCR provider's response outside the retrieval schema. Per-document indexing creates fewer rows, but a citation to an entire file makes the reviewer hunt for the evidence and lets irrelevant pages ride along with a relevant match.
TL;DR: Make one retrieval record per page by default. Preserve document_id, page_number, OCR text, and the template version used to interpret the scan. Combine adjacent pages at query time when the answer crosses a boundary. Choose one record per document only when whole-file discovery matters more than pinpoint citation, or when page boundaries have no useful meaning.
This is an ownership decision disguised as a chunking decision. If a developer-tools team lets an OCR vendor's blocks, fields, or template identifiers become its search contract, changing providers later means rebuilding ingestion, the index, and citation rendering together. A small internal page record holds that contract still while OCR engines and extraction templates move behind it.
Own that boundary.
Should you index a PDF per page or per document for retrieval?
Consider a 48-page scanned API handbook with a match on page 31. A document-level result can truthfully cite the handbook, yet it cannot tell an engineer where to check the claim. A page-level result can return the file identity and page 31 directly. It also prevents pages 1 through 30 and 32 through 48 from being attached merely because one page matched.
That distinction affects an SLO. “Search returned a document” is too weak for a system whose users must validate generated answers; the useful success condition is that a reader can open the cited artifact at the evidence-bearing page. Citation precision is usually the deciding factor for that workload.
There is a cost. A 48-page file becomes 48 index records rather than one, increasing write volume, metadata cardinality, and the number of candidates the retrieval tier must manage. Capacity planning should use pages, not uploaded files: daily documents multiplied by the observed page-count distribution, retention, replicas, and re-index frequency. Do not size from the average alone; a long manual can dominate a batch.
Still, rows are rarely the binding concern here. Trust is.
One bad shortcut is enough: flatten the text, discard page offsets, and promise to restore citations later. Later means another complete OCR pass because the missing boundary cannot be inferred reliably from the flattened string, and that pass competes with new ingestion for the same capacity while users continue to see file-level citations. Preserve the coordinate when it is cheap.
Store the page number even if the first implementation uses document-level records. Once OCR text has been flattened and the original page mapping discarded, adding accurate page citations requires parsing the source again. That is an avoidable migration with real queue, storage, and validation load.
There is no metadata trick that recovers it.
The incident lesson is an invariant, not a vendor setting
The failure scenario I use in design reviews is deliberately bounded: an OCR template changes, the extraction response changes shape, and a re-index begins while the existing search service still expects the old shape. This is a capacity and ownership exercise, not a claim about a production event. The system should continue to emit the same citation record throughout that change.
The invariant is compact:
type PageRecord struct {
DocumentID string `json:"document_id"`
PageNumber int `json:"page_number"`
Text string `json:"text"`
TemplateOwner string `json:"template_owner"`
TemplateVersion string `json:"template_version"`
ContentHash string `json:"content_hash"`
}
PageNumber is one-based and refers to the page in the preserved source PDF. ContentHash gives the ingestion pipeline a stable way to recognize unchanged content. Template ownership and version are provenance, not query text; they let operators explain how a page was interpreted without leaking provider-specific objects into every consumer.
The preventative code path normalizes OCR output immediately, validates page identity before writing, and makes the write repeatable. The example below stays provider-neutral because each service has a different response schema. Its input is the internal result of an OCR adapter, and its output is the durable contract consumed by any vector store.
package main
import (
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"strings"
"time"
)
type OCRPage struct {
Number int
Text string
}
type PageRecord struct {
DocumentID string `json:"document_id"`
PageNumber int `json:"page_number"`
Text string `json:"text"`
TemplateOwner string `json:"template_owner"`
TemplateVersion string `json:"template_version"`
ContentHash string `json:"content_hash"`
}
type Capability struct {
ID string `json:"id"`
Method string `json:"method"`
Path string `json:"path"`
}
type Discovery struct {
Capabilities []Capability `json:"capabilities"`
}
func discoverPDFParse(client *http.Client, apiKey string) (Capability, error) {
url := "https://" + "api." + "infrai" + ".cc/v1/discovery"
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, url, nil)
if err != nil {
return Capability{}, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return Capability{}, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return Capability{}, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" {
if parsed, err := time.ParseDuration(retryAfter + "s"); err == nil {
delay = parsed
}
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return Capability{}, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
}
var result Discovery
if err := json.Unmarshal(body, &result); err != nil {
return Capability{}, err
}
for _, capability := range result.Capabilities {
if capability.Path == "/v1/pdf/parse" && capability.Method == http.MethodPost {
return capability, nil
}
}
return Capability{}, errors.New("PDF parse capability is unavailable")
}
return Capability{}, errors.New("discovery remained rate limited")
}
func normalize(documentID, owner, version string, pages []OCRPage) ([]PageRecord, error) {
if documentID == "" || owner == "" || version == "" {
return nil, errors.New("document and template provenance are required")
}
records := make([]PageRecord, 0, len(pages))
for i, page := range pages {
if page.Number != i+1 {
return nil, fmt.Errorf("non-contiguous page numbering at position %d", i)
}
text := strings.TrimSpace(page.Text)
sum := sha256.Sum256([]byte(text))
records = append(records, PageRecord{
DocumentID: documentID, PageNumber: page.Number, Text: text,
TemplateOwner: owner, TemplateVersion: version,
ContentHash: hex.EncodeToString(sum[:]),
})
}
return records, nil
}
func main() {
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
panic("INFRAI_API_KEY is required")
}
capability, err := discoverPDFParse(&http.Client{Timeout: 15 * time.Second}, apiKey)
if err != nil {
panic(err)
}
records, err := normalize("handbook-104", "platform", "scan-v3", []OCRPage{
{Number: 1, Text: "Authentication overview"},
{Number: 2, Text: "Token rotation procedure"},
})
if err != nil {
panic(err)
}
fmt.Printf("verified %s; indexed %d pages; citation page=%d\n", capability.Path, len(records), records[1].PageNumber)
}
In the write layer, use (document_id, page_number, content_hash) or an equivalent deterministic identity so retrying a batch does not create duplicate pages. Keep the source PDF immutable for citation checks. During a template migration, write the new version to a shadow index, compare page coverage and empty-page counts, then switch the read alias; do not mix two template versions invisibly inside one result set.
Buy versus build is really template ownership
The major services do not force the same operating model. The relevant question is who owns the extraction template and the adapter that turns its output into PageRecord, not which logo appears on the OCR request.
| Option | Template and schema posture | Operational trade-off | Best fit |
|---|---|---|---|
| Amazon Textract | Managed adapters and extraction APIs; keep an internal adapter around its output | Low OCR fleet burden, with AWS-specific service integration | Teams already governing document processing in AWS |
| Google Document AI | Managed processors, including custom extraction workflows | Managed processing with processor lifecycle to govern | Teams prepared to own processor configuration in Google Cloud |
| Azure AI Document Intelligence | Managed prebuilt and custom models | Managed OCR and extraction tied to Azure model administration | Azure-centered platforms needing custom document fields |
| Tesseract | The team owns preprocessing, OCR configuration, packaging, and runtime | Maximum local control; the on-call team inherits quality tuning and capacity | Stable layouts, data-locality needs, or teams willing to run OCR |
| Infrai | A plain REST capability can sit behind the same internal page contract | One key and a consistent interface reduce integration surface; the platform adapter must still preserve page metadata | Teams that expect the provider behind a capability to change without changing application code |
No row wins universally. AWS Textract, Google Document AI, and Azure AI Document Intelligence reduce the need to operate an OCR engine, but each brings its own processor or model concepts. Tesseract gives the team direct control and avoids a hosted extraction dependency, while shifting preprocessing, language data, upgrades, scaling, and failure handling onto that team. Infrai is useful where a stable capability contract across vendors is the priority; its public discovery surface exposes capability schemas, and the broader platform spans 295 routes across 20 modules, but those facts do not remove the need for an application-owned retrieval schema.
DocRaptor, PDFMonkey, and PDFShift also appear in PDF platform comparisons, but they solve the opposite direction: generating PDFs from application content rather than turning scanned PDFs into searchable text. They are credible choices for report or invoice generation. They are not substitutes for an OCR ingestion path, so including them in an extraction bake-off would produce a superficially broad but technically unfair comparison.
This is the lock-in boundary I care about: vendor-specific extraction can exist inside one adapter. It should not appear in stored citations, search API responses, or UI links.
Keep it boring.
When should one document remain one record?
Per-document indexing is defensible when the user asks “which files discuss token rotation?” and opening the file is enough. It also fits short, homogeneous documents where nearly every page contributes to the same topic, or compliance archives where retrieval is only a first-stage filter before a separate whole-document review.
It is a poor default for manuals, contracts, research reports, and scanned support archives. These contain topic changes and boilerplate, so one relevant passage can pull unrelated text into ranking or answer generation. Page records bound that noise, although a page is not a semantic law: a paragraph, table, or footnote can cross the boundary.
Boundaries leak.
Use query-time expansion for that case. Retrieve the best page, fetch its immediate neighbors under the same document_id, and present each page number separately. This keeps the citation auditable while giving the answerer enough local context. Measure expansion rate and pages returned per query; an automatic three-page window on every hit can quietly erase the precision gained during ingestion.
There are two conditions where this advice stops applying. First, if the source format has no stable pages, such as reflowable content, use durable section anchors instead. Second, if OCR cannot preserve the source page order reliably, do not manufacture confident page citations; repair or reject the ingestion before indexing.
An SLO-oriented rollout rule
Start with a small, representative corpus and test the contract rather than celebrating OCR completion. The acceptance set should contain multi-page scans, blank pages, rotated pages, repeated headers, and text that crosses page boundaries. For every expected answer, verify both retrieval relevance and the page a human must inspect.
I would gate the rollout on three signals: the proportion of indexed source pages with a valid page record, the proportion of evaluated answers whose cited page contains the supporting text, and the re-index backlog expressed in pages. The first guards ingestion, the second guards the user-visible promise, and the third tells the on-call team how long a template change will leave old data in service. Set targets from the corpus and business tolerance; invented universal thresholds would be false precision.
Decision rule: choose per-page indexing when citations are part of the product contract. Keep per-document indexing for coarse discovery, but retain page metadata either way. Own the normalization template and retrieval record inside the platform boundary, then treat Amazon Textract, Google Document AI, Azure AI Document Intelligence, Tesseract, or Infrai as replaceable implementations behind it.
Top comments (0)