A customer-support document bundle can be merged for review and split again for escalation, so the page number visible in the current PDF is not a durable citation. Index text by source page, preserve that identity through every bundle operation, and translate it to the rendered page only at the presentation boundary. That keeps retrieval evidence explainable after templates change ownership.
TL;DR: extract one ordered text record per PDF page, attach an immutable document ID and source-page number, record every merge or split as a mapping rather than overwriting those coordinates, and return both the stable source coordinate and current display coordinate with each result. A Node.js service can call this indexing boundary; the contract matters more than the runtime.
The incident-shaped failure is easy to miss in a dashboard. A support agent searches a bundle containing a customer letter, troubleshooting transcript, and return authorization. The result says page 14. Operations later splits the return paperwork into a separate packet, and page 14 now contains something else. Search stayed green, latency stayed flat, and the citation became false. What page fired: the one in the result, or the source page that actually contained the sentence?
How should Node.js index PDF text per page?
PDF is a paginated document format, but an ordinal position inside one assembled file is not a business identifier. Merge two files and every sheet in the second file moves. Insert a cover sheet and nearly every visible ordinal moves. Split a packet and the same source sheet may become number 1 in a new artifact. The extracted text can remain useful while the locator silently rots.
It breaks quietly.
Template ownership makes this worse. In customer support, one team may own the correspondence template while another owns return or compliance forms. If the bundling pipeline treats the final PDF as the only record, a template team's harmless insertion changes retrieval coordinates for content it does not own. Ownership should stop at the source-document boundary; assembly owns mappings, not source identity.
This invariant is small: (source_document_id, source_page) never changes. A generated artifact gets its own ID and a mapping from each display page back to that pair. Search chunks inherit the stable pair. The UI may show Bundle page 14, but the audit record must still say which source document and page supplied the quoted text.
The indexing contract that survives assembly
Keep extraction, assembly, and retrieval separate. Extraction emits ordered page records. Assembly emits lineage. Retrieval stores chunks with page-level provenance and refuses to manufacture a citation when provenance is missing.
type PageRecord struct {
SourceDocumentID string
SourcePage int // One-based, as shown to reviewers.
Text string
TextSHA256 string
}
type DisplayPage struct {
ArtifactID string
DisplayPage int
SourceDocumentID string
SourcePage int
}
type Citation struct {
SourceDocumentID string
SourcePage int
ArtifactID string
DisplayPage int
Quote string
}
Do not use extracted text as identity. Normalization or template revisions can alter it, while duplicate boilerplate can make identical text appear on several pages. A digest can detect a changed extraction result, but document ID and page ordinal remain the locator.
For a Node.js ingestion endpoint, make the boundary accept a PDF plus caller-supplied document identity, then produce page records before chunking. Do not flatten the document into one string and reconstruct page breaks later. The extractor must preserve page order and represent a page with no extractable text as an explicit record; otherwise page 8 can be dropped and page 9 mislabeled as page 8.
Make invalid lineage fail before search
The preventative path belongs where an assembled artifact is registered. It should reject duplicate display pages, gaps, nonpositive ordinals, and references to pages absent from the extraction ledger.
func ValidateAssembly(pages []DisplayPage, known map[string]map[int]bool) error {
seen := make(map[int]bool, len(pages))
for _, p := range pages {
if p.DisplayPage < 1 || p.SourcePage < 1 {
return fmt.Errorf("page ordinals must be positive")
}
if seen[p.DisplayPage] {
return fmt.Errorf("duplicate display page %d", p.DisplayPage)
}
if !known[p.SourceDocumentID][p.SourcePage] {
return fmt.Errorf("unknown source page %s:%d", p.SourceDocumentID, p.SourcePage)
}
seen[p.DisplayPage] = true
}
for n := 1; n <= len(pages); n++ {
if !seen[n] {
return fmt.Errorf("missing display page %d", n)
}
}
return nil
}
Good.
At 3 a.m., a count of indexed chunks is weak evidence; the useful page names an artifact that references return-form:3 even though extraction produced two pages. Alert on rejected lineage, sustained extraction failures, and citations whose mapping cannot be resolved. The page should name the broken invariant and affected artifact.
Retries also need a stable key. Indexing the same source revision twice should replace or confirm the same page records, not create parallel identities. A new template revision gets a new source revision identity even if its filename is unchanged. Otherwise a late retry can attach old text to new mappings.
Test the operations users perform
A happy-path search test proves little. Build fixtures around transformations: merge a three-page transcript with a two-page authorization, insert a one-page cover, split pages 2–4 into a new artifact, and reorder a packet. Assert that each display page resolves to exactly one known source page and that a quoted result resolves back to the same extracted page text. The visible number should change in some tests. The source coordinate must not.
Use malformed fixtures too: a protected file the extractor cannot open, a scanned page with no embedded text, repeated boilerplate, and a mapping that omits one page. The expected behavior is explicit failure or an extraction_required state, never a confident empty result. OCR is a separate extraction mode and should record that distinction because its text can differ from embedded text; the citation still targets the page image a reviewer can inspect.
Deployment should be versioned around data contracts. Run a shadow index when changing extraction or chunking, compare page counts and unresolved mappings, then switch reads only after the new index passes lineage checks. Keep the old index long enough to resolve citations already issued. This costs storage and operational attention, but avoids turning an extraction upgrade into an untraceable evidence rewrite.
Alternatives and their limits
Indexing the final assembled PDF alone is acceptable when artifacts are immutable, never split, and citations only need to live as long as that exact file. It is the simplest model. It fails when a support workflow republishes the same documents in a different order.
Embedding page labels inside chunk text may improve retrieval context, but does not establish lineage; labels can be duplicated, omitted, or generated by a template. Storing only bounding boxes gives precise highlighting within a page, yet still needs stable page identity. Keeping source and display coordinates adds a join at read time. That is a real trade-off, but the join is testable and the alternative hides history.
This advice does not apply when the retrieval unit is an immutable page image with a content-addressed identifier and there is no concept of regenerated bundles. There, image identity already supplies a stable coordinate. For mutable support packets, the decision rule is blunt: the team owning a template may revise its pages; no revision may rewrite another team's evidence identity.
Operational decision rule
Choose an extraction component by verifying that it emits page boundaries deterministically for your fixture corpus, exposes failures rather than silently skipping pages, and can be pinned and regression-tested. Choose an index schema by asking whether a citation issued before a merge, split, or template revision can still be resolved afterward. Names and feature grids are secondary.
A useful result carries the quote, stable source document and page, current artifact and display page, extraction revision, and enough authorization context to prevent a user from resolving evidence from a document they cannot access. Log identifiers and outcomes, not customer text. Measure unresolved lineage and extraction-state counts alongside latency.
The durable design is a provenance ledger with search attached, not a search index with page numbers sprinkled on top. It makes bundle changes visible, keeps template ownership bounded, and gives the on-call engineer a concrete invariant when a citation cannot be resolved.
Sources
References:
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
Top comments (0)