TL;DR: Use born-digital text extraction when your B2B SaaS team owns the PDF templates; preserve the extractor's page boundary and index only redacted text. Choose OCR when customers control the templates, scans are valid input, or extraction quality cannot be enforced. Record which path produced every page so citations remain explainable.
| Input and ownership | Pick | Why | Required signal |
|---|---|---|---|
| Your team owns the export template | Native PDF extraction | Template changes can be tested | extracted characters and empty-page count |
| Customer uploads may be scans | OCR | Image-only pages lack a useful text layer | OCR path, confidence, and page count |
| Mixed portfolio | Native first, then explicit OCR fallback | Known-good exports avoid OCR while scans remain valid | extractor kind per page and fallback reason |
The firm rule is about control. Template ownership lets a team test the text layer before release. Without that control, visual correctness does not prove extractable text. For a document-sharing workflow, the pipeline is PDF ingestion, page-preserving extraction, personal-data redaction, chunking, embedding, and retrieval. The index never receives unredacted page text.
How should Node.js index PDF text per page?
The extraction adapter should. A PDF is a structured document defined by ISO 32000-2, not a bag of newline-delimited strings. If an application concatenates the whole file and guesses page breaks later, a retrieval hit cannot reliably point back to the page a reviewer must inspect.
Treat pageNumber as provenance. Keep it 1-based at the adapter boundary, validate it, and carry it through redaction, chunking, storage, and result rendering. One convention removes an ugly class of off-by-one citations.
Page numbers matter.
Text may appear in an order different from the order a person reads it. Multi-column layouts, headers, footers, and positioned glyphs expose that gap. Template owners can prevent regressions with fixture PDFs and expected page text. For customer-controlled files, use quality gates and route weak pages to OCR or review rather than pretending every successful parse is useful.
Diagram in words: upload enters quarantine; a parser emits numbered pages; a redactor replaces personal fields; a chunker emits page-scoped passages; the retrieval index stores those passages plus provenance; the sharing service renders only redacted results. Raw text takes a separate, access-controlled path to deletion.
Choose this path for invoices, account summaries, or reports generated from templates your team releases. The advantage is testability. Keep representative PDFs as fixtures, assert the page count, and assert stable anchor phrases on each page after every template change. A test should fail if page 3 becomes empty even when the rendered PDF still looks fine.
Count pages with no extracted characters. Track the ratio of redacted characters to extracted characters as a distribution, without logging the characters themselves. Alert on a sustained change by template version, because a global average can hide one broken export family.
Never log source text. OWASP's logging guidance calls out data that commonly needs removal, masking, sanitization, hashing, or encryption, including personal data. Here, document identifiers should be pseudonymous, and trace data should describe stages and counts rather than document contents.
Pick OCR when uploads vary
Choose OCR when valid inputs include scanned pages or another organization owns document creation. It adds an error surface, needs its own quality signal, and may produce plausible text that is wrong. A low-confidence page should not produce a confident-looking citation.
Keep the fallback visible. Store ocr as the extractor kind, retain page-level confidence when the engine supplies it, and define a review state below your acceptance threshold. Derive that threshold from labeled documents in the languages and layouts you accept. Do not borrow a magic number from a demo.
Scans change that.
For mixed PDFs, decide page by page. A ten-page upload can contain nine born-digital pages and one scanned signature page. Preserve all ten page numbers even if one yields no searchable chunk after redaction. Gaps are evidence.
Consider a concrete upload with pages 1 through 10. Pages 1 through 6 contain selectable account text, page 7 is a scanned authorization form, and pages 8 through 10 return to selectable text. Native extraction is the right first path, but it is not suitable for page 7. OCR is the right fallback there, but running OCR over all ten pages would replace known text with probabilistic output for no retrieval benefit. The trade-off is deliberate: a page-level branch makes the job state and monitoring more complex, while preserving the strongest available evidence for each citation. The index should still contain one provenance convention across both branches. A result from page 7 carries extractor: \"ocr\"; a result from page 8 carries extractor: \"native\". Reviewers can now distinguish those evidence paths without seeing private source text in logs.
Implement a redaction-first page contract
The parser or OCR adapter returns numbered pages; the policy layer owns redaction; the indexer accepts only redacted chunks. This TypeScript example leaves parsing, OCR, and embedding behind interfaces, so the contract does not depend on one product.
type ExtractorKind = "native" | "ocr";
type ExtractedPage = {
pageNumber: number;
text: string;
extractor: ExtractorKind;
confidence?: number;
};
type IndexedChunk = {
documentId: string;
pageNumber: number;
chunkNumber: number;
text: string;
extractor: ExtractorKind;
};
interface PageExtractor {
extract(pdf: Uint8Array): Promise<ExtractedPage[]>;
}
interface RetrievalIndex {
upsert(chunks: IndexedChunk[]): Promise<void>;
}
type Redact = (text: string) => string;
function chunkPage(text: string, maxChars = 1_200): string[] {
const paragraphs = text.split(/\n\s*\n/).map((part) => part.trim()).filter(Boolean);
const chunks: string[] = [];
let current = "";
for (const paragraph of paragraphs) {
if (current && current.length + paragraph.length + 2 > maxChars) {
chunks.push(current);
current = paragraph;
} else {
current = current ? `${current}\n\n${paragraph}` : paragraph;
}
}
if (current) chunks.push(current);
return chunks;
}
async function indexPdf(
documentId: string,
pdf: Uint8Array,
extractor: PageExtractor,
redact: Redact,
index: RetrievalIndex,
): Promise<{ pages: number; chunks: number; emptyPages: number }> {
const pages = await extractor.extract(pdf);
const seen = new Set<number>();
const chunks: IndexedChunk[] = [];
let emptyPages = 0;
for (const page of pages) {
if (!Number.isInteger(page.pageNumber) || page.pageNumber < 1) {
throw new Error("Extractor returned an invalid page number");
}
if (seen.has(page.pageNumber)) {
throw new Error(`Duplicate page number: ${page.pageNumber}`);
}
seen.add(page.pageNumber);
const safeText = redact(page.text).trim();
if (!safeText) emptyPages += 1;
chunkPage(safeText).forEach((text, chunkNumber) => {
chunks.push({ documentId, pageNumber: page.pageNumber, chunkNumber, text,
extractor: page.extractor });
});
}
await index.upsert(chunks);
return { pages: pages.length, chunks: chunks.length, emptyPages };
}
Chunk inside each page. This may create smaller chunks at page endings, but it makes the citation exact. A hit can render as Document 8f31, page 4 without reconstructing a cross-page span. If recall requires a passage across a boundary, store an explicit page range.
That boundary costs recall.
Test that names, email addresses, account identifiers, and other in-scope fields are absent from IndexedChunk.text. Also test text split across PDF drawing operations; it may not arrive as one convenient string. Redaction tests must run on extracted fixtures, not only hand-written sentences.
Idempotency belongs at the index boundary. Use a stable identity composed of document ID, source revision, page number, chunk number, and redaction-policy version. A retry should replace the same logical records. It must not create another searchable copy of yesterday's personal data.
Observe without observing the person
Emit a structured event per stage with a pseudonymous document ID, template version when known, page count, extractor kind, duration, empty-page count, chunk count, redaction-policy version, and final status. Avoid page text, matched personal values, filenames, and user-provided titles.
Build three views: extraction health from empty pages and OCR fallback; policy health from rejected documents and aggregated redaction counts; retrieval health from the share of returned chunks containing a valid document ID and page number. Logs explain individual jobs. Metrics reveal drift.
Keep both.
Alert on symptoms tied to an action. A sudden empty-page increase for one owned template goes to the template team. A broad OCR-confidence shift goes to the ingestion owner. Missing citation metadata blocks publication because the user cannot verify the result.
Limits
Page citations identify where extracted evidence came from; they do not prove extraction or redaction is correct. This approach has a hard limitation: it is not suitable for a workflow that cannot retain stable page identity from ingestion through retrieval. Validate against rendered pages, include difficult layouts and scans in the test set, and put low-quality inputs into review. Encrypted files, malformed documents, unsupported scripts, and pages with no useful text need explicit outcomes.
The choice stays simple: owned, tested templates favor native extraction; uncontrolled or scanned input requires an OCR path. In both cases, redact before indexing and make page provenance impossible to drop.
References
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- OWASP Logging Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html
- Node.js
Uint8Arraydocumentation: https://nodejs.org/api/buffer.html#class-uint8array - Unicode Standard Annex #29, Text Segmentation: https://unicode.org/reports/tr29/
Top comments (0)