DEV Community

Byte Chap
Byte Chap

Posted on

Ranking Scanned PDFs While Preserving OCR Uncertainty

Combine lexical and semantic retrieval, then use region evidence and page metadata to rank results without concealing weak extraction.

A scanned PDF presents two separate search problems: finding the relevant passage and deciding how much to trust the text extracted from its pixels. Hybrid retrieval helps with the first. It does not automatically solve the second. A strong semantic match can still point to a misread amount, an incomplete table, or a paragraph assembled from the wrong reading order.

This article was drafted with AI assistance, then fact-checked and edited by the developer behind DocBento.

I would treat OCR evidence as part of the retrieval contract. Search should return useful candidates while preserving enough provenance for the reader to inspect doubtful text. That requires a deliberate distinction between relevance, extraction quality, and human review.

What each retrieval channel can recover

Full-text retrieval operates on extracted tokens. It is useful when the query contains distinctive terminology, a clause heading, or an identifier that survives OCR. Ranking can incorporate frequency, proximity, and field weights. PostgreSQL provides ts_rank and ts_rank_cd for lexical ranking; its native full-text search should not be casually described as BM25. PostgreSQL’s text-search documentation explains these ranking controls.

Identifiers need additional care. Language-aware tokenization and stemming are designed for prose. A reference such as AB-1047 may deserve a separate normalized identifier field, with explicit rules for punctuation and case. Otherwise, an apparently exact search may be exact only after transformations the user cannot see.

Vector retrieval compares an embedded query with embedded passages. It can recover a passage whose wording differs from the query: “ending the agreement early” might retrieve a section headed “Termination.” But the embedding represents the extracted text, including its errors. It cannot establish whether a faint digit was recognized correctly. Corrupted reading order can also produce a plausible embedding of a passage that never existed as a coherent paragraph.

Both channels therefore fail differently. Lexical retrieval may miss a damaged query term. Semantic retrieval may overlook the distinction between neighboring reference numbers or related contractual concepts. Keeping both candidate lists gives the ranking stage more evidence, but agreement between them is evidence of relevance to the extraction, not proof of transcription accuracy.

Preserve evidence at the region boundary

A document-wide OCR average is too coarse for ranking. A clean title page says little about a faint payment schedule on page 18. Even a page average can hide one uncertain amount surrounded by clear boilerplate.

Store evidence at a region boundary that can be mapped back to the scan: a paragraph, table cell, or another bounded area. Search chunks may include several regions, but they should retain that mapping. Useful fields include the document version, extraction revision, physical page index, bounding box, OCR engine and configuration, raw confidence values, and assessment state.

Keep raw engine confidence separate from any derived quality estimate. A confidence value should not be treated as a portable probability of correctness. Its usefulness depends on the engine, language, layout, and material being processed. Missing confidence is also different from low confidence. A vision-based extraction that supplies no comparable score belongs in an unknown-quality category, not at an invented midpoint.

For ranking, derive a region-quality estimate only after evaluating representative scans. For lexical hits, inspect the regions containing the matched terms. For semantic hits, use the contributing chunk regions conservatively unless a later alignment step identifies a narrower passage. Do not claim word-level certainty from a chunk-level similarity score.

Retrieval path from scan regions through lexical and vector candidates to evidence-aware ranking, alongside unassessed, usable, uncertain, and reviewed evidence states.

Make review belong to a specific extraction

Use four evidence states: unassessed, usable, uncertain, and reviewed. These are workflow states, not four numerical confidence bands.

A new extraction starts unassessed. An automated assessment can mark a region usable or uncertain under a documented policy. A human can inspect either state against the scan and mark that region reviewed. A reviewer can also assess an unassessed region directly. Reviewed means the recorded transcription was checked within a defined scope; it does not certify the document’s factual claims or legal meaning.

The revision boundary matters. In the policy proposed here, every replacement extraction creates new unassessed regions, including when it incorporates a correction. Review status never transfers automatically. A reviewer must check the new region before it becomes reviewed. This costs an additional assessment, but avoids the ambiguity of attaching an old review to changed text or coordinates.

Retain the earlier extraction and review record for provenance. The active search index should point to the current extraction, while a result opened from an older search session can identify that its revision has been superseded. Updates to text, lexical index entries, embeddings, and evidence should become visible together; otherwise a result can combine an old embedding with a new snippet.

Fuse relevance before applying a bounded penalty

Lexical scores and vector similarity scores generally have different scales. Adding them directly makes ranking depend on incidental score ranges. Reciprocal rank fusion offers a practical starting point: each candidate receives contributions based on its position in each list, with absent candidates contributing zero. Elastic’s RRF documentation describes this method for combining independently ranked results.

Fuse candidates using stable chunk identities within one extraction revision. Then apply a bounded quality adjustment. I prefer a modest demotion over a hard confidence cutoff because the only relevant passage may be difficult to read.

The following Python example expresses that policy. base is an already-fused relevance score; quality is an application-derived region estimate between zero and one, not raw OCR confidence.

def evidence_score(base: float, quality: float | None) -> float:
    if quality is None:
        return base  # Preserve rank; display unknown quality separately.
    if not 0.0 <= quality <= 1.0:
        raise ValueError("quality must be between 0 and 1")
    max_penalty = 0.20
    return base * (1.0 - max_penalty * (1.0 - quality))
Enter fullscreen mode Exit fullscreen mode

The multiplier stays between 0.8 and 1.0. Missing quality leaves the relevance score unchanged and requires a visible unknown-quality label. That is an explicit starting policy, not an assertion that unknown extraction is reliable. Tune it against retrieval judgments and transcription checks.

Consider this hypothetical candidate set. The numbers demonstrate the formula; they are not measured performance or recommended parameters.

Candidate Fused relevance Region quality Adjusted score
A: uncertain clause 0.032 0.20 0.02688
B: clear related clause 0.030 0.95 0.02970
C: unique reference hit 0.024 0.10 0.01968

Before adjustment, the order is A, B, C. After adjustment, B overtakes A. C remains available despite weak extraction. Preserving that unique hit also requires candidate-generation and display policies: include exact identifier candidates in the union, retain a protected result slot when appropriate, and label their uncertainty. A bounded penalty alone cannot prevent a later top-k cutoff from discarding them.

Avoid giving reviewed passages an unconditional relevance boost. A carefully checked but irrelevant paragraph should not outrank the requested clause. Review describes evidence quality; relevance still determines whether the passage answers the query.

Page metadata changes both eligibility and presentation

Apply authorization and explicit user filters consistently to both retrieval channels before fusion. Folder access, selected document versions, and requested document types determine eligibility. A stale vector index must not reintroduce a passage excluded by the lexical path.

Other metadata can inform ranking more softly. Repeated headers and footers can generate many nearly identical hits; identifying those regions allows downweighting without erasing potentially meaningful text. Page position can support a heading or appendix heuristic, but “earlier page wins” is a poor general rule. A signature block or technical schedule may be the most relevant material near the end.

Group redundant chunk hits for presentation while keeping their underlying evidence. Ten overlapping chunks from one page should not crowd out a distinct page with a useful match. Deduplicate by region identity or overlap, not merely by equal text, because identical wording on two pages may have different significance.

Return the physical page index and, when available, the printed page label as separate fields. The third PDF page might be labeled “ii”; treating these as interchangeable produces misleading navigation. A snippet should also carry its extraction revision and uncertainty annotation. Readers need to know whether the doubtful text is the matched identifier, a neighboring sentence, or a table region used by the semantic candidate.

Test the path from uncertain hit to corrected result

A useful evaluation set includes exact references, paraphrased clauses, faint numbers, tables, rotated pages, and scans with no comparable confidence output. Judge retrieval separately from transcription: did the relevant page enter the candidate set, and does its displayed text accurately represent the scan?

Measure candidate recall, final ordering, page navigation, and the effect of quality penalties. Check whether low-quality unique hits survive. Also inspect how often an unknown-quality candidate outranks an assessed one and whether that reflects relevance or an avoidable scoring artifact.

The correction path needs its own test. In this hypothetical trace, an uncertain reference remains discoverable. A reviewer finds an OCR error, replacement extraction creates revision r2, and r2 requires assessment before review. The index then publishes the corrected text and evidence together.

Six-step hypothetical trace showing an uncertain hit, a bounded penalty, scan inspection, an unassessed replacement revision, review of the new region, and publication of matching indexes.

Test interruption between those steps. If extraction succeeds but embedding fails, keep the previous coherent index active or explicitly expose the new revision’s limited retrieval availability. Never present mixed revisions as one settled result. Log the policy version used for assessment and ranking so that a changed order can be explained later.

A concrete document-management implementation

My own product, DocBento — Python AI Document Management, is a self-hosted application built with FastAPI and React, using PostgreSQL with pgvector or MariaDB. Its OCR choices include offline PaddleOCR on ONNX Runtime, Tesseract, optional Docling, and OpenAI-compatible vision models. Hybrid keyword and semantic search returns filters, snippets, and page numbers. The optional assistant searches and reads documents before answering and provides page citations.

Those capabilities provide a concrete setting for the retrieval problems discussed here. The region evidence lifecycle and bounded ranking policy above are engineering proposals; the listed product facts do not establish that DocBento implements them. When assessing a document-search system, I would inspect how it preserves extraction revisions, represents unknown confidence, and connects a result to the scanned page before choosing a quality penalty.

DocBento on Gumroad

Top comments (0)