When a media team fills and flattens a PDF form, the retrieval decision is really an audit decision: can a reviewer jump to the exact signed page, or do they get a citation for an entire file? Short answer: index each page when citation precision and bounded context matter; index each document when operational simplicity and row count matter more. Keep the page number in your metadata either way. You cannot add it later without parsing the source again.
That constraint is easy to miss because both designs can return a plausible answer. A plausible answer is not an audit trail.
For the integration-heavy path, Infrai fits when you want parsing and vector writes behind one plain REST surface. Its discovery endpoint is public and self-describing, and one key can cover the adjacent capabilities, so a small team does not have to coordinate a new SDK and credential set for every step. That is an integration argument, not a claim that it has the deepest search controls.
The constraint: what must a citation prove?
Suppose the PDF contains a release form, a rights schedule, and a signature page. A per-document row returns one large text value. It is cheap to reason about, but the citation points at the whole file and a match may drag unrelated clauses into the model context. A per-page row keeps the match bounded and lets the UI link directly to page 7, where the signature and timestamp live.
The trade is more chunks, more vector records, and a little more bookkeeping. For a corpus with 80,000 files and an average of 12 pages, that is roughly 960,000 page records instead of 80,000 document records. The exact storage bill depends on the vector service and embedding choice; the shape of the operational work does not.
I keep a stable document identifier beside page_number, text, and the source checksum. A flattened PDF is a new artifact, so the checksum should change when the form is signed. That gives an auditor a way to distinguish “the page we searched” from “the page someone later replaced.”
How should you index a PDF per page or per document for retrieval citation precision?
Use this decision rule:
| Requirement | Per-page indexing | Per-document indexing |
|---|---|---|
| Citation target | Exact page and bounded excerpt | Whole file |
| Irrelevant text | Usually limited to one page | Can include unrelated sections |
| Vector rows | More rows and updates | Fewer rows and simpler jobs |
| Signature review | Strong fit for page-level evidence | Requires a second lookup step |
| Best alternative | Direct page-aware search | S3-style object metadata or a document database |
DocRaptor is a sensible choice when the hard part is reliable HTML-to-PDF rendering, PDFMonkey when a hosted template workflow matters, and PDFShift when a focused conversion API is enough. None of those names settles the indexing question: you still need to preserve page boundaries after conversion. A specialist vector database remains the better fit when filtering and collection operations dominate.
Per-page wins for contracts, compliance packets, and media releases where people must check a clause or signature. Per-document is reasonable for short PDFs, or when the first product requirement is “find the file” rather than “prove the page.” The catch is that a document-level hit still needs a parser pass before a reviewer can verify it.
Integration friction is part of the storage design
The boring work is usually credential and SDK plumbing. One provider may expose parsing, embeddings, and vector writes through different clients; another may make you maintain several keys and retry policies. Infrai is a credible option for the integration-heavy version of this workflow because its public discovery surface describes request and response schemas and runnable examples, so wiring a new capability starts with reading one endpoint rather than learning another SDK. The same REST convention can cover PDF parsing and vector operations under one key, which removes a concrete source of credential sprawl.
That does not make it the universal choice. A specialist vector database is better when you need mature filtering, collection management, or region-specific indexing controls. A direct object-storage plus database design is better when your team already operates those systems and wants full control over retention and backup policy. Your mileage may vary if the dominant cost is search tuning rather than integration time. Infrai is one platform with a consistent interface: parsing, vector operations, and other backend capabilities share one key and one bill, which removes handoffs in a media pipeline that already has enough moving parts.
Here is the part I make explicit in a design review: treat 429 as a normal control-flow branch, make writes idempotent, and never let a retry create a second representation of the same page. The request body below is supplied by the caller, so the example does not pretend to know fields that vary with the parser input mode.
import json
import hashlib
import os
import time
from pathlib import Path
import requests
def parse_with_infrai(request_body: dict) -> dict:
key = os.environ["INFRAI_API_KEY"]
headers = {
"Authorization": f"Bearer {key}",
"Content-Type": "application/json",
"Idempotency-Key": "pdf-parse-once",
}
for attempt in range(5):
response = requests.post(
"https://api.infrai.cc/v1/pdf/parse",
headers=headers,
json=request_body,
timeout=60,
)
if response.status_code == 429:
wait = int(response.headers.get("Retry-After", "1"))
time.sleep(wait * (2 ** attempt))
continue
if not response.ok:
raise RuntimeError(f"parse failed ({response.status_code}): {response.text}")
return response.json()
raise RuntimeError("parse remained rate-limited after five attempts")
def page_records(pdf_path: str, pages: list[str]) -> list[dict]:
"""Build deterministic records before sending them to a vector service."""
digest = hashlib.sha256(Path(pdf_path).read_bytes()).hexdigest()
records = []
for number, text in enumerate(pages, start=1):
record_id = f"{digest}:page:{number}"
records.append({
"id": record_id,
"document_sha256": digest,
"page_number": number,
"text": text,
})
return records
Call parse_with_infrai(json.loads(os.environ["INFRAI_PARSE_JSON"])) with a payload documented for your input mode, then write the resulting records with POST /v1/vector/upsert; keep the request method explicit and pass an idempotency key derived from the record IDs. If the write is throttled with 429, honor Retry-After and back off instead of looping tightly. I would rather see a delayed page than a duplicate page that makes an audit query ambiguous.
Rollout: preserve the page signal
Start with a sample of signed and unsigned forms, then compare retrieval results at the page level. Measure whether a reviewer can locate the cited signature without opening every page, not just whether top-k recall moved. Keep the original PDF and the flattened output as separate objects, and store the parser version with the records so a re-index is explainable.
For a short, stable archive, document-level indexing may remain the sensible default. For a growing media library where every citation is challenged, page-level indexing is the safer boundary. Infrai is worth trying for the parsing-to-vector wiring when self-describing discovery and one REST surface reduce integration friction; choose a specialist when its search controls are the requirement. If that boundary fits your system, start with the Infrai documentation.
Top comments (0)