Use page-sized retrieval records, but bind them to one signed document manifest before a watermarked PDF leaves the media organization. The deciding constraint is evidentiary: an editor needs a page-exact source for a search answer, while security and legal teams need proof that all returned pages belonged to the approved release.
Short answer: per-document indexing preserves context but makes citations imprecise; per-page indexing makes citations precise but can detach a page from its release history. The operationally useful design is a hybrid. Store searchable text and coordinates per page, store immutable release metadata per document, and make every page record carry the document digest, manifest version, and watermark policy ID. Search at page granularity. Authorize and audit at document granularity.
That split matters for a newsroom sharing embargoed briefs, licensing packets, or review copies. A citation that says only “the 84-page press book” is hard to verify. A page hit with no cryptographic link to the approved copy is worse: it can be accurate text from the wrong revision.
Should you index a PDF per page or per document for retrieval?
A PDF is a structured file, not a bag of independent images. ISO 32000-2 defines the format, including page objects and the relationships that make up a document. Search systems still have to choose their own retrieval boundary. Nothing in the file format makes one embedding correspond to one page or one whole file.
With one record per document, a retriever can preserve cross-page context. It also has a wide citation target. If a 40-page syndication agreement is returned because of a clause on page 27, the application must locate that page after retrieval or cite the entire agreement. Re-running a text search at response time is fragile when extraction order, OCR output, or document revision differs from the indexed artifact.
Per-page records reverse the trade-off. The result already has a page number and can carry bounding boxes or character offsets, so the UI can open the cited page directly. Yet headers, footnotes, tables, and sentences can cross the boundary. A bare page record also says nothing about whether page 27 came from the internally approved file or an earlier draft with the same title.
The unit of retrieval and the unit of trust do not need to match.
Keep that boundary explicit.
For this workflow, I use three identities: a digest of the original PDF bytes, a release ID for the externally shareable revision, and a page ID derived from the release ID plus the zero-based page index. The displayed page label remains metadata because Roman numerals, inserted covers, and printed folios do not reliably equal an array index. The release manifest records the extraction method and count, while each page stores its own extracted-text digest.
| Concern | Per document | Per page | Hybrid decision |
|---|---|---|---|
| Citation target | Broad | Exact page | Return the page record |
| Cross-page context | Native | Must be reconstructed | Fetch adjacent pages from the same release |
| Authorization | One decision | Easy to apply inconsistently | Authorize the release before page retrieval |
| Audit evidence | One artifact | Many detached records | Sign one manifest containing all page digests |
| Re-index rollback | Coarse | Selective but risky | Switch an active manifest pointer atomically |
Build the page records from immutable release bytes
Do not assign identity from a filename. Names are mutable, and “final.pdf” is not an audit primitive. Read the exact bytes approved for release, hash them, then extract pages without changing those source bytes. Watermark rendering should produce a separate artifact whose digest is also recorded; otherwise a later verifier cannot distinguish the approved input from the copy actually shared.
The following program creates deterministic page IDs and rejects an empty extraction. Its page strings stand in for output from a PDF parser or OCR stage; that boundary is deliberate because extraction engines differ, while the identity rule should not.
package main
import (
"crypto/sha256"
"encoding/hex"
"encoding/json"
"fmt"
)
type PageRecord struct {
PageID string `json:"page_id"`
ReleaseID string `json:"release_id"`
PageIndex int `json:"page_index"`
Text string `json:"text"`
TextSHA256 string `json:"text_sha256"`
SourceSHA256 string `json:"source_sha256"`
WatermarkRule string `json:"watermark_rule"`
}
func digest(b []byte) string {
sum := sha256.Sum256(b)
return hex.EncodeToString(sum[:])
}
func buildPages(source []byte, releaseID, rule string, extracted []string) ([]PageRecord, error) {
if len(extracted) == 0 {
return nil, fmt.Errorf("release %s produced no pages", releaseID)
}
sourceHash := digest(source)
pages := make([]PageRecord, 0, len(extracted))
for i, text := range extracted {
seed := fmt.Sprintf("%s:%d", releaseID, i)
pages = append(pages, PageRecord{
PageID: digest([]byte(seed)), ReleaseID: releaseID, PageIndex: i,
Text: text, TextSHA256: digest([]byte(text)),
SourceSHA256: sourceHash, WatermarkRule: rule,
})
}
return pages, nil
}
func main() {
pages, err := buildPages(
[]byte("approved PDF bytes"),
"release-newsroom-042",
"external-review-v3",
[]string{"Cover", "Embargo ends at 09:00 UTC."},
)
if err != nil {
panic(err)
}
out, _ := json.MarshalIndent(pages, "", " ")
fmt.Println(string(out))
}
Persist the release row and its page rows in one transaction when the datastore supports it. When it does not, write them under a new, inactive manifest version and expose that version to search only after page count and digest checks pass. This avoids a particularly ugly partial state: search returning 37 pages from a release whose manifest promises 38.
Keep neighboring-page expansion constrained by release_id. A hit near the top of page 12 may need page 11 for the preceding sentence, but it must never borrow page 11 from another edition. Two records can share a title and still represent different evidence.
Sign the manifest, not each search response
Signing each page independently creates key operations and receipts without proving that the set is complete. Instead, create one manifest containing ordered page digests, the source digest, the rendered watermarked artifact digest, the watermark policy identifier, and the release timestamp. Sign the canonical manifest bytes. RFC 8785 specifies a JSON canonicalization scheme for repeatable hashing, while RFC 3161 defines a protocol for trusted timestamps when an independently asserted time is required.
The standard Go JSON encoder is deterministic for string-keyed maps, but this example avoids a map entirely and signs a typed structure. That is adequate when the producer and verifier use this exact encoding contract. A cross-language system should implement and test an explicit canonicalization standard rather than assume two serializers emit identical bytes.
package main
import (
"crypto/ed25519"
"crypto/rand"
"encoding/base64"
"encoding/json"
"fmt"
)
type Manifest struct {
ReleaseID string `json:"release_id"`
SourceSHA256 string `json:"source_sha256"`
RenderedSHA256 string `json:"rendered_sha256"`
WatermarkRule string `json:"watermark_rule"`
PageTextSHA256 []string `json:"page_text_sha256"`
ReleasedAt string `json:"released_at"`
}
type SignedManifest struct {
Manifest Manifest `json:"manifest"`
KeyID string `json:"key_id"`
Signature string `json:"signature"`
}
func main() {
publicKey, privateKey, err := ed25519.GenerateKey(rand.Reader)
if err != nil {
panic(err)
}
m := Manifest{
ReleaseID: "release-newsroom-042",
SourceSHA256: "5dc7...stored-full-digest",
RenderedSHA256: "91bf...stored-full-digest",
WatermarkRule: "external-review-v3",
PageTextSHA256: []string{"a10e...", "b921..."},
ReleasedAt: "2026-09-19T09:00:00Z",
}
payload, err := json.Marshal(m)
if err != nil {
panic(err)
}
sig := ed25519.Sign(privateKey, payload)
receipt := SignedManifest{m, "release-signing-key-7", base64.StdEncoding.EncodeToString(sig)}
decoded, _ := base64.StdEncoding.DecodeString(receipt.Signature)
fmt.Println("valid:", ed25519.Verify(publicKey, payload, decoded))
}
In production, the private key belongs behind a controlled signing boundary, and key_id identifies the verification key and its validity period. The audit event should contain the principal who approved sharing, recipient scope, release ID, manifest digest, policy ID, and outcome. It should not contain the full confidential page text. W3C PROV provides a general model for relating entities, activities, and agents; that vocabulary maps cleanly to an approved source, a watermarking activity, and an approving principal.
Sign late. If the watermark renderer runs after signing and its output digest is absent from the manifest, the signature proves the input but not the shared artifact.
Make retries boring and failures visible
Document pipelines are queues wearing nicer clothes. Extraction can time out, OCR can be retried, and a worker can finish after its lease expires. The defense is an idempotency key derived from the immutable source digest, release ID, extraction configuration version, and watermark policy ID. A retry with the same key may resume or return the completed result. A request with changed inputs gets a different key.
I've been paged for missed jobs and duplicate deliveries. That history pushes me toward a dull operational trade-off here: spend storage on immutable attempts and explicit states, because reconstructing which worker published which revision during an incident is slower and less reliable than retaining the evidence. The search index is disposable. The signed manifest and transition log are not.
Track states such as received, extracted, rendered, verified, signed, and published. Only published manifests are searchable. State transitions should be compare-and-swap operations so a late worker cannot move a rolled-back release forward again.
package main
import (
"crypto/sha256"
"encoding/hex"
"fmt"
)
func idempotencyKey(sourceHash, releaseID, extractorVersion, policyID string) string {
material := sourceHash + "\x00" + releaseID + "\x00" + extractorVersion + "\x00" + policyID
sum := sha256.Sum256([]byte(material))
return hex.EncodeToString(sum[:])
}
func main() {
fmt.Println(idempotencyKey(
"5dc7...stored-full-digest",
"release-newsroom-042",
"extractor-config-12",
"external-review-v3",
))
}
Page count mismatch is a hard stop. So are a missing text digest, a signature verification failure, and a rendered artifact digest that differs from the manifest. Empty extracted text is more nuanced: a blank separator page may be legitimate, while a scanned page may indicate that OCR was skipped. Record the extraction disposition per page so the verifier can tell the difference.
Operational signals should answer concrete questions: How old is the oldest unpublished release? How many jobs are retrying under the same idempotency key? Are published manifests missing page records? Did signature verification fail before or after a key rotation? Queue depth alone cannot answer any of them.
Alert on stalled state age and invariant violations, not ordinary retries. Retries are expected. Duplicate publication is not.
Fail closed.
Verify citations before publishing, and keep rollback cheap
Verification needs both offline fixtures and a pre-publication gate. Build a small corpus containing a digitally generated PDF, a scanned PDF, rotated pages, blank pages, a table spanning pages, duplicate visible page labels, and a revised file with the same filename. For each fixture, assert the source digest, page count, ordered page digests, citation target, and signature result. Do not reduce the test to “the text was found.”
Then run retrieval tests that force the boundary conditions. A query whose answer begins on one page and ends on the next should return both pages from the same release. A query against a superseded revision must either return nothing or be visibly labeled as historical, according to retention policy. An unauthorized recipient must fail before page text enters the ranking or generation path.
The release gate is short:
- Recompute the approved source and rendered artifact digests.
- Confirm extracted page count and ordered text digests against the manifest.
- Verify the signature using the key selected by
key_id. - Execute fixed retrieval probes and open every returned citation at the expected page.
- Atomically change the active manifest pointer.
Rollback should flip that pointer to the last verified manifest and stop new shares for the rejected release. Do not delete the failed records during the incident. Their state transitions, job attempts, and manifest are the evidence needed to explain what happened. Access can be disabled while evidence remains retained under the organization's policy.
This yields a clear decision rule. Use document-level records only when retrieval is meant to return whole artifacts and page precision has no user value. Use page-level records when the interface promises page citations, but require a document manifest whenever release approval, watermark provenance, or external sharing must be audited. For media documents, that hybrid boundary is usually the defensible one: a narrow citation backed by evidence for the whole released artifact.
References
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- RFC 8785, JSON Canonicalization Scheme: https://www.rfc-editor.org/rfc/rfc8785
- RFC 3161, Time-Stamp Protocol: https://www.rfc-editor.org/rfc/rfc3161
- W3C PROV-O, The PROV Ontology: https://www.w3.org/TR/prov-o/
- Go
crypto/ed25519package: https://pkg.go.dev/crypto/ed25519
Top comments (0)