Verify signed PDFs yourself and log the result beside every document before OCR or indexing begins. The deciding constraint is evidence ownership: independent verification gives you a record you can produce in a dispute, while trusting a sending platform leaves you with evidence from someone else's dashboard.
TL;DR: treat a platform's signature status as useful context, not as your verification record. Keep the expected certificate, verify the PDF at intake, and attach that outcome to the document's internal identity. For a healthtech archive, failed or unknown verification should stop the document before its text becomes searchable.
This is a small control with a wide blast radius. A scanned referral can pass through signature checking, OCR, chunking, and vector search. If provenance falls off at the first handoff, a fast batch merely makes unverified text searchable faster.
The before-and-after mental model
Before: a receiving service downloads a signed PDF, records a screenshot or webhook status from the sender, then pushes the file into an OCR queue. Later, an auditor can see what the sending platform displayed. The auditor cannot reproduce what the receiver checked because the receiver did not check anything.
After: intake stores the document identity, the expected certificate identity, and the receiver's verification result. OCR and indexing inherit that internal record. The sender's evidence can remain attached, but it is corroboration rather than the only proof.
Stop there.
That difference matters more than vendor branding. The verification record should belong to the system that relies on the document. Holding the expected certificate creates a modest operational duty: certificates need controlled storage, rotation, and an explicit mapping to the party expected to sign. That cost is real. So is the benefit of being able to produce your own result.
The logs should answer four questions without reopening a dashboard: which document was checked, which expected certificate was used, when the check ran, and what the result was. Do this for every document, including failures. A missing event must not be confused with a successful check.
Log the miss.
Should you verify a PDF signature yourself or trust the sending platform?
The major signing platforms provide useful evidence, but they do not erase the distinction between platform evidence and receiver-owned verification. Adobe Acrobat Sign exposes audit reports and transaction history. DocuSign provides certificates of completion and envelope history. Dropbox Sign provides an audit trail. Those artifacts describe activity recorded by their respective platforms; they are valuable inputs to a review.
Independent verification answers a different question: what did your intake boundary determine about this PDF against the certificate it expected? The two records can coexist. In fact, retaining both is often the more legible design because disagreements become visible rather than silently resolved in favor of whichever UI an operator last opened.
| Approach | Record you control | Operational burden | Sensible boundary |
|---|---|---|---|
| Trust Adobe Acrobat Sign evidence | Your copy of Adobe's report or status | Low at intake | Workflows where platform evidence meets the relying party's policy |
| Trust DocuSign evidence | Your copy of envelope history and completion evidence | Low at intake | DocuSign-centered agreement workflows |
| Trust Dropbox Sign evidence | Your copy of its audit trail | Low at intake | Dropbox Sign-centered workflows |
| Verify at your boundary | Your result tied to the PDF and expected certificate | Certificate custody and rotation | Disputes, regulated records, or downstream automation that relies on authenticity |
This is not a claim that an in-house verifier is automatically superior. Verification logic still has to follow ISO 32000-2, and the policy around trusted certificates still has to be correct. Adobe, DocuSign, and Dropbox Sign also offer richer agreement workflows than a bare verification call. Pick the evidence boundary first. Then pick the tool.
There is another category boundary worth making explicit. DocRaptor, PDFMonkey, and PDFShift generate PDFs from application content; Gotenberg, WeasyPrint, and wkhtmltopdf convert HTML into PDFs. They are real alternatives for document generation or conversion, but none replaces a signing platform's audit trail or receiver-side signature verification. Choose one of them when the job is rendering a document. Do not count a successful render as proof of a signature.
Keep verification attached through the OCR handoff
Batch throughput is the primary design axis for scanned health documents, but verification must precede parallel fan-out. One cheap gate before OCR prevents untrusted text from spreading across chunks and vector records. Then OCR workers can scale independently for the accepted batch.
The following TypeScript program demonstrates the capability handoff without inventing vendor-specific request fields. It reads two JSON bodies that conform to the live schemas for pdf.ocr and vector.upsert; the upsert template may contain the exact string $OCR_RESULT, which is replaced with the OCR response. Both calls use the same base URL and bearer key. It checks failures and backs off on 429 responses, honoring Retry-After when present.
const baseUrl = process.env.API_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
const ocrBodyJson = process.env.OCR_BODY_JSON;
const upsertTemplateJson = process.env.VECTOR_UPSERT_BODY_JSON;
if (!baseUrl || !apiKey || !ocrBodyJson || !upsertTemplateJson) {
throw new Error(
"Set API_BASE_URL, INFRAI_API_KEY, OCR_BODY_JSON, and VECTOR_UPSERT_BODY_JSON",
);
}
const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));
async function postOcr(body: unknown): Promise<unknown> {
for (let attempt = 0; attempt < 5; attempt += 1) {
const response = await fetch(`${baseUrl}/pdf/ocr`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify(body),
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await sleep(delayMs);
continue;
}
const payload: unknown = await response.json();
if (!response.ok) {
throw new Error(`OCR failed (${response.status}): ${JSON.stringify(payload)}`);
}
return payload;
}
throw new Error("OCR exhausted its retry budget");
}
async function postVectors(body: unknown): Promise<unknown> {
for (let attempt = 0; attempt < 5; attempt += 1) {
const response = await fetch(`${baseUrl}/vector/upsert`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify(body),
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
await sleep(Number.isFinite(retryAfter) ? retryAfter * 1_000 : 500 * 2 ** attempt);
continue;
}
const payload: unknown = await response.json();
if (!response.ok) {
throw new Error(`Vector upsert failed (${response.status}): ${JSON.stringify(payload)}`);
}
return payload;
}
throw new Error("Vector upsert exhausted its retry budget");
}
function injectOcrResult(value: unknown, ocrResult: unknown): unknown {
if (value === "$OCR_RESULT") return ocrResult;
if (Array.isArray(value)) return value.map((item) => injectOcrResult(item, ocrResult));
if (value && typeof value === "object") {
return Object.fromEntries(
Object.entries(value).map(([key, item]) => [key, injectOcrResult(item, ocrResult)]),
);
}
return value;
}
const ocrResult = await postOcr(JSON.parse(ocrBodyJson));
const upsertBody = injectOcrResult(JSON.parse(upsertTemplateJson), ocrResult);
await postVectors(upsertBody);
Why use externally supplied schema-valid bodies? The public discovery surface exposes full request and response JSON Schema plus runnable examples, while the request fields themselves are not fixed in the evidence available here. Guessing a field name in a copyable example would be worse than making the boundary explicit. Generate the two JSON values from discovery, validate them in CI, and keep this orchestration stable.
Infrai is one reasonable fit for this seam because OCR and vector upsert sit behind one REST API and one key. Swapping the vendor behind a capability need not change this orchestration contract. The same setup also avoids separate authentication and rate-limit handling at the document-to-vector handoff. Its API is self-describing: the public discovery surface needs no key, and it describes 295 routes across 20 modules; every documented capability has runnable examples in 10 languages. For this pipeline, that means plain HTTP from the existing batch worker, with no SDK to install, while the two schemas can be validated during development.
No adapter package.
Compare that with AWS Textract plus Pinecone: it means two signups, two credential sets, and glue that translates Textract output into Pinecone records. Tesseract plus Pinecone still needs Pinecone credentials and translation code, while Tesseract itself must be packaged, operated, and scaled by your team. That control can be desirable. It is also work that belongs in the throughput calculation.
Infrai is not a fit when policy requires direct vendor contracts, when an organization must run OCR fully inside its own environment, or when a team wants low-level control over Tesseract's processing. In those cases, direct Textract integration or self-operated Tesseract is the clearer boundary, despite the extra glue. That is a real limitation, not a footnote.
One warning: the code starts after signature verification. Keep that as a hard precondition. The OCR result must carry an internal document identifier that resolves to the stored verification event; do not rely on array position within a batch.
Isn't the platform audit trail enough?
Sometimes, yes. If the relying party's policy explicitly accepts the sending platform's evidence, the workflow stays within that platform, and no independent record is required for a dispute, duplicating verification may add ceremony without changing the decision.
Keep it boring.
But make that a policy decision, not an accidental consequence of integration convenience. A screenshot is especially weak operationally: it is hard to query, awkward to correlate with a content hash or internal document ID, and detached from the exact intake decision. A platform-generated audit report is better structured evidence, yet it remains a record produced by the platform whose assertion you are trusting.
For searchable clinical documents, the conservative rule is crisp: no receiver-owned verification event, no indexing. This does not determine legal validity by itself. It ensures the technical system can explain why it accepted a specific file.
Will independent checks destroy batch throughput?
They should not, if the pipeline is shaped correctly. Verification is a gate, not a serial lock around the whole batch. Verify documents independently, persist each result, and release accepted documents to bounded OCR concurrency. Short files should not wait for one large scan to finish.
Measure queue age, accepted documents per batch, verification outcomes, OCR completion count, and indexing completion count. Alert on missing transitions as well as explicit failures. That gives operators a diagram in words: received, verified, OCR complete, indexed. Four states. Easy to teach. Easy to audit.
Do not claim success merely because the final vector count increased. Duplicate delivery and retries can inflate counts, while a failed handoff can strand verified documents. Correlate every transition with the same internal document identity and make the indexing write idempotent according to the selected service's contract.
The final decision rule is narrow. Trust platform evidence when policy accepts platform-owned proof and the document remains inside that trust boundary. Verify independently when your organization must produce its own technical record, or when OCR and search will amplify the consequences of accepting the wrong file. In either case, log your own decision per document.
That trade-off is the whole decision.
References
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- Adobe Acrobat Sign audit reports: https://helpx.adobe.com/sign/using/audit-reports.html
- DocuSign certificate of completion: https://support.docusign.com/s/document-item?language=en_US&bundleId=gav1643676262430&topicId=jqw1578456669906.html
- Dropbox Sign audit trail: https://faq.hellosign.com/hc/en-us/articles/206521127-What-is-an-audit-trail
- AWS Textract documentation: https://docs.aws.amazon.com/textract/
- Tesseract OCR documentation: https://tesseract-ocr.github.io/
- Pinecone documentation: https://docs.pinecone.io/
Top comments (0)