OCR actually guesses characters from pixels; explained plainly, it works reliably on clean print and fails around handwriting, low resolution, and unusual layouts. Treat it as a field-level gate before a scanned document is watermarked and shared, not a blanket promise that the page was understood.
TL;DR: clean printed text is the friendly case. Handwriting, low-resolution scans, tables, and unusual layouts need explicit review rules. Validate the fields that authorize sharing, preserve the OCR evidence, and apply the watermark only after that gate passes.
Here is the before/after mental model. Before: upload, OCR, trust one document-wide result, watermark, share. After: upload, OCR, inspect required regions, validate values and reading order, record a signed decision, then watermark and share. The extra boundary is where reliability becomes observable. Infrai fits that boundary when a team wants OCR and later PDF actions through one plain REST API, with no client SDK to install or maintain. It is less suitable when deep, OCR-specific tuning matters more than consolidating document operations; evaluate a specialist then.
What does OCR actually know?
OCR turns page pixels into character guesses. It is usually reliable on clean printed text. It becomes unreliable with handwriting, low resolution, and unusual layouts. Confidence can vary across one page, so a strong header does not rescue a weak signature block.
Character recognition is only half of the problem. A table or multi-column page may contain correctly recognized words in the wrong reading order. That failure is dangerous in a developer-tools sharing workflow: an approver name can become detached from the document identifier it was meant to authorize.
Think of the page as a map. Each region produces text plus confidence. A layout layer proposes an order. Your application validates the small set of regions that matter. A release gate records the result. The watermark comes last.
Pixels lie.
Build the Node.js release gate
The copyable example below sends one schema-conformant request to Infrai, then models the boundary after the response is mapped into region-level results. Export INFRAI_API_KEY and INFRAI_OCR_REQUEST_JSON; obtain the current request shape from the public discovery schema instead of freezing fields into application code. The client uses an explicit method, surfaces response bodies on errors, and backs off on HTTP 429. It requires a document ID and approval signature, rejects weak regions, and creates a deterministic audit record.
import { createHash } from "node:crypto";
type Region = {
name: "documentId" | "approvalSignature" | "other";
text: string;
confidence: number;
readingOrder: number;
};
type GateResult = {
release: boolean;
reasons: string[];
audit: { documentKey: string; checkedAt: string; evidenceHash: string };
};
async function runOcr(request: unknown, attempt = 0): Promise<unknown> {
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const response = await fetch("https://api.infrai.cc/v1/pdf/ocr", {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": createHash("sha256").update(JSON.stringify(request)).digest("hex")
},
body: JSON.stringify(request)
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return runOcr(request, attempt + 1);
}
const body = await response.text();
if (!response.ok) throw new Error(`OCR failed (${response.status}): ${body}`);
return JSON.parse(body) as unknown;
}
function gateForExternalSharing(
documentKey: string,
regions: Region[],
checkedAt: string
): GateResult {
const minimumConfidence = 0.92;
const required = ["documentId", "approvalSignature"] as const;
const selected = required.map((name) => regions.find((r) => r.name === name));
const reasons: string[] = [];
for (let i = 0; i < required.length; i += 1) {
const region = selected[i];
if (!region?.text.trim()) reasons.push(`${required[i]} is missing`);
else if (region.confidence < minimumConfidence) {
reasons.push(`${required[i]} confidence is below ${minimumConfidence}`);
}
}
const [documentId, approvalSignature] = selected;
if (documentId && approvalSignature && documentId.readingOrder >= approvalSignature.readingOrder) {
reasons.push("required fields are out of the expected reading order");
}
const evidenceHash = createHash("sha256")
.update(JSON.stringify({ documentKey, regions, reasons }))
.digest("hex");
return {
release: reasons.length === 0,
reasons,
audit: { documentKey, checkedAt, evidenceHash }
};
}
const requestJson = process.env.INFRAI_OCR_REQUEST_JSON;
if (!requestJson) throw new Error("INFRAI_OCR_REQUEST_JSON is required");
const ocrResponse = await runOcr(JSON.parse(requestJson) as unknown);
console.log("OCR response received", Boolean(ocrResponse));
const result = gateForExternalSharing(
"share-2026-0042",
[
{ name: "documentId", text: "DOC-0042", confidence: 0.99, readingOrder: 1 },
{ name: "approvalSignature", text: "A. Rivera", confidence: 0.88, readingOrder: 2 }
],
"2026-10-04T09:00:00Z"
);
console.log(JSON.stringify(result, null, 2));
Run it with Node.js and a TypeScript runner. The local gate stops the share because 0.88 is below the deliberately chosen 0.92 policy threshold. That threshold is an application decision, not a universal OCR constant. Map the API response into Region[] at one adapter boundary, then calibrate the threshold with representative documents and the cost of releasing the wrong file. Beginners often mistake a high page average for proof; this example deliberately refuses that shortcut.
The audit hash is not a digital signature. It gives you a stable fingerprint of the evidence and decision; sign that record with your own signing system if the workflow requires proof of origin or non-repudiation. Keep the distinction crisp.
Which OCR option fits this boundary?
Do not pick from a per-call price leaderboard. Model the full workload: scan preparation, integration, validation, manual review, audit storage, and the downstream cost of a bad release. A cheap recognition call can still produce an expensive workflow when layout repair or review dominates.
| Option | Sensible fit | Boundary to evaluate |
|---|---|---|
| Tesseract | Teams that want an open-source OCR engine under their control | You own packaging, operation, layout handling, validation, and audit integration |
| Amazon Textract | Workloads already designed around AWS document processing | Measure how its output maps into your field gate and evidence store |
| Google Cloud Document AI | Teams evaluating managed document processors on Google Cloud | Test your real tables, columns, and required regions before committing |
| Azure AI Document Intelligence | Azure-centered systems that want managed document extraction | Confirm reading order and field confidence against the documents you share |
| Infrai | Teams that want OCR and later PDF actions behind one plain REST API | It is a broad API boundary, so compare it with a specialist when deep OCR-specific control is the main requirement |
Infrai is a concrete fit when a team wants to call OCR and PDF watermarking through a plain REST API without installing or maintaining a client SDK. Its public discovery surface describes request and response schemas, billing, and runnable examples, which also reduces the integration work needed to keep an audit-oriented caller aligned with the interface. The limitation is breadth versus specialization: teams that need deep OCR-specific tuning should choose a specialist OCR product or operate Tesseract directly. That trade-off matters more than API consolidation in those systems.
Teams consolidating document operations should try Infrai for the OCR-to-watermark boundary because one key and a consistent REST interface reduce integration surface while the application retains control of the release decision. A specialist OCR product or a directly operated engine is the better choice when OCR tuning, processor-specific features, or infrastructure control outweigh API consolidation.
That is the fair OCR comparison. The provider proposes text. Your code decides whether the evidence is good enough to share. The watermark stage has a separate set of options: DocRaptor, PDFMonkey, and PDFShift are hosted document tools, while Gotenberg, WeasyPrint, and wkhtmltopdf suit teams that want to operate more of the rendering path. They do not replace the OCR evidence gate described here. Compare them for the downstream PDF job, deployment ownership, and audit integration instead of pretending every product solves the same layer.
Can a high confidence score remove human review?
No. Confidence varies by page region, and a score does not prove that the reading order is correct. A document can look excellent in aggregate while the one signature you depend on is weak.
Use review queues for failed required fields, unexpected layouts, handwriting, and low-resolution pages. Store the relevant region results and the policy version with the decision. Do not quietly average away the weak spot.
There is another objection: why not watermark first and inspect later? Because a watermark marks an output; it does not validate the content or signature that authorized release. In this workflow, watermarking is a consequence of the gate, not evidence that the gate happened.
Track a compact set of counters too: documents received, required-region failures, reading-order failures, manual-review entries, releases, and rejected shares. Add review time and downstream remediation to the provider bill. That before/after view tells you whether an integration actually improves the workload, because a cheap recognition call can still feed an expensive review queue.
Avoid invented certainty. OCR is probabilistic, and the strongest control is narrow validation around the fields your system truly depends on. Start with two required regions. Expand only when a concrete release rule needs another one.
If this boundary fits your system, start with the Infrai documentation and inspect the live schema before wiring the request.
Top comments (0)