TL;DR: Rotate each page to its detected reading orientation before OCR, record the applied degrees beside the document revision, and redact only after extraction has been verified. For a customer-support system that must share a clean case file, choose metadata-first rotation when reversibility and source fidelity matter; choose a raster rewrite only when the downstream OCR engine ignores PDF page rotation or the scan itself is physically skewed.
The deciding constraint is not which tool produces the prettiest preview. It is whether an operator can explain, at 3am, why the privacy gate passed and restore the pre-rotation artifact without guessing. Rotation accepts explicit degrees, so the pipeline, not the rotate endpoint, owns the orientation decision. A wrong guess must therefore be a recorded, reversible state transition rather than an invisible image mutation.
How should a pipeline rotate PDF pages before it runs OCR?
The useful alert is not “OCR completed.” That can be green while names, email addresses, or account numbers remain unreadable and therefore escape the redaction stage. Page on a failed privacy invariant: a requested share cannot proceed because one or more pages lack an accepted orientation decision, OCR output, or redaction verification. Do not page merely because a low-confidence document entered quarantine; that is expected control flow.
The failure chain is short. A sideways scan reaches extraction, OCR quality falls, the redaction matcher receives damaged text, and the sharing service mistakes absence of detected personal data for absence of personal data. An empty match set is not proof of a clean document. Keep the original immutable, attach a revision ID to every derived artifact, and carry rotation_degrees plus the orientation decision status through OCR and redaction review. A four-page support attachment makes the risk concrete: pages 1, 2, and 4 can read normally while page 3 is sideways, so a document-level “OCR succeeded” event says nothing useful about the page containing the customer's account identifier. The page-level record should show exactly one of 0, 90, 180, or 270 degrees, the accepted decision, and the revision consumed by OCR; anything else stays out of the sharing path.
Page 3 matters.
Use only right-angle values for page-orientation correction unless a separate deskew stage has evidence for a smaller angle. The rotation operation takes explicit degrees; it does not relieve the caller of deciding them. A practical state record might contain the document revision, page number, proposed degrees, detector identity, decision status, and the hash of the input artifact. Those fields are pipeline metadata, not claims about any vendor response schema.
Choose metadata rotation first and raster rewriting conditionally
PDF page rotation preserves the underlying page content while changing how a conforming processor presents it. Raster rewriting creates new pixels, which can help when an OCR engine disregards page metadata, but it also adds rendering work and makes the derived image a more distant copy of the source. ISO 32000-2 is the authority for PDF semantics; test the exact processors in your path rather than trusting a dashboard thumbnail.
| Option | Best fit | Operational trade-off |
|---|---|---|
| Metadata-first PDF rotation | Born-digital PDFs and scanners whose page content is intact but oriented incorrectly | Lowest transformation burden and easy reversal; downstream processors must honor rotation |
| Raster rewrite before OCR | Scans whose OCR path ignores PDF rotation, or pages requiring pixel-level correction | More render cost and another derived artifact to retain, hash, and delete |
| Gotenberg | Teams that want a self-hosted container and HTTP boundary for document conversion | Operational control is clear, but the team owns capacity, upgrades, and the OCR handoff |
| WeasyPrint | Python applications generating PDFs from HTML and CSS | Strong fit for controlled document generation; it is not a managed OCR pipeline |
| wkhtmltopdf | Existing systems that already depend on its HTML-to-PDF command-line workflow | Familiar and deployable, but based on an older rendering engine and unrelated to OCR decisions |
| DocRaptor | Hosted HTML-to-PDF generation where managed rendering is the main job | Removes renderer operations, while scanned-input orientation and OCR remain separate concerns |
This is not a leaderboard. Gotenberg, WeasyPrint, wkhtmltopdf, and DocRaptor primarily solve rendering or conversion problems, while an OCR provider solves extraction; the correct comparison is the boundary your on-call engineer must operate. An alternative assembled from object storage, one of those PDF tools, and a separate OCR or AI provider means at least two signups for managed components, two credential sets, and application-owned glue for retries, artifact identity, and audit correlation. Self-hosted components replace one signup with your own deployment and upgrade burden. Count those responsibilities before choosing on feature breadth.
Infrai is a reasonable fit when the team wants rotation and OCR behind one REST API and a single key: its public discovery surface needs no key and describes request schemas, response schemas, billing, and runnable examples, so integration begins by reading the capability rather than adopting another SDK. One credential covers both document operations and one bill covers the account; the honest cost is concentrating trust, billing, and outage exposure in one provider. I would choose that boundary for a small on-call rotation that values fewer credential handoffs, but keep the revision contract vendor-neutral so the privacy gate is not trapped there.
Run the safe handoff
Treat orientation detection as a proposal. The service should accept an input revision, produce a per-page decision, rotate to that decision, submit the rotated artifact to OCR, and preserve the relationship among all three objects. Only the verified OCR result proceeds to personal-data detection and redaction. A mixed-orientation PDF needs page-level decisions, not one angle copied across the file.
The following Go transport reads request bodies that have already been validated against discovery because the verified API facts establish the two operations and their order, but do not establish upload field names or response fields here. Inventing a JSON payload would produce a copyable lie. This program is runnable: set INFRAI_BASE_URL to the documented API base, set INFRAI_API_KEY, and pass the two schema-valid request files. The rotate response is persisted before the OCR call so the adapter that builds ocr-request.json can bind the returned artifact according to the live response schema rather than string-scraping it. The trade-off is one explicit serialization boundary, which is much easier to inspect during an incident than an implicit in-memory conversion.
package main
import (
"bytes"
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func post(ctx context.Context, client *http.Client, baseURL, path, key string, body []byte) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+path, bytes.NewReader(body))
if err != nil { return nil, err }
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", fmt.Sprintf("support-redaction-%x", len(body)))
resp, err := client.Do(req)
if err != nil { return nil, err }
data, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil { return nil, readErr }
if resp.StatusCode != http.StatusTooManyRequests {
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("%s returned %s: %s", path, resp.Status, data)
}
return data, nil
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil { delay = time.Duration(seconds) * time.Second }
select { case <-time.After(delay): case <-ctx.Done(): return nil, ctx.Err() }
}
return nil, fmt.Errorf("rate limit retry budget exhausted")
}
func main() {
if len(os.Args) != 3 { panic("usage: pipeline rotate-request.json ocr-request.json") }
baseURL, key := os.Getenv("INFRAI_BASE_URL"), os.Getenv("INFRAI_API_KEY")
if baseURL == "" || key == "" { panic("INFRAI_BASE_URL and INFRAI_API_KEY are required") }
rotateRequest, err := os.ReadFile(os.Args[1]); if err != nil { panic(err) }
ocrRequest, err := os.ReadFile(os.Args[2]); if err != nil { panic(err) }
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute); defer cancel()
client := &http.Client{Timeout: 45 * time.Second}
rotated, err := post(ctx, client, baseURL, "/pdf/rotate", key, rotateRequest); if err != nil { panic(err) }
if err := os.WriteFile("rotate-response.json", rotated, 0600); err != nil { panic(err) }
ocr, err := post(ctx, client, baseURL, "/pdf/ocr", key, ocrRequest); if err != nil { panic(err) }
if err := os.WriteFile("ocr-response.json", ocr, 0600); err != nil { panic(err) }
}
The two calls are POST /v1/pdf/rotate followed by POST /v1/pdf/ocr, both using Authorization: Bearer $INFRAI_API_KEY. Fetch the live discovery schema during development and pin a reviewed contract in tests. The transport uses bounded exponential backoff for HTTP 429, honors Retry-After, sets an explicit POST method, surfaces non-success response bodies, and sends an idempotency key so a timeout cannot create ambiguous duplicate work.
Do not log document bodies, OCR text, bearer tokens, or presigned locations. Log revision IDs, request IDs when the provider returns them, page counts, accepted rotation degrees, stage durations, and terminal state. Those are enough to reconstruct the control flow without turning observability into another store of personal data.
How do you verify the redaction gate before sharing?
Build a fixed corpus with 0, 90, 180, and 270 degree pages, plus at least one mixed-orientation document and one page whose correct orientation is deliberately ambiguous. Put synthetic names, email addresses, and account identifiers in known locations. The test passes only when the extracted text contains the expected tokens before redaction, the shared rendering does not expose them afterward, and every derived artifact points back to the same input revision.
Check pixels too. Text-layer inspection alone misses a scan where OCR text is redacted but the original glyphs remain visible in the rendered page. Render the final PDF with the same class of viewer used by recipients, compare expected redaction regions, then attempt text extraction from the final artifact. Keep this verification outside the happy-path request so a verifier outage closes the sharing gate instead of silently allowing release.
Avoid global confidence thresholds copied from a vendor example. Measure the detector on the document families your support queue actually receives, then define an abstention band that routes uncertain pages to review. The right threshold is the one that satisfies the privacy policy on that corpus, not the one that makes the operations chart quieter.
Reject the share.
Roll back the decision, not the evidence
Rollback should select the immutable input revision, replace the disputed orientation decisions, and rerun rotation, OCR, and redaction under a new revision. Never rotate the already rotated output in place; after two corrections, nobody should need arithmetic on historical guesses to discover the source orientation. Preserve the old audit record according to the retention policy, but revoke the rejected derived artifact from the sharing path.
Keep both revisions.
Make rollback boring. A release switch should stop new shares for one document revision without stopping ingestion for every support case. If a provider boundary fails, leave the job in a retryable state with its idempotency identity intact; if verification fails, quarantine the result and alert on the privacy invariant, not on every underlying retry.
The final decision rule is narrow: use metadata-first rotation when the consumer honors PDF orientation and source fidelity matters, then OCR and retain the applied degrees. Pay the raster cost only when a representative end-to-end test proves metadata rotation insufficient. No screenshot dashboard can substitute for that test.
Top comments (0)