Short answer: preserve the uploaded PDF, rotate only the pages that are provably sideways, render the corrected pages, and send those images to OCR. For a B2B SaaS monthly-report archive, that is the least complex path that keeps charts and tables faithful while making extracted text searchable. Put rotation before rasterization. Record the decision for every page.
One pass.
| Pick this approach | When it fits | Fidelity risk | Render cost |
|---|---|---|---|
| Apply PDF page rotation | The correct quarter-turn is known | Low; original page content stays intact | One corrected render per page |
| Normalize during rasterization | Rotation is needed only for OCR | The archived PDF and OCR view can disagree | One render per candidate angle |
| Quarantine for review | Orientation is uncertain or mixed | Lowest chance of silently changing the document | Extra queue and operator time |
The table is the operating rule. A rotation pipeline should not be a blind 90-degree transform over the whole file. Monthly reports often mix portrait summaries, landscape tables, and imported appendices. Treat orientation as page data, not document data.
How should Node.js rotate PDF pages before running OCR?
OCR consumes pixels, not the intent behind a PDF page. PDF defines pages, coordinate systems, and rotation as document semantics; the renderer turns those semantics into the image seen by the recognizer. If the pipeline renders first and repairs orientation later, it either runs recognition on sideways pixels or pays for another render and another recognition pass.
The useful mental diagram is: upload -> validate -> decide per page -> rotate -> serialize corrected PDF -> render page images -> recognize -> archive PDF plus text. Logs follow the same path. Each event carries a report ID, page number, chosen angle, decision source, render duration, and OCR duration.
This ordering also protects a subtle requirement: the archived artifact and the searchable representation should describe the same reading orientation. A support engineer opening page 12 should see the table in the orientation used to produce page 12's extracted text.
Start there.
Pick PDF rotation when the answer is known
Use page rotation when the application already has a trustworthy answer: a user corrected the preview, an upstream scanner supplied verified orientation, or a deterministic report template marks its landscape pages. Keep the angle to quarter-turns: 0, 90, 180, or 270. Reject everything else at the boundary.
This is the best fidelity-versus-cost trade for recurring monthly reports because it changes the viewing transform rather than rebuilding page content. It also creates one canonical corrected PDF for rendering and archiving. The trade-off is strictness. A wrong input angle produces a consistently wrong result, so the decision must remain visible in logs and job metadata.
Do not infer one angle for all pages from the first page. A 24-page report can legitimately contain both portrait narrative and landscape revenue tables. Per-page decisions are more verbose, but they prevent a common class of quiet corruption.
Mixed pages are normal.
Pick raster normalization when the archive must remain byte-stable
Sometimes retention policy requires preserving the uploaded bytes exactly. In that case, leave the source PDF untouched and rotate only the raster image passed to recognition. Archive the source, the orientation decision, and the OCR output as related artifacts.
There is a sharp edge: anyone opening the source may see a sideways page while search results reflect the corrected image. Make that distinction explicit in the data model. This option can also increase work if orientation detection tries several candidate angles. Measure rendered pages and recognition attempts, not only request latency.
Quarantine is the honest third option. If orientation confidence falls below your documented threshold, or pages contain conflicting visual signals, stop automatic processing and ask for review. Guessing creates searchable nonsense that can look successful from an HTTP status alone.
A minimal Express pipeline in TypeScript
The route below accepts a PDF body and a compact per-page rotation map such as 1:90,4:270. Page numbers are one-based at the API boundary and converted once. The PDF library applies page rotation; renderer and recognizer remain interfaces because those components depend on deployment constraints and should be tested independently.
import express, { Request, Response } from "express";
import { PDFDocument, degrees } from "pdf-lib";
type QuarterTurn = 0 | 90 | 180 | 270;
type PageImage = { page: number; png: Uint8Array };
type OcrPage = { page: number; text: string };
interface PdfRenderer {
render(pdf: Uint8Array): Promise<PageImage[]>;
}
interface TextRecognizer {
recognize(image: Uint8Array): Promise<string>;
}
interface Archive {
put(reportId: string, pdf: Uint8Array, pages: OcrPage[]): Promise<void>;
}
const allowed = new Set<number>([0, 90, 180, 270]);
function parseRotations(value: string | undefined): Map<number, QuarterTurn> {
const result = new Map<number, QuarterTurn>();
if (!value) return result;
for (const item of value.split(",")) {
const [rawPage, rawAngle] = item.split(":");
const page = Number(rawPage);
const angle = Number(rawAngle);
if (!Number.isInteger(page) || page < 1 || !allowed.has(angle)) {
throw new Error(`Invalid rotation: ${item}`);
}
result.set(page - 1, angle as QuarterTurn);
}
return result;
}
async function rotatePages(
input: Uint8Array,
rotations: Map<number, QuarterTurn>,
): Promise<Uint8Array> {
const pdf = await PDFDocument.load(input);
const pages = pdf.getPages();
for (const [index, angle] of rotations) {
const page = pages[index];
if (!page) throw new Error(`Page ${index + 1} does not exist`);
page.setRotation(degrees(angle));
}
return pdf.save();
}
export function buildApp(
renderer: PdfRenderer,
recognizer: TextRecognizer,
archive: Archive,
) {
const app = express();
app.post(
"/reports/:reportId/archive",
express.raw({ type: "application/pdf", limit: "25mb" }),
async (req: Request, res: Response) => {
try {
if (!Buffer.isBuffer(req.body) || req.body.length === 0) {
return res.status(400).json({ error: "A PDF body is required" });
}
const header = req.header("x-page-rotations");
const rotations = parseRotations(header);
const correctedPdf = await rotatePages(req.body, rotations);
const images = await renderer.render(correctedPdf);
const pages: OcrPage[] = [];
for (const image of images) {
const started = performance.now();
const text = await recognizer.recognize(image.png);
pages.push({ page: image.page, text });
console.info("ocr_page_complete", {
reportId: req.params.reportId,
page: image.page,
rotation: rotations.get(image.page - 1) ?? 0,
durationMs: Math.round(performance.now() - started),
});
}
await archive.put(req.params.reportId, correctedPdf, pages);
return res.status(202).json({
reportId: req.params.reportId,
pages: pages.length,
});
} catch (error) {
const message = error instanceof Error ? error.message : "Unknown error";
return res.status(422).json({ error: message });
}
},
);
return app;
}
The 25mb limit is an application policy in this example, not a PDF standard. Choose a limit from observed report sizes and infrastructure constraints. For larger files, replace request buffering with durable object storage and a queued worker; the ordering of rotate, render, and recognize stays the same.
The response is 202 because archival work can outlive the HTTP request. In production, make the report ID an idempotency boundary, cap page count and decompressed output, and validate that the bytes are actually a supported PDF before scheduling expensive work. Never use a client filename as a storage path. Test the rotation function with a fixture containing portrait and landscape pages. Assert page count, each resulting rotation value, and rejection of page zero, missing pages, and 45 degrees. Then run a visual regression on a page with a chart, small labels, and a table; text equality alone cannot reveal clipped axes or shifted annotations. Observability needs one counter for accepted reports, one for rejected reports by reason, histograms for render and recognition duration, and a count of OCR attempts per page. A useful dashboard puts queue age beside page throughput, then splits failures into input validation, rendering, recognition, and archival groups so an on-call engineer can locate the stalled stage without reading report contents. Alert on sustained changes in failure ratio or queue age rather than a single slow page. Do not log extracted report text because B2B reports can contain customer and financial data.
Keep payloads private.
Limits worth keeping visible
Page rotation cannot repair skew, perspective distortion, low resolution, a broken font mapping, or a scan whose content is already baked sideways into a larger image. Those require different normalization steps and should have distinct metrics. Rotation also does not prove OCR quality. Sample outputs against labeled pages and track character or word error using a stable evaluation set.
Keep both artifacts only when retention rules permit it. Otherwise, define which PDF is authoritative, delete temporary page images after recognition, and record the deletion outcome. Fidelity wins when a human can open the archived report and reconcile it with the extracted text; render cost stays controlled when each accepted page follows one measured path.
Further reading
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- Express API reference: https://expressjs.com/en/4x/api.html
- Node.js performance measurement APIs: https://nodejs.org/api/perf_hooks.html
- pdf-lib API documentation: https://pdf-lib.js.org/docs/api/
Top comments (0)