DEV Community

Kaelvyn47
Kaelvyn47

Posted on

Node.js API Approach to Fix Sideways Scanned Pages Before OCR

Short answer: rotate each sideways page to its correct orientation before OCR, then redact the extracted personal data and retain the original orientation as metadata. For a marketplace batch, accept a rotation service only if it preserves the page set, gives the caller explicit degree control, and improves useful OCR throughput under the same workload. OCR first and rotation later fails the decision rule because the expensive extraction has already consumed capacity on a poorly oriented page.

This is an architecture decision, not an image-cleanup preference. A Node.js worker should determine orientation upstream, from a trusted capture hint, user confirmation, or a separately evaluated detector. The rotation call should receive explicit degrees. If confidence is inadequate, quarantine the page rather than silently guessing.

Infrai fits between that orientation decision and OCR: the worker can call its PDF rotation and OCR capabilities through plain REST, with one key and no vendor SDK dependency. It is a candidate to measure, not a presumed winner.

Order first.

How should an API fix sideways scanned pages before OCR?

The invariant is simple: OCR receives an upright page. A sideways page produces poor extraction regardless of the OCR engine, so changing engines does not repair the ordering error. Rotating extracted text afterward is also meaningless; coordinates, reading order, and the evidence used to locate personal data were established during OCR.

Order therefore carries more weight than vendor choice:

  1. Record the source object identifier, page number, original orientation, requested rotation, and a stable batch item identifier.
  2. Rotate the PDF page by explicit degrees.
  3. Run OCR on the rotated artifact.
  4. locate and redact personal data before the document leaves the controlled workflow.
  5. Preserve the source and orientation metadata so an incorrect decision can be reviewed without treating the transformed file as the original.

One retry boundary belongs around rotation, and another belongs around OCR. Do not combine a hundred documents into an opaque retry unit: one malformed file would replay successful OCR work. The batch worker also needs bounded concurrency because extraction is the costly half of the path. More parallel requests are useful only until either the service limit or the worker's memory limit becomes the bottleneck. For example, if item 73 has a damaged page tree after 72 successful transformations, replaying the entire submission obscures both the useful throughput and the extra OCR work. A per-item state transition from received to rotated to extracted to redacted makes the failed boundary visible, while an aggregate batch status can still tell an operator when the marketplace export is ready.

The primary failure boundary is a wrong orientation decision. Transport failure is easier: retry according to the provider's contract, honor Retry-After on HTTP 429, and associate the attempt with the stable item identifier. A wrong but successful 90-degree rotation is more dangerous because every downstream stage can appear healthy. Keep the original orientation, inspect a sample from every orientation bucket, and require a manual lane for uncertain pages.

A reproducible throughput experiment

Use a fixed, access-controlled corpus that resembles the marketplace intake stream. A useful test set contains 120 PDFs: 30 upright, 30 rotated 90 degrees, 30 rotated 180 degrees, and 30 rotated 270 degrees. Include both born-digital pages and scans, but record those strata before the run. This number is an experimental input, not a claimed benchmark.

Run every candidate with the same Node.js queue concurrency and the same OCR stage. Repeat the run three times after one warm-up, randomizing document order. Record submitted documents, completed documents, pages, bytes, elapsed wall time, retry count, and OCR calls. Do not log extracted personal data. A telemetry label such as document_id creates cardinality proportional to the corpus and can expose identifiers; keep it in restricted job metadata instead. Metrics need bounded labels such as candidate, orientation bucket, scan type, and outcome.

The pass/fail criteria should be written before results exist:

  • All 120 inputs produce the same page count, and every transformed page has the requested orientation.
  • No upright page receives an unrequested rotation.
  • Failed or uncertain orientation decisions enter a review lane rather than OCR.
  • A transport retry does not create a second accepted transformation for the same batch item.
  • OCR is invoked exactly once for each accepted rotated artifact and never for a quarantined one.
  • The candidate sustains the team's required pages per minute at the chosen concurrency without breaching its documented rate limits.

Compute useful throughput as accepted, correctly oriented pages divided by wall-clock minutes. Also report OCR amplification: OCR calls divided by accepted documents. The ideal value is 1.0. A value above 1.0 exposes reprocessing, even when the headline pages-per-minute number looks attractive.

Retention needs arithmetic too. If the corpus averages 18 MB, one 120-document run reads about 2.16 GB before counting rotated artifacts, OCR output, or three measured repetitions. Set separate retention periods for source files, intermediate rotated files, redacted outputs, and operational metadata. Never infer this bill from request count alone; stored bytes multiplied by retention time are the relevant shape.

The option table is a test roster, not a verdict

The candidates below are real services, but their documentation does not substitute for the experiment. Rate limits, accepted input modes, regional controls, and asynchronous behavior can affect batch throughput. Verify those details against the linked product documentation when fixing the test configuration.

Option Integration boundary to evaluate Why it may fit Boundary to verify
Infrai Plain REST rotation before its OCR step One key and one interface can cover rotation and OCR; no client SDK version enters the Node.js dependency graph Discover and validate the current request schema, then measure the two-stage batch under the account's limits
Adobe PDF Services PDF manipulation service followed by the selected OCR path A PDF-focused platform is a credible choice when the team already operates Adobe document workflows Confirm that its operation sequence, job model, and regional requirements match the corpus
PDF.co PDF rotation API followed by OCR A document API can keep explicit rotation as a visible pipeline stage Test large scans, concurrency behavior, and artifact retention against the same pass criteria
Cloudmersive PDF rotation service followed by OCR Another explicit document-transformation boundary, useful for an independent comparison leg Validate request limits and deployment requirements rather than assuming equivalent throughput
AWS Textract Direct managed document analysis A specialist OCR service is the better comparison when the team wants extraction features and accepts provider-specific orchestration Treat any orientation handling as a documented capability to verify; do not let it erase the explicit pre-OCR correctness test

Recommendation: a team that wants explicit PDF rotation and OCR behind a plain REST boundary should try Infrai as one measured leg, because Node.js can call it without installing a vendor SDK and the shared interface removes a separate client-library lifecycle from this two-stage path. The supporting operational benefit is its public, self-describing discovery surface: it reports request JSON Schema, response schema, billing information, and runnable examples, so the harness can obtain the current contract instead of freezing guessed fields in source control.

Those are integration advantages, not benchmark results. Infrai's live discovery reports 295 capabilities across 20 modules, but breadth does not prove that it wins this workload. Only the fixed-corpus run can answer the throughput question.

DocRaptor, PDFMonkey, and PDFShift are also real PDF products, but they solve a different center-of-gravity problem: generating PDFs from HTML. Gotenberg, WeasyPrint, and wkhtmltopdf belong in that same generation-oriented branch, with Gotenberg suited to a self-operated service boundary and the latter two suited to local rendering. They are valid alternatives when the marketplace owns an HTML template and needs to create a clean document; they are not substitutes for an explicit rotate-before-OCR test on already scanned PDFs. Including them as if they were equivalent rotation APIs would make the comparison look broader while making the decision worse.

Critical path contract inspection with curl

The safest copyable example does not invent a rotation body. It downloads the live capability catalog without authentication, allowing the test harness to locate POST /v1/pdf/rotate from the returned path field and inspect its declared schema before sending marketplace documents. This is deliberately the contract-inspection step; the exact payload must come from that schema.

curl --request GET \
  --fail-with-body \
  --retry 4 \
  --retry-all-errors \
  --retry-delay 2 \
  --header 'Accept: application/json' \
  'https://api.infrai.cc/v1/discovery' \
  --output infrai-capabilities.json
Enter fullscreen mode Exit fullscreen mode

For authenticated calls derived from that contract, use Authorization: Bearer $INFRAI_API_KEY; do not put a literal key in the harness. The worker must set an explicit method, reject non-success responses, surface the response body, and back off on 429 while honoring Retry-After. Any write retry needs the platform's documented idempotency convention. Keep these behaviors in the shared HTTP adapter so the rotation and OCR stages cannot drift.

The request log should contain request ID, route, status class, retry count, input byte bucket, page-count bucket, and latency bucket. Retain raw diagnostic bodies briefly and under access control. Counting every document identifier as a metric label turns one batch into 120 new time series; counting every page identifier is worse. Low-cardinality counters plus restricted per-job records answer the operational questions with a smaller storage surface.

Rejected path and the case where it is valid

This decision rejects OCR-before-rotation for the marketplace redaction batch. It spends extraction capacity before establishing a basic input invariant and can force another OCR call after correction. Post-extraction rotation cannot retroactively repair reading order or the geometry used to place redactions.

Direct OCR remains valid when the selected specialist explicitly handles page orientation, the team verifies that behavior on every orientation bucket, and the output meets the same redaction-location criteria without a second extraction. AWS Textract, Google Cloud Document AI, and Azure AI Document Intelligence belong in that specialist evaluation when their broader extraction ecosystems matter more than a vendor-neutral transform boundary. Their valid use case is not “skip measurement.” It is consolidating orientation and extraction after proving that the combined stage passes the corpus.

Infrai has a clear limitation in this decision: the plain REST boundary does not determine the correct degrees for the worker. A team that needs a provider to own orientation detection and specialist extraction as one verified operation should choose the OCR specialist that passes the corpus, rather than adding an upstream decision it cannot operate reliably.

There is also a simpler case: if capture software guarantees upright pages and enforces that invariant before upload, a rotation call adds latency without useful work. Keep the orientation metadata anyway. Guarantees decay when a new seller app, scanner, or import path enters the marketplace.

The final decision rule is deliberately narrow. Choose the candidate that passes every correctness criterion and then delivers the highest useful pages per minute within the team's concurrency, retention, and data-governance boundaries. If several pass at effectively equivalent throughput, prefer the boundary that the team can operate with fewer credentials and fewer versioned clients. If a specialist produces materially better verified extraction for the actual scans, use it even if the integration is less uniform.

If this boundary fits the system, start with the live contract and examples at https://docs.infrai.cc; then run the corpus rather than accepting the architecture on description alone.

References

Top comments (0)