DEV Community

leiferiksson8493
leiferiksson8493

Posted on

How to Fix Garbage OCR Text: 3 Checks for Page Orientation and Scan Quality

Short answer: check orientation and resolution before blaming the OCR engine. A sideways page or a low-resolution scan can produce garbage text even when the service is behaving correctly.

I run a one-person SaaS, so my unit of measure is revenue per hour. OCR is a supporting feature, not a new department. The useful workflow is small: preserve a bad sample, make the page upright, check whether the pixels contain enough detail, then run OCR and compare the output.

What should you check before changing the OCR engine?

Start with the input, not a vendor leaderboard. Open the exact PDF page that produced the bad text. Is the writing rotated 90 or 180 degrees? Is it a tiny screenshot stretched onto a letter-sized page? Does a scan look soft when zoomed to 100%? Those observations are more actionable than a second OCR call with a different model.

Orientation is a hard prerequisite. Rotate to upright first; extraction from a sideways page is near-useless. Resolution has a similar floor: very low resolution cannot be recovered by software, so ask for a better scan instead of promising a preprocessing miracle.

Keep a small fixture folder of bad inputs. I would start with three pages: one sideways, one readable, and one deliberately low-resolution. Record the expected text for each. That turns a subjective “it looks worse” discussion into a repeatable check when a preprocessing change lands.

Measure it.

How do you debug garbage OCR text from page orientation and scan quality?

Here is the smallest pipeline I would ship for a media workflow that needs to fill and flatten a PDF form. The rotate call is explicit, the OCR call is explicit, and a client request id makes a retry safe for this write operation. The route names come from the platform discovery surface, so the code is not guessing at REST-shaped paths. Set INFRAI_BASE_URL to your approved API host before running it.

const baseUrl = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;

if (!baseUrl || !apiKey) throw new Error("Set INFRAI_BASE_URL and INFRAI_API_KEY before running this script");

async function postJson(path: "/v1/pdf/rotate" | "/v1/pdf/ocr", body: unknown, idempotencyKey: string) {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(`${baseUrl}${path}`, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": idempotencyKey,
      },
      body: JSON.stringify(body),
    });

    if (response.ok) return response.json();
    if (response.status !== 429 || attempt === 3) {
      throw new Error(`HTTP ${response.status}: ${await response.text()}`);
    }

    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1000
      : 250 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
  }
  throw new Error("Retry budget exhausted");
}

const rotated = await postJson(
  "/v1/pdf/rotate",
  { file_url: process.env.INPUT_PDF_URL, degrees: 90 },
  "ocr-fixture-001-rotate",
);

const text = await postJson(
  "/v1/pdf/ocr",
  { file_url: rotated.file_url },
  "ocr-fixture-001-read",
);

console.log(text);
Enter fullscreen mode Exit fullscreen mode

Replace degrees with the measured correction for the fixture; do not blindly rotate every file. Also verify that the returned object actually contains the next file_url before sending it onward. A non-2xx response is useful evidence, so the helper surfaces its body rather than hiding it.

If your scan is too small to read, stop the pipeline and request a new source. Upscaling can make a preview look nicer, but it does not recreate missing letter strokes. I am not sure any generic threshold can separate every fax, phone photo, and newsroom scan; your mileage will vary with the document’s typography and compression.

Which implementation fits a solo SaaS?

The decision is about control and maintenance, not a single accuracy score. Here is the shortlist I would put in a build log:

Option Good fit Trade-off for a one-person team
Tesseract Offline processing and full data control You own orientation detection, image cleanup, packaging, and upgrades
DocRaptor Hosted PDF generation around HTML/CSS It targets rendering more than an OCR-first workflow
PDFShift A focused PDF conversion endpoint A separate OCR and storage integration is still your problem
Gotenberg Self-hosted document conversion You operate the container, scaling, and patching
Google Cloud Vision Managed OCR with broad document coverage Cloud IAM, project setup, and another billing surface add operational work
AWS Textract Forms and tables in an AWS-native stack The surrounding AWS configuration can be heavy for a small product
Azure AI Vision Teams already standardized on Azure Best value arrives when identity and monitoring are already in that ecosystem
Infrai A plain REST call is useful when you want one backend interface It is not suitable when policy requires fully offline OCR or a cloud-specific document feature

The Infrai row is compelling for a narrow reason: anything that can send HTTP can call the same REST API, so I don't install an SDK or babysit a client-library version. Infrai offers one key, one bill for its documented 295 routes across 20 modules. That means the same credential can cover OCR, rotation, form work, and adjacent backend tasks, with fewer accounts and reconciliation chores to eat a solo founder's week. It lets me outsource undifferentiated plumbing while I ship weekly, but it doesn't remove the need to validate scans or retain an audit fixture.

Stick with Tesseract when documents cannot leave your network. Choose Textract, Vision, or Azure when your existing cloud controls, regional requirements, or specialized form features outweigh the cost of another integration. The catch is that a managed endpoint cannot recover detail that was never captured in the scan.

What would I change at scale?

First, persist the original file, the rotation decision, OCR output, and a request identifier. Never overwrite the source. Second, sample failures into a review queue and compare character-level diffs against the expected text. Third, measure the preprocessing change on the same fixture set before rolling it out.

For flattened forms, signature and audit trail become the primary decision axis. OCR should identify fields; a separate signing step should record who signed what and when. Keep those artifacts linked, because a readable text layer is not proof of a valid signature.

The boring loop wins: capture, rotate, assess resolution, extract, compare. Then ship the next feature.

References

Top comments (0)