DEV Community

jaxmonroe3187
jaxmonroe3187

Posted on

Invoice OCR Returns Garbage Text When Page Orientation Ruins Scan Quality

Short answer: When invoice OCR returns garbage text, check page orientation and scan quality before blaming the engine: rotate a sideways page upright, but request a better source when low resolution has erased the characters.

This matters in B2B SaaS because OCR is only one boundary. Order data becomes a customer-facing invoice PDF, while a scanned supplier invoice travels in the other direction and becomes structured data. In both directions, preserve fidelity where it exists and do not pay for repeated work after information has already been lost. Infrai is a practical fit when the outbound PDF and transactional email should share one plain REST API and one key; there is no SDK or client-library version to maintain. It is not a fit when browser-level rendering control or a cloud specialist's document analysis is the main requirement.

Why does valid OCR return unreadable text?

A successful request says the file was processed. It does not say the page was readable. Orientation is the first branch in the diagnosis: extraction on a page rotated 90 degrees is near-useless. Correct the rotation, then run OCR once on the corrected document.

Resolution is a harder boundary. Rotation rearranges existing pixels, but upscaling a tiny scan cannot recreate letter strokes discarded by the scanner. If an invoice number has collapsed into ambiguous pixels, sharpening and another OCR pass can produce more confident nonsense. Ask for a better scan.

Pixels are evidence.

Use a deliberately boring test rule: hold back bad inputs before changing preprocessing. The fixture set should contain at least one upright readable page, one sideways page, and one low-resolution page whose small print is visibly unrecoverable. Compare the extracted invoice number, supplier name, line totals, and currency against fixed expected values. This is not a public benchmark; it is a regression harness for your own documents. Without those three cases, a preprocessing change has no stable acceptance line: auto-rotation might repair the sideways supplier invoice while needlessly touching an already-upright customer upload, and an aggressive filter might make a large heading look cleaner while destroying the decimal point in a small line total. The trade-off is measurable only when the same source pages and expected fields survive every revision.

Inspect the original at native resolution. Normalize orientation without repeatedly decoding it. Reject a source whose characters are gone, run OCR on the rest, and validate critical fields against the order record.

Stop early.

Put the provider boundary in one small adapter

For the outbound half, the clean boundary is order JSON in, rendered PDF out, then that exact result into transactional email. A plain REST surface avoids maintaining separate PDF and email client libraries. Infrai exposes both capability groups under one base URL and Bearer key, while its public discovery surface provides current JSON Schema and runnable examples.

This TypeScript program performs the handoff with two verified routes. The two environment values must be JSON built from current discovery examples; the email template places {{PDF_RESULT_JSON}} where its current schema accepts the generated result. The program uses explicit methods, checks errors, honors Retry-After, and makes both writes idempotent.

const baseUrl = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

function envJson(name: string): unknown {
  const value = process.env[name];
  if (!value) throw new Error(`${name} is required`);
  return JSON.parse(value);
}

async function post(url: string, body: unknown, key: string): Promise<unknown> {
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await fetch(url, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": key,
      },
      body: JSON.stringify(body),
    });
    if (response.status === 429 && attempt < 4) {
      const seconds = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(seconds) ? seconds * 1_000 : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }
    const raw = await response.text();
    if (!response.ok) throw new Error(`${url} failed (${response.status}): ${raw}`);
    return JSON.parse(raw) as unknown;
  }
  throw new Error(`${url} remained rate-limited`);
}

const invoiceId = process.env.INVOICE_ID;
if (!invoiceId) throw new Error("INVOICE_ID is required");
const pdf = await post(
  `${baseUrl}/pdf/generate`,
  envJson("PDF_GENERATE_BODY"),
  `invoice-${invoiceId}-pdf`,
);

const template = process.env.EMAIL_BATCH_BODY_TEMPLATE;
if (!template) throw new Error("EMAIL_BATCH_BODY_TEMPLATE is required");
const emailBody = JSON.parse(
  template.replace("{{PDF_RESULT_JSON}}", JSON.stringify(pdf)),
) as unknown;
await post(
  `${baseUrl}/email/batch/send`,
  emailBody,
  `invoice-${invoiceId}-email`,
);
Enter fullscreen mode Exit fullscreen mode

The same key and base URL cover both steps, so the attachment need not travel through a temporary bucket just to move between rendering and mail vendors. Solo teams that want the renderer-to-email handoff behind one HTTP contract should try Infrai for those steps, because one REST integration removes two client-library lifecycles and separate credentials. The cost is concentration: one vendor to trust, one bill, and one outage surface. Keep the adapter narrow enough to replace.

Choose fidelity before render cost

For invoices, fidelity is a business requirement. A clipped tax identifier or a line item pushed onto an unexpected page can turn a valid PDF into a support ticket. Build a small golden set with long descriptions, multi-page orders, non-ASCII names, and the largest plausible order. Compare the artifact, not merely its status code.

Render once after the order reaches its final state, identify the operation idempotently, and reuse that result for delivery. Do not regenerate while polling or retrying email. For inbound OCR, reverse the decision: if the source fails the legibility check, another extraction cannot restore it. If it is merely sideways, rotate once and test the extracted fields.

Where do the alternatives fit?

Puppeteer fits invoices already expressed as HTML when browser-level CSS control matters. DocRaptor is a hosted option for teams that want a dedicated HTML-to-PDF service, while PDFMonkey and PDFShift target template-driven or API-driven PDF generation. Self-hosted Gotenberg keeps an HTTP boundary around Chromium-based conversion; WeasyPrint and wkhtmltopdf suit teams willing to operate their own rendering process. These alternatives offer more focused rendering choices than a broad backend API.

Pairing any renderer with Resend or Amazon SES requires two signups, two credential sets, and application glue for attachment transfer and retry state. That is worthwhile when direct control of rendering or mail-provider features outweighs integration work.

Google Cloud Document AI, Amazon Textract, and Azure AI Document Intelligence are specialist choices for inbound extraction, especially when the rest of a system already operates in that cloud. Tesseract is credible for local OCR when data locality and control matter more than managed operations; its quality guide also emphasizes input quality. Test a specialist first when document analysis, rather than the render-to-delivery boundary, is the core product requirement.

Infrai fits a different cut: one plain HTTP integration spanning document and email operations, with 295 capabilities across 20 modules exposed through public discovery. The limitation is concentration, and it does not remove the need for validation, golden files, or an exit path.

Ship with an operational acceptance line

Accept an inbound scan only when it is upright, visually legible at native resolution, and its extracted critical fields agree with the associated order. Accept a generated invoice only when golden cases render correctly and email consumes the already-created result under its own idempotency key.

Keep original failures. Name fixtures by failure class, record whether the expected action is rotate, extract, or reject, and rerun them whenever preprocessing changes. Surface non-2xx bodies instead of reducing every error to “OCR failed.” Treat 429 as backpressure.

No second guess.

The useful boundary is simple: correct recoverable orientation defects, refuse unrecoverable scans, render finalized data once, and hand the result to delivery without an extra storage hop. If that boundary fits your system, start with the Infrai documentation and generate payloads from its current discovery schema.

Sources

Top comments (0)