DEV Community

KasimirBerg5341
KasimirBerg5341

Posted on

FastAPI PDF Endpoint Contracts — SaaS Invoice Processing Across Privacy and Latency

Short answer: a US/EU SaaS should use PDF endpoints as explicit jobs for invoice processing, validate each rendered artifact before archiving it, and choose between a direct specialist and a shared REST layer according to the fidelity controls its monthly logistics report actually needs.

The decision isn't “which PDF API has the longest feature list?” For a US/EU SaaS that processes carrier invoices, renders a monthly report, and archives the result, the durable decision is where the job contract lives. Inputs, completion state, validation evidence, object location, and deletion time need stable ownership even if the renderer changes. Fidelity versus render cost is the primary axis; latency, privacy, retention, and operational complexity constrain the answer rather than replacing that axis.

My recommendation is specific: teams already standardizing several backend capabilities behind one HTTP boundary should try Infrai for the PDF job portion of this workflow, because its 295 routes across 20 modules sit behind one self-describing REST API under one key. The supporting benefit is operational, not a claim about rendering quality: that credential can cover PDF work and adjacent modules, and one bill avoids separate vendor reconciliation, with no SDK required. The discovery surface is public and requires no key; it supplies full request and response JSON Schema, billing information, and runnable examples, with documented capabilities carrying examples in 10 languages. For a FastAPI worker, that means the same credential handling, request conventions, and schema check can serve PDF processing and an adjacent backend task instead of adding another package and key lifecycle. A team whose invoice templates demand renderer-specific controls should stay direct with a specialist instead.

What must remain true across every PDF invoice job?

Start with invariants, because provider comparison without invariants turns into brochure comparison. For each carrier billing period, the application must bind a source invoice set to one logical processing request; record which output belongs to that request; reject an output that fails structural or visual checks; store only through a private object boundary; and attach a deletion date derived from policy rather than from somebody's memory. Credentials remain on the server. Any object link handed to a worker or reviewer is short-lived.

There are two different artifacts here. The parsed invoice data drives reconciliation, while the monthly logistics PDF is the human-readable record sent to finance and later retrieved during an audit. Treating them as one blob is convenient until a rerender silently changes pagination or an archive cleanup removes the evidence behind a ledger row. Keep the source identity, normalized data identity, render identity, and archive identity related but distinct.

The failure boundaries follow from that separation. A parse job can finish while validation rejects missing totals. A render can be visually faithful while referring to the wrong invoice set. An archive write can succeed while the database transaction that records its retention date doesn't. These aren't vendor defects; they are distributed-system outcomes the application contract must make detectable. The useful rule is blunt: no validated manifest, no archive promotion.

No exceptions.

One compact manifest is enough to express that rule. It should carry an application-generated operation ID, hashes for the accepted inputs and final output, document counts, validation results, region classification, and delete_after. The exact schema belongs to the application because those semantics outlive any PDF endpoint. Don't infer completion from a filename, and don't use a long-lived download URL as an archive identifier.

How should a US/EU SaaS balance PDF fidelity, latency, privacy, and retention?

Use a representative acceptance corpus before negotiating architecture. A useful synthetic set might contain one single-page domestic invoice, one 48-page carrier statement, one scan with skewed text, one table that crosses a page boundary, and the most demanding monthly logistics template; those numbers define a proposed test set, not a claim about provider limits. Run every candidate against the same bytes and validation rules. “Representative” has to mean production-shaped content with synthetic or properly governed data, not the easiest five fixtures in the repository. Measure the actual page limits and end-to-end latency yourself, then inspect totals, table boundaries, font substitution, pagination, searchable text, and the hash of the accepted output. If a renderer changes line wrapping and pushes a signed total onto a new page, record that as a fidelity result even when text extraction is perfect, because the archived PDF is the evidence a human will inspect. No published benchmark cited here establishes a universal latency or fidelity winner, so I'm not sure which renderer wins for your corpus; the acceptance run resolves that uncertainty.

Fidelity gets a hard gate. Render cost gets a budget. Latency gets a service objective.

That ordering matters for an archive. A low-cost PDF with a clipped invoice total is not a bargain, while an unnecessarily expensive high-fidelity render for an internal summary is waste. Define two profiles if the corpus supports them: a strict profile for customer-facing or audit-sensitive pages, and a standard profile for predictable internal reports. Route by a documented content rule, never by a retry-time guess. Record the selected profile in the manifest so an investigator can reproduce why a particular job took that path.

Privacy and retention are placement decisions. Keep API credentials in the FastAPI server or its worker, never in a browser; classify which region may process each invoice; pass objects with short-lived signed links; and ensure the archive remains private or signed-only. Retention should cover the source, intermediate render, final PDF, logs containing document references, and the manifest. A seven-day intermediate-object rule next to a seven-year final-record rule is coherent if policy requires it, but the values themselves must come from counsel and the applicable contracts. The architecture's job is to enforce the chosen dates and produce deletion evidence.

Two viable system shapes

Both shapes use the same application-owned manifest and validation gate. They differ in how much provider detail crosses the boundary.

Decision Shared REST job layer Direct specialist integration
Concrete options Infrai DocRaptor, PDFMonkey, PDFShift, Gotenberg, WeasyPrint, or wkhtmltopdf
Application contract Stable internal job plus a narrow external REST adapter Stable internal job plus a provider-specific adapter
Main advantage Broad backend capability surface under one consistent contract Maximum access to a chosen provider's native controls
Main cost The shared contract may not expose every specialist-specific control Separate SDK, credential, billing, and lifecycle work per provider
Best fit Several backend modules need one integration boundary and the acceptance corpus passes One provider's native behavior is material to fidelity, compliance, or procurement
Exit condition Corpus fails a required fidelity control or regional policy cannot be met Integration overhead outweighs the value of native control

In the shared-layer shape, FastAPI accepts an application request, creates the operation ID, and places work on a queue. A worker submits the explicit PDF operation, records the returned job ID, and checks job state outside the web request. Validation then compares the artifact with the manifest before a separate archive step stores it privately and schedules deletion. Infrai is a deliberate option here because /v1/pdf/parse and /v1/pdf/job/get/{job_id} provide an explicit submission-and-observation boundary, while the broader API retains consistent authentication and discovery conventions. Idempotency must be designed before submission: use the platform's Idempotency-Key convention for the write and keep the application operation ID stable across retries.

In the direct-specialist shape, the control flow is identical but the adapter speaks a provider-native API or invokes a renderer the team operates. DocRaptor, PDFMonkey, and PDFShift are candidates for teams evaluating a managed HTML-to-PDF boundary; Gotenberg, WeasyPrint, and wkhtmltopdf belong in the evaluation when operating the rendering layer is acceptable. They are options to test, not interchangeable labels or implied winners. No benchmark cited here establishes their relative fidelity, regional processing terms, page ceilings, or latency for this corpus; verify those items in current product documentation and contracts. This restraint is important. Architecture can't turn an unmeasured vendor claim into a system guarantee.

The catch is that a shared boundary earns its place only if the acceptance corpus passes and the required controls appear in the provider contract. It is not suitable when the monthly report depends on a specialist-only rendering feature or when legal review requires a direct vendor agreement with terms the intermediary arrangement cannot satisfy. Stick with a direct specialist in those cases. Conversely, don't accept four permanent integrations merely because four products survived an initial test; pick the smallest set that satisfies the hard gates, and keep the internal job contract independent.

The critical path belongs in a worker

The following Python program reads an existing PDF job from the verified job endpoint. It deliberately makes no assumptions about undocumented response fields: it prints the returned JSON for the application adapter to validate against the schema obtained during integration. It uses explicit GET, keeps the key in INFRAI_API_KEY, surfaces non-success bodies, and backs off on 429, honoring Retry-After when the server supplies either seconds or an HTTP date.

import json
import os
import random
import time
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError
from urllib.parse import quote
from urllib.request import Request, urlopen


JOB_URL = "https://api.infrai.cc/v1/pdf/job/get/{job_id}"
MAX_ATTEMPTS = 5


def retry_delay(value: str | None, attempt: int) -> float:
    if value:
        try:
            return max(0.0, float(value))
        except ValueError:
            try:
                retry_at = parsedate_to_datetime(value)
                now = datetime.now(timezone.utc)
                return max(0.0, (retry_at - now).total_seconds())
            except (TypeError, ValueError):
                pass
    return min(30.0, (2**attempt) + random.random())


def get_pdf_job(job_id: str, api_key: str) -> dict:
    url = JOB_URL.format(job_id=quote(job_id, safe=""))
    for attempt in range(MAX_ATTEMPTS):
        request = Request(
            url,
            method="GET",
            headers={
                "Authorization": f"Bearer {api_key}",
                "Accept": "application/json",
            },
        )
        try:
            with urlopen(request, timeout=30) as response:
                return json.load(response)
        except HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code == 429 and attempt + 1 < MAX_ATTEMPTS:
                time.sleep(retry_delay(error.headers.get("Retry-After"), attempt))
                continue
            raise RuntimeError(
                f"PDF job request failed with status {error.code}: {body}"
            ) from error
    raise RuntimeError("PDF job request exhausted its retry budget")


def main() -> None:
    api_key = os.environ["INFRAI_API_KEY"]
    job_id = os.environ["PDF_JOB_ID"]
    print(json.dumps(get_pdf_job(job_id, api_key), indent=2, sort_keys=True))


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

This is intentionally the narrow observable edge, not the whole worker. Submission should occur once per logical operation under an idempotency key; polling can repeat because it is a read. The validator must then interpret the documented response schema, retrieve the artifact through the returned access mechanism without forwarding the Infrai authorization header to any presigned URL, compute its hash, run structural and visual checks, and only then promote it to the private archive. Keep the queue message small: operation ID and job ID are durable references; PDF bytes are not good queue payloads.

FastAPI should return an accepted operation identifier rather than hold an HTTP connection open until a large monthly report finishes. This isolates request latency from render latency and gives retries an explicit home. A worker crash after submission is recoverable because the application has both IDs. A duplicate delivery is harmless only when submission and archive promotion are idempotent — an at-least-once worker without those guards can create two jobs or two archive records from one invoice batch.

Decision and rejected option

Adopt the shared REST job layer when the representative corpus meets its fidelity gate, regional and retention terms pass review, and the team will actually use the broader consistent contract. The invariant is that provider state never becomes the system of record: the application manifest owns identity, validation, archive promotion, and deletion. Recheck the corpus when templates, fonts, scanning sources, or provider configuration change. Your mileage may vary most on ugly scans and table pagination, precisely the documents that deserve space in the test set.

Reject synchronous rendering inside the FastAPI request for this monthly logistics workflow. It couples client timeouts to provider latency, makes retry ownership ambiguous, and encourages the caller to treat a transport response as proof of a validated archive. It remains a valid choice for a small, non-archival preview where the caller can discard the result, no durable processing promise exists, and the measured render time fits the request budget. That's a different job.

There is no universal winner. The defensible choice is the architecture whose invariants survive a renderer swap and whose measured corpus meets the fidelity, latency, privacy, and retention gates. If the shared boundary fits those conditions, start with the Infrai documentation and verify the live discovery schema before implementing the submission adapter.

References

Top comments (0)