DEV Community

SyltharWave2946
SyltharWave2946

Posted on

How to Use PDF Endpoints: US/EU SaaS Image Assets Under Load

Short answer: A US/EU SaaS extracting image assets from PDFs should use an explicit asynchronous job contract, validate both the input and every returned artifact, benchmark representative documents at the intended concurrency, and delete source and derived objects on a written retention schedule.

For an e-commerce system turning scanned invoices, packing slips, and return forms into searchable text, endpoint selection starts with the bill rather than the API syntax. The bill contains PDF-processing calls, downstream OCR, object storage, data transfer, and retries. The dominant term is whichever product of units x unit price is largest at production volume; don't assume it is the PDF call. Embedded catalog photography can make retained image bytes dominate, while repeated OCR after ambiguous job outcomes can make retry work dominate. Measure both before signing a contract.

The practical change is to retain fewer ambiguous artifacts: keep the private source PDF long enough to replay a failed validation, keep validated extracted images only until OCR and indexing are accepted, and retain the searchable text plus an audit manifest for the business period. This removes routine storage and reprocessing work. The catch is real: after deletion, a parser regression or a newly discovered OCR requirement may force the customer to upload the document again.

What does the bill say about retention and throughput?

Build the estimate from workload variables, not a vendor's smallest demo file. For each document class, record pages per document, source bytes, extracted-image bytes, processing calls, OCR minutes or pages, retry rate, and days retained. Separate US and EU cohorts because residency decisions can change where objects and processing live, even when the application presents one queue. No supplied evidence establishes a particular provider's regional fit, so residency, subprocessors, and transfer behavior must be verified in the provider's current contract and documentation.

This small calculator is deliberately vendor-neutral. Put the rates from current quotes into environment variables and replace the planning values with measurements from your own corpus. Its useful output is the ranking of cost terms, not the sample total.

import os
from dataclasses import dataclass


@dataclass(frozen=True)
class Workload:
    documents: int
    pages_per_document: float
    processing_calls_per_document: float
    retained_gib_months: float


def env_rate(name: str) -> float:
    return float(os.environ.get(name, "0"))


def estimate(workload: Workload) -> dict[str, float]:
    pages = workload.documents * workload.pages_per_document
    calls = workload.documents * workload.processing_calls_per_document
    return {
        "pdf_processing": calls * env_rate("PDF_USD_PER_CALL"),
        "downstream_ocr": pages * env_rate("OCR_USD_PER_PAGE"),
        "private_storage": workload.retained_gib_months
        * env_rate("STORAGE_USD_PER_GIB_MONTH"),
    }


if __name__ == "__main__":
    planning_case = Workload(
        documents=10_000,
        pages_per_document=4.0,
        processing_calls_per_document=1.0,
        retained_gib_months=80.0,
    )
    costs = estimate(planning_case)
    for name, amount in sorted(costs.items(), key=lambda item: item[1], reverse=True):
        print(f"{name}: ${amount:,.2f}")
Enter fullscreen mode Exit fullscreen mode

Run at least three retention cases: source plus images for the full business period, source briefly plus images until OCR acceptance, and audit manifest plus text only after acceptance. The middle case is usually the first policy worth testing because it preserves a bounded replay window without treating every intermediate bitmap as a permanent record. "Usually" matters here. I'm not sure which term will dominate your bill until document sizes, OCR charging units, retry rates, and live quotes are filled in; a spreadsheet without those four inputs is theater.

Consider what the calculator's 10_000-document planning case is meant to expose. Suppose the sample set contains clean one-page return slips alongside four-page invoices and long supplier catalogs; do not average those into one synthetic PDF and call the result representative. Run each class separately, capture source and extracted bytes at the object boundary, count every OCR page accepted, and record a retry only when it actually creates charged or retained work. Then multiply each class by its forecast share and compare the terms. If retained images lead, shorten only their lifetime and rerun the estimate; if OCR pages lead, retention tuning is a distraction, so test whether validation can reject blank or duplicate pages before OCR; if processing calls lead because a timeout causes resubmission, fix reconciliation and idempotency before negotiating rates. Finally, rerun the same calculation for a peak day rather than a monthly average, since throughput capacity and monthly cost answer different questions. None of these planning values is a benchmark or a price claim. They are controls for discovering which measurement is missing, and they prevent an attractive per-call quote from hiding a larger downstream term.

How should a US/EU SaaS balance PDF image extraction latency under load?

Use a two-dimensional acceptance test: fidelity gates decide whether an output is usable, and latency distributions decide whether the system can keep up. A median hides queue saturation. Capture p50, p95, and p99 from client-observed submission-to-acceptance time at each concurrency level, plus completed pages per minute and the count of retries. Stop increasing concurrency when throughput flattens, p95 climbs beyond the service objective, or 429 responses become sustained rather than occasional.

Keep the sample stratified. Scanned return labels, image-heavy supplier catalogs, digitally generated invoices, rotated pages, and password-protected inputs exercise different failure modes. Validation should check the PDF before submission, then check artifact count, decodability, dimensions, checksums, and the mapping from each artifact back to its document and page. Fidelity isn't "the request returned 200"; it is evidence that the assets needed by OCR arrived intact and can be audited.

The following harness is runnable as-is against a local test double or any provider adapter exposing the same internal contract. It avoids encoding a vendor-specific request schema that may change. The adapter must return 202 with a job location for accepted work, 429 with an optional Retry-After, or another status with a useful body. Set BENCHMARK_URL to a non-production endpoint containing a representative, non-sensitive corpus.

import asyncio
import json
import os
import random
import time
import urllib.error
import urllib.request


def submit_once(url: str, document_id: str) -> tuple[int, dict, dict]:
    body = json.dumps({"document_id": document_id}).encode("utf-8")
    request = urllib.request.Request(
        url,
        data=body,
        headers={
            "Content-Type": "application/json",
            "Idempotency-Key": f"extract-{document_id}",
        },
        method="POST",
    )
    try:
        with urllib.request.urlopen(request, timeout=30) as response:
            return response.status, dict(response.headers), json.load(response)
    except urllib.error.HTTPError as error:
        payload = json.loads(error.read().decode("utf-8") or "{}")
        return error.code, dict(error.headers), payload


async def submit(url: str, document_id: str) -> float:
    started = time.monotonic()
    for attempt in range(6):
        status, headers, payload = await asyncio.to_thread(
            submit_once, url, document_id
        )
        if status == 202:
            if "job_url" not in payload:
                raise ValueError("accepted response omitted job_url")
            return time.monotonic() - started
        if status != 429:
            raise RuntimeError(f"submission failed: status={status}, body={payload}")
        retry_after = headers.get("Retry-After")
        delay = float(retry_after) if retry_after else min(2**attempt, 30)
        await asyncio.sleep(delay + random.uniform(0, 0.25))
    raise RuntimeError("rate limit retry budget exhausted")


async def main() -> None:
    url = os.environ["BENCHMARK_URL"]
    document_ids = [f"representative-{index:04d}" for index in range(100)]
    latencies = await asyncio.gather(*(submit(url, item) for item in document_ids))
    ordered = sorted(latencies)
    for percentile in (50, 95, 99):
        index = min(len(ordered) - 1, round((percentile / 100) * len(ordered)) - 1)
        print(f"p{percentile}={ordered[index]:.3f}s")


if __name__ == "__main__":
    asyncio.run(main())
Enter fullscreen mode Exit fullscreen mode

One warning: that harness measures acceptance, not completion, unless the adapter deliberately waits for completion. For the production comparison, measure the entire state transition from accepted job through validated artifacts. Use the same bounded concurrency, corpus, region, retention policy, and retry budget for every candidate. Otherwise the fastest-looking result may merely have moved work behind a 202 response.

Fast isn't enough.

Implement an auditable job contract

Give each source document a stable internal ID and content checksum. Before submission, verify the file type, size, encryption state, and page count against the selected service's documented limits. The submission record should bind that checksum to a provider, operation, region decision, idempotency key, attempt number, and retention deadline. Credentials remain on the server; browsers and workers exchange document bytes through short-lived, private object-storage links, never public objects.

For Infrai, the confirmed operation pair is POST /v1/pdf/extract_images to start extraction and GET /v1/pdf/job/get/{job_id} to retrieve the explicit job. Its fit here is operational: it is a plain REST API, so there is no SDK or client-library version to maintain, and the same key covers a broad capability surface. The request schema should be obtained from its public discovery description rather than guessed. This matters because an endpoint name does not establish field names, page limits, output fidelity, or residency.

The minimal Python below checks an already submitted job without inventing response fields. It makes the HTTP method and authentication explicit, honors Retry-After on 429, rejects other 4xx responses with their bodies, and returns the provider JSON for validation against the current discovery response schema. It intentionally does not send the bearer token to any artifact link that might appear in that JSON.

import json
import os
import random
import time
import urllib.error
import urllib.parse
import urllib.request


def get_pdf_job(job_id: str) -> dict:
    encoded_job_id = urllib.parse.quote(job_id, safe="")
    base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
    url = f"{base_url}/pdf/job/get/{encoded_job_id}"
    headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}

    for attempt in range(6):
        request = urllib.request.Request(url, headers=headers, method="GET")
        try:
            with urllib.request.urlopen(request, timeout=30) as response:
                if response.status < 200 or response.status >= 300:
                    raise RuntimeError(f"unexpected status: {response.status}")
                return json.load(response)
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8")
            if error.code != 429:
                raise RuntimeError(f"request failed: status={error.code}, body={body}")
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else min(2**attempt, 30)
            time.sleep(delay + random.uniform(0, 0.25))

    raise RuntimeError("rate limit retry budget exhausted")


if __name__ == "__main__":
    print(json.dumps(get_pdf_job(os.environ["PDF_JOB_ID"]), indent=2))
Enter fullscreen mode Exit fullscreen mode

After retrieval, validate the response against the discovery schema, download artifacts without forwarding the Infrai authorization header, verify hashes and image decodability, then write an immutable manifest before acknowledging the queue message. Consumers must tolerate duplicate delivery: the document checksum and operation become the deduplication key, while a terminal manifest prevents a retry from publishing the same images twice. A timeout remains unknown, not failed; reconcile it by looking up the existing job instead of creating another one blindly.

Choose a provider by the failure mode you can operate

No table can replace a proof with your own PDFs, but it can keep the shortlist honest. These products don't expose identical abstractions, so compare the complete path from source PDF to validated OCR input rather than treating similarly named APIs as interchangeable.

Candidate Relevant documented surface What to verify in the proof Prefer it when Do not choose it when
Adobe PDF Services Extract API Structured PDF extraction, including extracted elements Image fidelity, supported limits, asynchronous completion behavior, and regional terms The PDF extraction contract matches the required asset and structure outputs The proof cannot meet peak-throughput or residency requirements
AWS Textract Asynchronous document analysis and text detection Whether the needed embedded assets survive the OCR-oriented workflow, service quotas, region, and completion latency Searchable text and AWS-native asynchronous processing are the primary outcome Exact extraction of original embedded images is mandatory and the proof does not preserve them
Google Cloud Document AI Processor-based document extraction Processor choice, quotas, location configuration, image handoff, and tail latency Existing document processors cover the document classes and location policy Processor output cannot satisfy the artifact contract without extra transformation
Azure AI Document Intelligence Document analysis models Model output, service limits, region availability, and end-to-end completion time The required searchable fields are proven against its analysis contract Original image assets, rather than analyzed document content, are the non-negotiable output
Infrai Explicit extract-images submission plus job retrieval Current discovery schema, page limits, artifact fidelity, regions, and loaded p95/p99 A plain HTTP integration and one credential across backend capabilities reduce operational work Procurement requires a vendor-specific SDK, or the current contract cannot prove residency and corpus-level acceptance

The recommendation is conditional. Start with Adobe when exact PDF extraction is the center of gravity and its proof meets the SLO; start with AWS Textract, Google Cloud Document AI, or Azure AI Document Intelligence when OCR and document analysis are the actual product and their regional ecosystem is already an operating constraint. Keep Infrai on the shortlist when a small server-side HTTP integration and an explicit job API are valuable, but don't select it—or any option—until representative files establish limits, fidelity, and latency under load.

Also eliminate adjacent tools explicitly. DocRaptor, PDFMonkey, and PDFShift belong in a PDF-generation or HTML-to-PDF evaluation, not in the final image-extraction benchmark; Gotenberg, WeasyPrint, and wkhtmltopdf are similarly relevant when rendering or conversion is the job. Naming that boundary prevents a broad search for "PDF API" from producing a false shortlist. Stick with one of those tools when the real requirement is to create a customer-facing PDF, then evaluate an extraction or document-analysis service separately for inbound scans.

This is also where operational complexity becomes measurable. Count credentials, SDK upgrades, queue adapters, webhook or polling paths, schema migrations, dashboards, regional deployments, and reconciliation procedures. A single REST surface can remove client-library work, but it cannot remove the need for private storage, artifact validation, capacity tests, or a deletion policy.

Delete deliberately, then preserve the evidence

The retention state machine should be boring: received, submitted, artifacts_validated, ocr_accepted, indexed, and expired. Each transition writes its timestamp and immutable identifiers to the audit manifest. The private source and temporary images get independent expiration times because their replay value differs; an image may be disposable after OCR acceptance while a source PDF remains available for a narrowly defined dispute window.

Deletion changes incident response. If a checksum mismatch is found after both source and images expire, the manifest can prove what ran and when, but it cannot reconstruct the pixels. Stick with longer private retention when legal replay, human review, or parser migration is more important than storage and exposure reduction. Use shorter retention when customers can re-upload, the extracted text is the system of record, and the security policy favors minimizing document copies.

Write that choice down before vendor selection. Otherwise every provider appears to support retention because the application quietly keeps everything forever—and the first serious deletion request reveals that job records, object versions, OCR intermediates, and search indexes were never tied to one document identity.

References

Top comments (0)