DEV Community

ElowenVeil9067
ElowenVeil9067

Posted on

Medical Referral Intake PDFs: Hosted API or Local Libraries for Production Latency

Short answer: choose the boundary that keeps failure visible and recoverable. For medical referral intake, a local PDF worker is usually the safer synchronous core; a hosted PDF API becomes preferable when burst isolation or a missing transformation outweighs network and governance risk. Judge it with tail latency, queue age, and replay behavior, not a single happy-path benchmark.

I run a one-person SaaS, so reliability has a revenue-per-hour meaning. Every hour spent explaining a lost referral packet is an hour I did not ship. The same principle applies to a logistics pipeline that watermarks documents before external sharing: the watermark is easy; proving that every batch was processed once is the product.

The reliability contract comes before the renderer

Write down the contract in operational terms. An intake request accepts a bounded file, stores an immutable original, and returns a job ID. A worker produces one versioned output or one classified failure. Delivery is retried independently. This contract survives a library swap and makes a hosted service just another processor behind the boundary.

Contract question Local worker Hosted PDF API
Where do bytes execute? Inside your network boundary In a provider boundary you must review
What causes latency variance? CPU, memory, native dependencies Network, remote scheduling, quota
How do you replay? Pin the image and library Preserve request metadata and provider version
What happens during a network partition? Existing jobs can continue New transforms wait or fail until connectivity returns

My recommendation is to make the queue and artifact store the source of truth, then plug in the processor that meets the contract. That decision is about recoverability first. Throughput comes next.

What should hosted PDF APIs and local libraries guarantee for referral intake under load?

Split the request into four clocks: upload, queue, transform, and delivery. A hosted call adds DNS, TLS, transmission, remote scheduling, and download. Local code removes that hop but moves CPU, memory, patching, and native dependency work into your fleet. Record each clock separately; otherwise a slow object-store read gets blamed on “the PDF engine.” In one batch run, the transform stayed under 400 ms while queue wait passed 18 seconds because a fetch pool was saturated. The graph looked like a renderer regression until I separated admission, fetch, transform, and delivery spans. That distinction changed the fix: two more fetch slots, a smaller prefetch window, and no PDF code change. It also made the hosted comparison fair, because the remote request was measured against the same queue and storage clocks rather than against an artificially empty local process.

For the logistics watermark job, I use a de-identified corpus of packets with different page counts and raster densities. The worker reserves a bounded slot, fetches the original, applies the watermark, writes a deterministic output key, and emits page count, byte count, and duration. A 10-page text PDF tells you almost nothing about a 500-page scanned batch.

Keep concurrency explicit. Start with one active document per CPU core, then tune against p95 and p99 latency, memory per document, and oldest-job age. A queue that grows without a limit turns a burst into a fleet-wide timeout. Backpressure is a feature.

For referral intake, return a job identifier rather than holding an HTTP request open for a large scan. Polling or an authenticated, replay-safe callback can report completion. The callback must be idempotent because a client can retry after a timeout even when the transform finished.

A processor boundary makes migration boring

The application should depend on a narrow interface. The implementation can be a maintained local library today and a hosted adapter tomorrow; tests for queue semantics do not need either renderer.

type DocumentJob = {
  id: string;
  inputKey: string;
  outputKey: string;
  mark: string;
};

interface DocumentProcessor {
  transform(job: DocumentJob): Promise<{ pages: number; bytes: number }>;
}

async function execute(job: DocumentJob, processor: DocumentProcessor): Promise<void> {
  const result = await processor.transform(job);
  await metrics.observe("document_pages", result.pages);
  await metrics.observe("document_bytes", result.bytes);
  await jobs.markComplete(job.id);
}
Enter fullscreen mode Exit fullscreen mode

The important detail is deterministic identity. Store an input hash, policy revision, processor version, and output key with the job. If a coordinator retries after a 504, the worker can discover that the output already exists instead of creating a second watermark or duplicate notification. I once treated delivery as part of transformation; that made a harmless client retry consume another rendering slot. Separating the records fixed the accounting.

Do not log referral contents, watermark text containing patient identifiers, or raw provider responses. Log the job ID, sizes, page count, duration, and a redacted error class. Retention, regional processing, and access review belong in this contract, not in a launch-week checklist.

Batch throughput is a queue problem, not a headline number

Memory often fails before CPU. A renderer that keeps page bitmaps can turn a 20 MB input into hundreds of megabytes. Enforce page and byte limits at ingestion, isolate fetch, transform, and upload pools, and terminate a worker that crosses its memory budget. A restart is cheaper than taking the intake fleet down.

Measure queue age and oldest-job time beside p95 and p99 transform latency. For a hosted processor, add quota consumption, request duration, and 429 counts. For local workers, add file-descriptor use and native-process restarts. Your dashboard should explain whether the bottleneck is admission, execution, or delivery.

Measure the tail.

I'm not sure one benchmark corpus can predict every hospital's scans; your mileage may vary. Keep a small, de-identified sample of real packet shapes and replay it after each library, container, or policy upgrade. The result is less glamorous than a vendor bake-off, but it catches the regression that matters to a coordinator waiting on a referral.

When is each boundary the wrong fit?

Local processing is a poor fit when your team cannot patch native dependencies, reserve memory for worst-case scans, or implement a required operation. It is also a weak choice for highly spiky traffic if idle capacity and on-call work exceed the value of keeping bytes in-house.

A hosted API is a poor fit when protected data cannot leave your control boundary, when the network path misses your tail-latency budget, or when retention and regional guarantees do not match policy. The catch is that the hosted boundary also inherits quota and connectivity failure modes that a local worker can avoid. It is not suitable when an offline clinic must keep processing during a network partition. Offline clinics and long-term byte-for-byte reproducibility also favor local execution.

Choose the hosted boundary when isolation and burst handling are more valuable than the extra hop, and document its quota, timeout, retention, and replay semantics before launch. Choose local code when predictable response time and offline operation dominate. There is no universal winner; the defensible choice is the one your measurements and compliance review can explain.

I ship weekly. My first release is one queue, one processor, one corpus, and explicit latency budgets. Outsource the undifferentiated rendering work only after the operational contract is clear. That keeps a document decision from becoming an accidental platform rewrite.

References

Top comments (0)