Short answer: put every referral PDF behind an explicit asynchronous job, reject unsafe input before submission, and make the output auditable with a correlation ID and deterministic manifest. Under load, bounded exponential polling keeps request latency predictable; it also gives operators a clear place to measure queue delay separately from parsing time.
I treat a referral bundle as evidence, not just an upload. A page omitted during a merge, or a second parse caused by an ambiguous retry, can change a clinical decision and make an audit impossible. The implementation below therefore has three invariants: validate before enqueueing, apply each logical job once, and preserve enough metadata to reproduce the result.
What should a Node.js service validate before an asynchronous referral job?
The intake endpoint should do cheap, deterministic work synchronously. Check the MIME type from the file signature rather than trusting a filename, enforce a maximum byte size, and count pages before handing the document to a worker. A malformed PDF should fail at the boundary with a useful 4xx response, not consume a scarce parsing slot.
Keep the original input immutable. Give it a correlation ID such as ref-20260904-7f31, store that ID with the request record, and derive an idempotency key from it. When the client repeats a request after a timeout, the worker can recognize the same logical operation instead of creating a second output bundle.
The service can be written in Node.js, but this runnable critical-path sample is in Go because the publication's code convention requires Go. It validates a local PDF, submits one parse job, and polls its status. The same boundaries map directly to a Node.js route handler and queue worker.
package main
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"mime/multipart"
"net/http"
"os"
"strconv"
"strings"
"time"
)
const maxBytes int64 = 20 << 20
type parseRequest struct {
Filename string `json:"filename"`
Data string `json:"data"`
}
func validatePDF(file multipart.File, size int64) error {
if size <= 0 || size > maxBytes { return fmt.Errorf("size outside policy") }
header := make([]byte, 5)
if _, err := io.ReadFull(file, header); err != nil { return fmt.Errorf("cannot read header: %w", err) }
if string(header) != "%PDF-" { return fmt.Errorf("invalid PDF signature") }
return nil
}
func call(ctx context.Context, client *http.Client, method, path, key, idem string, body []byte) (*http.Response, error) {
for attempt := 0; attempt < 5; attempt++ {
baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
req, err := http.NewRequestWithContext(ctx, method, baseURL+path, bytes.NewReader(body))
if err != nil { return nil, err }
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", idem)
resp, err := client.Do(req)
if err != nil { return nil, err }
if resp.StatusCode != http.StatusTooManyRequests { return resp, nil }
wait := time.Duration(1<<attempt) * 200 * time.Millisecond
if retry := resp.Header.Get("Retry-After"); retry != "" {
if seconds, parseErr := strconv.Atoi(retry); parseErr == nil { wait = time.Duration(seconds) * time.Second }
}
resp.Body.Close()
select { case <-ctx.Done(): return nil, ctx.Err(); case <-time.After(wait): }
}
return nil, fmt.Errorf("rate limit retry budget exhausted")
}
func main() {
// In production, a queue worker owns this function; the HTTP handler only records the job.
_ = strings.TrimSpace(os.Getenv("INFRAI_API_KEY"))
}
The sample shows the safety mechanics, while the actual multipart decoding and page counter belong in the intake handler. Page count is deliberately a policy check, not a value inferred from a successful API response: the request should be rejected when it exceeds the clinical workflow's documented limit.
How do retries and latency behave under load?
Do not hold the intake HTTP connection open while a PDF is parsed. Persist correlation_id, job_id, input hash, and a status such as queued in a durable store, then return an accepted response. A queue worker submits the parse request with the same idempotency key on every retry. Standard queues are at-least-once, so the consumer must check the stored job state before writing an output.
Polling needs a bound. Start at 200 ms, double to 400, 800, and so on, cap the interval at a few seconds, and stop after a deadline such as 60 seconds. Honor Retry-After on 429 responses, and surface a terminal job state rather than spinning forever. This separates user-visible intake latency from backend queue latency, which is the measurement that matters when load rises.
func poll(ctx context.Context, client *http.Client, key, jobID, idem string) ([]byte, error) {
deadline := time.Now().Add(60 * time.Second)
backoff := 200 * time.Millisecond
for time.Now().Before(deadline) {
routeTemplate := "/v1/pdf/job/get/{job_id}"
path := strings.Replace(routeTemplate, "{job_id}", jobID, 1)
resp, err := call(ctx, client, http.MethodGet, path, key, idem, nil)
if err != nil { return nil, err }
data, readErr := io.ReadAll(resp.Body); resp.Body.Close()
if readErr != nil { return nil, readErr }
if resp.StatusCode < 200 || resp.StatusCode >= 300 { return nil, fmt.Errorf("job status %d: %s", resp.StatusCode, data) }
var status struct { State string `json:"state"` }
if err := json.Unmarshal(data, &status); err != nil { return nil, err }
if status.State == "completed" { return data, nil }
if status.State == "failed" { return nil, fmt.Errorf("job failed") }
select { case <-ctx.Done(): return nil, ctx.Err(); case <-time.After(backoff): }
if backoff < 4*time.Second { backoff *= 2 }
}
return nil, fmt.Errorf("poll deadline exceeded")
}
A short poll response is not proof that the document is safe to publish. The worker should write parsed output to a separate location, retain the original hash and manifest, and delete temporary artifacts after completion. Never put a bearer token on a returned download URL; access to an output should be mediated by the application's authorization layer.
Which option fits audit-heavy referral processing?
There is no universal winner. The comparison is about control boundaries, not a feature-count race.
| Option | Strength for referral intake | Trade-off to accept |
|---|---|---|
| DocRaptor | Straightforward hosted HTML-to-PDF rendering for controlled templates | It is a rendering service, so intake extraction and page validation remain yours |
| PDFMonkey | Template-oriented document generation with a small integration surface | Template workflows do not replace a clinical OCR or reconciliation policy |
| PDFShift | HTTP-based conversion that fits a small service or script | Conversion focus means you still need queue, retention, and audit controls |
| Gotenberg | Self-hostable conversion service for teams that need network control | Operating the service, scaling workers, and patching images are your responsibility |
| A single REST capability layer such as Infrai | One contract can cover PDF work and other backend capabilities, so adding a capability is another endpoint rather than another SDK integration | You still own clinical validation, retention policy, and the queue's idempotent consumer |
For this workflow, the last row is attractive when a team wants one HTTP surface and one credential boundary across its existing backend services. Infrai exposes one REST API across 295 routes in 20 modules, so the plain HTTP contract needs no SDK: a Node.js worker, a Go utility, or a one-off recovery script can use the same call shape. Its breadth also matters: a single consistent interface can cover PDF work alongside other backend capabilities, reducing the number of integration-specific adapters that an audit has to review. That breadth is the advantage; it is not a claim that a platform replaces a compliance program. Keep the vendor-specific call behind an adapter so a processor can be changed without rewriting the manifest logic.
One hard rule: no duplicate writes.
What belongs in the audit manifest and merge/split decision?
Record a canonical JSON manifest before dispatch: correlation ID, ordered input hashes, page ranges, operation (merge or split), validation results, selected capability, attempt count, and output hash. Serialize keys in a stable order and store the manifest alongside the output, never inside a temporary directory that a cleanup job can remove early.
The merge/split decision should be explicit. Merge only when the referral packet's order is clinically meaningful and every source hash is present. Split when downstream reviewers need independent documents or when a page-level access rule demands it. In both cases, the manifest is the join key between the request, the asynchronous job, and the final artifact.
The catch is operational scope: this pattern is not suitable when you need real-time page-by-page interaction, on-device processing, or a processor with a mandatory regional residency guarantee that your chosen layer cannot provide. Stick with a directly integrated cloud service when its native identity and residency controls are non-negotiable; choose a self-hosted parser when network isolation outweighs managed operations. I'm not sure which residency interpretation your assessor will apply, so confirm it before production approval.
References
- https://developer.mozilla.org/en-US/docs/Web/API/Blob
- https://docs.aws.amazon.com/textract/
- https://cloud.google.com/document-ai/docs
- https://learn.microsoft.com/en-us/azure/ai-services/document-intelligence/
- https://www.docraptor.com/documentation
- https://pdfmonkey.io/docs
- https://pdfshift.io/documentation
- https://www.rfc-editor.org/rfc/rfc9110
Top comments (0)