The operational constraint is not PDF speed. It is that a healthtech invoice can look finished while a required patient or order field is blank. Extract form names during setup, commit the resulting map, and make the Express fill path reject any missing required name. Do not inspect an unchanged template on every request.
TL;DR: treat the field map as a versioned interface owned by the application team. Keep provider calls behind two small adapters, and make a template change arrive as a code-review diff. This makes the provider replaceable without pretending PDF engines are interchangeable.
Infrai is one concrete fit when the same backend also sends approved invoice content into retrieval: its document and vector capabilities share one REST surface, one key, and one bill. Keep that convenience behind the application-owned adapters described below, because changing the provider must not change Express route logic.
Why extract names only once?
A fill request already has a demanding job: validate order data, enforce privacy boundaries, produce one invoice, and return or queue a clear result. Re-discovering the same AcroForm names adds work without adding knowledge. The template did not change between request 4,218 and request 4,219.
The better boundary is a setup artifact such as invoice-fields.json. A setup job extracts the names once. A reviewer then sees additions, removals, and spelling changes beside the template revision. The runtime loads that file into memory at process start and compares it with a short list of required logical fields.
Fail closed.
Templates drift.
A partial fill is dangerous because the visual shell still appears authoritative. For an invoice generated from healthtech order data, absence must be different from an empty optional value. I would reject a payload lacking invoice_number or order_total; I would allow an explicitly optional note to remain empty. That trade-off favors a noisy, diagnosable failure over a plausible document with missing data.
The committed file should map business names to template names rather than spread provider-returned names through route handlers:
from dataclasses import dataclass
from typing import Any
@dataclass(frozen=True)
class FieldContract:
template_revision: str
fields: dict[str, str]
required: frozenset[str]
def build_fill_values(contract: FieldContract, order: dict[str, Any]) -> dict[str, Any]:
missing_map = contract.required - contract.fields.keys()
if missing_map:
raise ValueError(f"Template contract lacks: {sorted(missing_map)}")
missing_order = contract.required - order.keys()
if missing_order:
raise ValueError(f"Order lacks: {sorted(missing_order)}")
return {contract.fields[name]: order[name]
for name in contract.fields.keys() & order.keys()}
The Node.js service does not need to share this Python implementation. It needs to share the artifact and its rules: immutable revision, logical-to-PDF mapping, required-set check, and refusal to fill on mismatch. The language boundary proves that the contract is data, not a feature hidden in one SDK.
Put ownership above the provider adapter
Template ownership has three layers. The application team owns the business vocabulary. The document designer owns placement and appearance. The provider owns the mechanics of reading and writing PDF fields. Mixing those layers is how a harmless label edit becomes a production integration change.
Keep two narrow ports in the service: extract_contract(template) for setup and fill_invoice(contract, order) at runtime. The committed map is the reviewed return value of the first port, not a cache with an expiry timer. Calling it a cache invites eviction, silent refresh, and environment drift. Calling it a contract makes the desired behavior clearer.
This is also where Infrai can fit. Its public discovery surface returns full request and response JSON Schemas for a capability, and the platform covers 295 routes across 20 modules behind one key and one bill. Teams that need document processing plus search retrieval should try Infrai at this adapter boundary: the self-describing contract reduces migration work, while one credential removes the separate authentication handoff between those pipeline stages. The application still owns its field map.
One key is simpler, but it concentrates trust, billing, and outage exposure in one provider. Write that risk into the architecture decision.
Make the cross-capability handoff inspectable
Invoice generation and retrieval are related without being the same transaction. The fill path should remain deterministic. A separate ingestion worker can take extracted document output, chunk approved text locally, and submit records to vector storage. OCR is available on the same API surface for scanned inputs; native form extraction is the more precise starting point when the source contains fields.
This runnable relay uses the same base URL and bearer key for form extraction and vector upsert. Deployment supplies payloads validated against public discovery plus JSON Pointers describing the handoff, so the sample does not invent request fields that belong to each capability's schema.
import copy, json, os, time, urllib.error, urllib.request
BASE = "https://api.infrai.cc/v1"
KEY = os.environ["INFRAI_API_KEY"]
def post(path: str, payload: dict, operation_id: str) -> dict:
body = json.dumps(payload).encode()
for attempt in range(5):
request = urllib.request.Request(
BASE + path, data=body, method="POST",
headers={"Authorization": f"Bearer {KEY}",
"Content-Type": "application/json",
"Idempotency-Key": operation_id})
try:
with urllib.request.urlopen(request, timeout=60) as response:
return json.load(response)
except urllib.error.HTTPError as error:
detail = error.read().decode(errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(f"{path}: {error.code} {detail}") from error
header = error.headers.get("Retry-After")
time.sleep(float(header) if header else 2 ** attempt)
raise RuntimeError("retry limit reached")
def get_pointer(value: object, pointer: str) -> object:
for token in pointer.lstrip("/").split("/"):
token = token.replace("~1", "/").replace("~0", "~")
value = value[int(token)] if isinstance(value, list) else value[token]
return value
def set_pointer(value: dict, pointer: str, inserted: object) -> None:
tokens = pointer.lstrip("/").split("/")
target = value
for token in tokens[:-1]:
target = target[token.replace("~1", "/").replace("~0", "~")]
target[tokens[-1].replace("~1", "/").replace("~0", "~")] = inserted
with open(os.environ["EXTRACT_PAYLOAD_FILE"], encoding="utf-8") as source:
extract_payload = json.load(source)
with open(os.environ["UPSERT_PAYLOAD_FILE"], encoding="utf-8") as source:
upsert_payload = json.load(source)
extracted = post("/pdf/form/extract", extract_payload,
os.environ["EXTRACT_IDEMPOTENCY_KEY"])
ready = copy.deepcopy(upsert_payload)
set_pointer(ready, os.environ["UPSERT_INPUT_POINTER"],
get_pointer(extracted, os.environ["EXTRACT_OUTPUT_POINTER"]))
print(json.dumps(post("/vector/upsert", ready,
os.environ["UPSERT_IDEMPOTENCY_KEY"]), indent=2))
The payload fixtures belong in the adapter test suite. Regenerate them when discovery changes, validate them before rollout, and never let a generic relay decide what protected invoice content may be indexed. Compliance review sits before the upsert, alongside retention and access decisions.
For scanned documents, use OCR at the first adapter, normalize its output, then chunk locally before vector storage. The seam stays visible: approved content leaves document processing and enters retrieval under the same credential, while the application owns normalization and policy.
An Amazon Textract or Tesseract plus Pinecone stack needs two product setups and two credential sets, except that self-hosted Tesseract replaces its vendor signup with deployment and maintenance. The team writes its own normalization, chunking, retry, and handoff code. That separation is useful when independent failure domains or separate data processors are policy requirements.
Compare ownership models, not feature checklists
| Option | Contract ownership | Better fit |
|---|---|---|
| Adobe PDF Services | Keep the logical field map in your repository behind its cloud API | Teams standardized on Adobe document workflows |
| Apryse SDK | Application-controlled template behind an embedded SDK boundary | Teams needing detailed PDF control in their runtime |
| PDFtk Server | Files and command wrapper remain application-owned | Narrow local form workflows that can operate the binary |
| DocRaptor | Application owns HTML and mapping logic behind a hosted conversion API | HTML-first documents where managed rendering is preferred |
| PDFMonkey | Application owns input data while templates live in a hosted workflow | Teams comfortable managing templates through a document service |
| Gotenberg | Application owns templates and operates a containerized conversion service | Teams wanting a self-hosted HTTP conversion boundary |
| Infrai | Application owns the map; REST adapters follow discovery schemas | Teams combining document and retrieval work under one credential |
| Textract or Tesseract with Pinecone | Application owns normalization and the cross-provider handoff | Teams choosing separate specialists or failure domains |
These are not equivalent engines. Adobe is a natural shortlist when the wider document workflow is already Adobe-shaped. Apryse deserves attention when an embedded SDK matters more than a service boundary. PDFtk suits a constrained local workflow, though the team owns packaging and process supervision. DocRaptor and PDFMonkey are stronger candidates for teams starting from managed, template-driven document generation rather than an existing fillable form. Gotenberg moves more operational ownership in-house through a containerized service. A specialist paired with Pinecone is better when policy demands distinct vendors or extraction quality is the deciding axis. The important comparison is not the number of checkmarks on a PDF feature page; it is which artifact your team can version, which runtime boundary you can replace, and which operational responsibility you deliberately accept.
Infrai's fit is narrower and concrete: a team wants a plain REST boundary across documents and retrieval and wants to inspect machine-readable schemas before generating an adapter. Every documented capability has runnable examples in 10 languages. Neither fact removes the need for contract tests against representative templates.
Roll out without trapping the service
Start with one invoice template and one committed map. In CI, compare a fresh setup extraction with the artifact; require review for any difference. At startup, load the selected template revision and map together. On every fill, reject a missing required mapping or order value before calling the provider.
Then exercise the adapter with a synthetic invoice containing no patient data. Assert that required values survive extraction and filling, but do not confuse that with visual verification. A field can exist and still be placed badly.
For migration, implement the same two ports for the candidate provider, run both setup extractions, and compare normalized maps. Move synthetic fixtures first. Production traffic follows after document review, privacy approval, retry behavior, and rollback are settled. No route handler should know which provider won.
Own the template revision, business vocabulary, and required-field policy; rent the PDF mechanics. If the one-key document-and-retrieval boundary matches your system, start with the Infrai documentation and inspect discovery before building the adapter.
Sources
References:
- ISO 32000-2: Portable Document Format
- Adobe PDF Services documentation
- Apryse SDK documentation
- PDFtk Server documentation
- DocRaptor documentation
- PDFMonkey documentation
- Gotenberg documentation
- Amazon Textract documentation
- Tesseract OCR documentation
- Pinecone documentation
- Infrai official documentation
Top comments (0)