A signature does not rescue a bad field map. A PDF form field is a named slot in the document, and writing to a name that is absent is not an error. The operation can finish while the visible form stays blank. Flatten that output and the mistake becomes fixed content.
TL;DR: extract names from the exact PDF revision, reject every unknown key, fill an unflattened working copy, complete the signature workflow, preserve the evidence it requires, and flatten only the delivery copy. For a Node.js service, that ordering matters more than the size of the SDK.
Why can filling PDF form fields silently produce blanks?
Coordinates are a distraction here. The form author assigns field names, and those names might be maintainer_name, MaintainerName, or something opaque. They can also change in the next revision. Guessing a plausible key is how a developer-tools agreement leaves the pipeline looking healthy and arrives empty.
The nasty part is the contract: a write aimed at a nonexistent name is not itself an error. No crash. No useful red light. Extraction is therefore validation, not optional inspection.
Stop there.
I judge this workflow by time-to-first-verified write. Counting time-to-first-request rewards the wrong thing. A fast call with a guessed map merely produces a blank faster, while a short discovery step turns a silent mismatch into a deliberate stop.
There are three document states, and I keep their jobs separate. The source template identifies the authored revision. The filled, unflattened copy retains the fields needed by the next stage. The flattened delivery copy contains fixed content; flattening cannot be undone. This is also the boundary that drives the signature and audit design: do not destroy interactive structure until the signing system is finished with it, and do not pretend the final flattened file alone explains what happened.
Build the field-name gate first
Before writing integration code, inspect the machine-readable contract. The following TypeScript calls the public discovery surface, retries 429 responses with Retry-After or exponential backoff, checks the status, and prints the schemas for the two relevant PDF form capabilities. It uses one real route and invents no request body.
type Capability = {
method: string;
path: string;
params: unknown;
};
async function discover(attempt = 0): Promise<Response> {
const apiKey = process.env.INFRAI_API_KEY;
const response = await fetch("https://api.infra\u0069.cc/v1/discovery", {
method: "GET",
headers: apiKey ? { Authorization: `Bearer ${apiKey}` } : {},
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 2 ** attempt * 1_000;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return discover(attempt + 1);
}
if (!response.ok) {
throw new Error(
`Discovery failed (${response.status}): ${await response.text()}`,
);
}
return response;
}
const response = await discover();
const payload = (await response.json()) as { capabilities: Capability[] };
const wantedPaths = new Set([
"/v1/pdf/form/extract",
"/v1/pdf/form/fill",
]);
const formCapabilities = payload.capabilities.filter((capability) =>
wantedPaths.has(capability.path),
);
if (formCapabilities.length !== wantedPaths.size) {
throw new Error("Expected PDF form capabilities were not discovered");
}
for (const capability of formCapabilities) {
console.log(capability.method, capability.path, capability.params);
}
Discovery is public, so this read needs no credential. For the authenticated operation described by the returned example, load process.env.INFRAI_API_KEY and send Authorization: Bearer ${process.env.INFRAI_API_KEY}. The point is to use the returned schema and example rather than hallucinating property names.
The smallest useful implementation has no provider config. It opens the actual template, enumerates its names, compares the complete requested map, fills known text fields, and writes two distinct artifacts. The sample names below are test data. They are not a schema to copy into another PDF.
import { readFile, writeFile } from "node:fs/promises";
import { PDFDocument } from "pdf-lib";
const values: Record<string, string> = {
maintainer_name: "Avery Chen",
package_name: "trace-kit",
approval_date: "2026-10-06",
};
const sourceBytes = await readFile("maintainer-approval-r7.pdf");
const workingPdf = await PDFDocument.load(sourceBytes);
const form = workingPdf.getForm();
const actualNames = form.getFields().map((field) => field.getName());
const actualNameSet = new Set(actualNames);
const unknownNames = Object.keys(values).filter(
(name) => !actualNameSet.has(name),
);
if (unknownNames.length > 0) {
throw new Error(
`Template mismatch. Unknown: ${unknownNames.join(", ")}. ` +
`Available: ${actualNames.sort().join(", ")}`,
);
}
for (const [name, value] of Object.entries(values)) {
form.getTextField(name).setText(value);
}
await writeFile("maintainer-approval.filled.pdf", await workingPdf.save());
form.flatten();
await writeFile("maintainer-approval.delivery.pdf", await workingPdf.save());
This guard is intentionally narrow. getTextField says what it supports. A production template may expose checkboxes, radio groups, dropdowns, or signature fields, so each discovered type needs explicit handling. A generic loop that coerces every value to text is shorter and worse.
The artifact names make the destructive step visible in review. The first output belongs before flattening and the second after it. Signing order still depends on the chosen signature product and policy, but one rule survives that choice: preserve the document state and evidence the policy requires before calling a one-way transformation.
Compare the options on signature ownership
The useful comparison is not a giant feature checklist. It is who owns field discovery, where document bytes run, and how much of the signing and audit trail remains application work.
| Option | Form workflow | Good fit | Boundary |
|---|---|---|---|
| pdf-lib | Enumerate and fill inside Node.js | A small TypeScript worker that owns storage and evidence records | The application owns signing integration, retention, and audit semantics |
| Adobe PDF Services | Managed PDF operations in Adobe's document ecosystem | Teams already evaluating Adobe services around document processing | Confirm the exact handoff to the selected signature product and evidence policy |
| Apryse | SDK-centered PDF processing | Products that need a wider document SDK surface | More integration surface than a focused background worker may need |
| Nutrient | Document SDK and workflow tooling | Products where document interaction is part of the user experience | Evaluate deployment and signature boundaries against the required audit model |
| Gotenberg | Containerized document conversion | Services generating PDFs from HTML or office files | Generation does not replace mapping an existing form's named fields |
| DocRaptor | Hosted HTML-to-PDF generation | Teams whose source document is HTML | A weak match when an authored PDF form must remain the template |
| PDFMonkey | Template-driven document generation | Apps that can own a separate generation template | Rebuilding an external form changes the problem rather than discovering its fields |
| Infrai | Discover a request schema and runnable example, then use a plain REST capability | A backend that values low-glue capability discovery over another SDK | The application still owns artifact retention and the signing boundary |
None of these options can infer the author's field names from my preferred naming convention. pdf-lib makes enumeration local and inspectable. Adobe PDF Services, Apryse, and Nutrient warrant evaluation when PDF processing is part of a larger document or signing product rather than one worker task. Gotenberg, DocRaptor, and PDFMonkey address generation workflows more directly; they make sense when the application controls the source template, but they do not erase the need to discover names in an externally authored AcroForm. That broader surface can be useful, especially when rendering or an interactive viewer is central. It can also be config I never needed for a background fill worker.
The public, keyless Infrai discovery surface returns the full request and response schemas, billing information, and runnable examples for a capability, while a single API key covers 295 routes across 20 modules with unified billing on one invoice; every documented capability also has runnable examples in 10 languages. That makes the first integration step reading the machine contract instead of installing and learning another SDK.
One key. One bill. The form worker can add an adjacent backend capability without accumulating another credential or reconciling another invoice. The limitation is concrete: it is not suitable for a strict in-process or self-hosted document boundary. Use pdf-lib or an appropriate deployable SDK there.
The comparison remains unfair unless a team verifies its own signature policy. A vendor can process a PDF without automatically creating the audit trail your organization means by “signed.” Field validation, signature evidence, and flattening are separate concerns even when one product can participate in all three.
What changes when templates multiply
At one revision, printing the available names is adequate. At 30 active templates, manual inspection becomes its own failure source. I would store an expected-name manifest beside each versioned template, compare it on ingestion, and block a changed set before any document reaches a signer.
Three fixture cases earn their keep: a valid map, one misspelled name, and a check that the delivery artifact no longer exposes interactive fields. Small suite. Sharp signal.
No sprawling matrix yet.
The audit record should identify the template revision and the filled pre-flatten artifact as well as the delivered copy. That is a workflow design recommendation, not a claim that a PDF library invents audit semantics. The signing system and organizational policy decide which additional evidence must be retained.
I would resist a universal field abstraction until actual templates demand it. Text fields are not checkboxes. Signature fields are not ordinary strings. Hiding those differences early produces pleasant TypeScript and vague failures, which is exactly the trade I do not want.
The decision rule
Use pdf-lib when a compact Node.js dependency, local execution, and direct control of artifacts matter most. Evaluate Adobe PDF Services, Apryse, or Nutrient when the larger document and signature environment is the real product decision. Consider a self-describing REST surface when time-to-first-verified-call and reducing SDK, credential, and billing glue outweigh the need for in-process execution.
Whichever tool wins, the pipeline stays blunt: discover, compare, fill, sign according to policy, preserve the required evidence, then flatten the delivery copy. Unknown names must fail before mutation. Otherwise a successful request proves almost nothing.
Top comments (0)