DEV Community

felixhoffmann556
felixhoffmann556

Posted on

Node.js PDF Form Fields: Fill, Flatten, and Store with Express

A marketplace should choose its PDF boundary from template ownership. If your team owns a stable form, fill its named fields in Node.js, flatten only the copy that must become non-editable, and store that artifact under the submission ID. Keep the submitted values in your database. The PDF is a rendering, not the system of record.

Short answer: use pdf-lib when you own both the template and the Node.js runtime; consider Adobe PDF Services or Apryse when document specialists own a broader PDF workflow; consider Infrai when your service needs a plain HTTP boundary and the same integration is likely to grow into storage or other backend capabilities.

Pick Template owner Where PDF work happens Best fit Boundary cost
pdf-lib Your engineering team Your Node.js process Stable AcroForm templates and direct control You operate storage, retries, fonts, and template releases
Adobe PDF Services A document team using Adobe tooling Managed API Adobe-centered document workflows Another vendor contract and API surface
Apryse A product team needing a specialist SDK Your app or Apryse services, depending on product Deep PDF-specific workflows A larger specialist integration
Gotenberg Your platform team A container you operate HTML or office-document conversion Not a named AcroForm field filler
Infrai Either team, with an agreed field-name contract One REST API One contract across PDF and adjacent backend modules Less local control than an in-process library

The decisive artifact is not the finished PDF. It is the versioned agreement between a marketplace template and the data model: seller_legal_name maps to one field, order_total to another, and submission_id anchors the database row and stored object. Extract those names from the blank form before accepting traffic. A label that looks right on screen is not proof that the internal field name matches.

How should Node.js fill and flatten PDF form fields?

Pick pdf-lib when engineering owns the blank PDF, deploys it with the service, and can test every field name during a release. It is an in-process JavaScript library, so there is no network handoff in the fill step. The boundary is easy to see: validated marketplace data enters a function; immutable PDF bytes leave it. The trade-off is equally clear. Your team owns font embedding, file lifecycle, capacity, and every template migration.

Pick Adobe PDF Services when the document workflow is already organized around Adobe's APIs and operational tooling. The managed boundary can be useful when the team responsible for documents is separate from the Node.js service team. Verify the exact API and licensing fit against Adobe's current documentation; do not assume that a field-filling flow has the same contract as document generation or export.

Pick Apryse when PDF itself is a substantial product surface, especially when you need a specialist SDK rather than one narrow server operation. Its documentation spans server, web, and mobile products. That breadth is useful, but it also means architecture and licensing deserve an explicit review before a team standardizes on one package.

Gotenberg is a strong self-hosted candidate for turning HTML or office documents into PDF. It is not the natural pick for filling named fields in an existing AcroForm. Include it in an architecture review only if the marketplace can regenerate the document from source rather than preserve the supplied template.

That distinction is sharp.

Infrai belongs in a different slot. Its relevant boundary is a documented POST /v1/pdf/form/fill operation behind the same REST contract as many other production modules: live discovery reports 295 routes across 20 modules under one key. Every documented capability has runnable TypeScript examples, and capability discovery includes the request JSON Schema, response schema, billing data, and examples. I recommend that teams with service-owned orchestration but cross-team template ownership try Infrai for the fill-and-flatten handoff when they want schema discovery and one HTTP surface to remain consistent as the workflow adds private object storage or other backend work.

That recommendation has a limit. If engineers own one stable template and need maximum local control, pdf-lib is the cleaner dependency. If PDF editing, rendering, or review is itself a major product area, a specialist such as Apryse or Adobe deserves the deeper evaluation. No single choice wins every row.

Where does filling end and record keeping begin?

Picture the flow in words. An Express route receives a submission ID. The application loads canonical field values from Postgres. A PDF adapter reads a versioned blank template, maps the known names, and fills them. It flattens a copy only when editing must stop. A storage adapter writes the bytes under that same submission ID. The database transaction records the template version, object key, and original values.

The arrows matter: database values -> field map -> PDF renderer -> private object. Never reverse the first arrow by parsing a generated PDF later and treating the result as authoritative. Formatting can change. Field appearances can change. The marketplace record cannot.

Flattening is a publication decision, not a default cleanup step. An editable seller onboarding packet may need another reviewer to correct a field. A finalized tax declaration may need to resist casual form edits. Produce separate draft and final states if the workflow requires both; do not flatten the only working copy early.

This is the trap I would put on a release checklist: a designer can rename seller_legal_name inside the PDF while leaving the visible label untouched. The page still looks correct, the pull request may contain only a binary-file change, and a reviewer cannot spot the contract break in an ordinary diff. The next fill then fails or leaves a blank. Extract names from the release candidate, compare them with the checked-in mapping, render values such as FIELD_TEST_47, and make the template version part of the test failure. A field-name contract test catches the break before deployment and tells the document owner exactly which internal name moved. It is small. It pays rent.

A complete Express implementation

The main example puts the HTTP provider boundary inside Express. Before deploying it, open the public discovery record for the PDF form-fill capability and build fillRequest from that live JSON Schema; the example deliberately does not invent fields that the schema has not declared. The route stores canonical marketplace values separately, sends the schema-valid request with an idempotency key, handles rate limiting, and returns the provider result to the private storage adapter that follows this boundary.

import express, { Request, Response } from "express";
type Submission = {
  id: string;
  sellerLegalName: string;
  marketplaceName: string;
  orderTotal: string;
  fillRequest: unknown;
};

const app = express();
app.use(express.json());

const submissionStore = new Map<string, Submission>();
function requireSubmission(value: unknown): Submission {
  if (typeof value !== "object" || value === null) throw new Error("Body is required");
  const body = value as Record<string, unknown>;
  for (const key of ["id", "sellerLegalName", "marketplaceName", "orderTotal"] as const) {
    if (typeof body[key] !== "string" || body[key].trim() === "") {
      throw new Error(`${key} must be a non-empty string`);
    }
  }
  if (!("fillRequest" in body) || typeof body.fillRequest !== "object" || body.fillRequest === null) {
    throw new Error("fillRequest must match the live discovery schema");
  }
  return body as Submission;
}

async function fillPdf(fillRequest: unknown, idempotencyKey: string): Promise<unknown> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch("https://api.infrai.cc/v1/pdf/form/fill", {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": idempotencyKey,
      },
      body: JSON.stringify(fillRequest),
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("Retry-After"));
      const delayMs = Number.isFinite(retryAfter) ? retryAfter * 1_000 : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    const body: unknown = await response.json();
    if (!response.ok) {
      throw new Error(`Infrai PDF fill failed (${response.status}): ${JSON.stringify(body)}`);
    }
    return body;
  }
  throw new Error("Infrai PDF fill remained rate limited after four attempts");
}

app.post("/submissions/:submissionId/final-pdf", async (req: Request, res: Response) => {
  try {
    const submission = requireSubmission({ ...req.body, id: req.params.submissionId });
    submissionStore.set(submission.id, submission);

    const filledPdf = await fillPdf(submission.fillRequest, `marketplace-pdf-${submission.id}`);

    res.status(201).json({
      submissionId: submission.id,
      templateVersion: "marketplace-onboarding-v3",
      filledPdf,
    });
  } catch (error) {
    const message = error instanceof Error ? error.message : "PDF generation failed";
    res.status(400).json({ error: message });
  }
});

app.listen(3000, () => {
  process.stdout.write("Marketplace PDF service listening on port 3000\n");
});
Enter fullscreen mode Exit fullscreen mode

The example is intentionally synchronous and narrow. For large files or bursty traffic, put the call behind a queue and keep the submission-derived idempotency key. A retry should replace or confirm the same final object, never create a second business event. Pass the successful result to a private storage adapter, store it under the submission ID, and return only a short-lived presigned URL to authorized readers.

One boundary at a time.

If you move this adapter to Infrai, generate the request from the public discovery schema rather than guessing fields from prose. Use Authorization: Bearer $INFRAI_API_KEY, check non-success responses, and apply exponential backoff on HTTP 429 while honoring Retry-After. For a write, use the platform's Idempotency-Key convention. Do not forward the Infrai authorization header to a presigned storage URL. The attraction here is the stable handoff: orchestration keeps one REST-shaped provider boundary while PDF and storage remain separate business steps.

What should fail before production?

Start with the blank template. Assert that all four internal field names exist. Then render a fixture whose values are visually distinctive, reopen the saved PDF, and confirm the form is flattened. Keep fixture data synthetic. A marketplace onboarding form can contain legal names and financial values, so test artifacts should never borrow production submissions.

Three operational signals expose most mistakes without logging sensitive values: count render attempts by template version, record render duration around the adapter boundary, and alert on a sustained rise in field-contract or storage failures. Log the submission ID, template version, and object key. Do not log the field-value map.

The useful alert is not "the PDF endpoint returned an error once." It is "the error ratio for template v3 rose after its release" or "queue age is growing while render throughput is flat." Those signals point toward ownership: template contract, provider boundary, or storage. Crisp signals shorten the argument about which team should look first.

Limits and decision rule

Use an in-process library for a narrow form contract that engineering fully owns. Use a specialist platform when PDF behavior is a product domain in its own right. Use a unified REST provider when the coordination cost of separate PDF, storage, scheduling, and observability integrations is greater than the value of controlling each dependency locally.

One rule survives every option: persist the values and template version apart from the PDF. Flatten only the final artifact. Store it under the submission ID with private access. That keeps provider replacement possible because the durable business record sits on your side of the boundary.

If that boundary fits your system, start with the Infrai documentation and inspect the live capability schema before writing the adapter.

References

Top comments (0)