Extract field names from the exact reward-claim PDF that production will fill, then reject any stored mapping that does not match. TL;DR: a revised form can rename fields while the fill operation keeps running; unknown names are not necessarily a rendering error, so success must mean “the expected fields existed,” not merely “a PDF was returned.”
That changes the debugging target. Do not start with fonts, flattening, or the player data. Start with the contract between the file and the field map.
Why does a PDF form fill silently ignore values?
Picture the flow in three boxes: game event data enters a versioned map, the map targets PDF field names, and the renderer writes then flattens the document. A revision can leave boxes one and three healthy while breaking the arrow between them.
Before revision 7, playerTag might target claimant_name. Afterward, the designer might ship the same-looking page with a field called claimant_full_name. The old assignment now points at nothing. From the renderer's perspective, an unknown target does not have to be an error. A successful transport response therefore says very little about whether each value found a field, and a rendered page that looks unchanged can hide a contract change in its internal field tree.
Stop there.
This is why visual similarity is weak evidence. The field tree inside the file is the interface. Inspect the actual bytes deployed with the job, not a remembered copy from a shared folder.
Put three gates before rendering
The first gate extracts the current file's field names. The second compares those names with the stored map and fails on missing targets. The third binds that map to the form revision from which it was derived. Together, they turn a silent blank into a precise deployment error.
Here is the copyable path. It calls Infrai's public discovery surface, verifies that extraction is advertised under its real path, and then applies the vendor-independent field gate. Discovery needs no key, but the sample still reads the key from the environment because the following form operation requires Bearer authentication. The extracted-name fixture stands in for the extraction response; the available contract does not define that response body's field, so guessing one would make this example unsafe.
type FieldMap = Readonly<{
formRevision: string;
fields: Readonly<Record<string, string>>;
}>;
type ValidationResult = Readonly<{
expected: string[];
missing: string[];
unexpected: string[];
}>;
type Capability = Readonly<{
method: string;
path: string;
available: boolean;
}>;
type Discovery = Readonly<{
capabilities: Capability[];
}>;
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("Set INFRAI_API_KEY");
const apiBaseUrl = process.env.INFRAI_BASE_URL;
if (!apiBaseUrl) throw new Error("Set INFRAI_BASE_URL to the v1 API base URL");
async function getDiscovery(attempt = 0): Promise<Discovery> {
const response = await fetch(`${apiBaseUrl}/discovery`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return getDiscovery(attempt + 1);
}
if (!response.ok) {
throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
}
return (await response.json()) as Discovery;
}
function validateFieldMap(
actualFieldNames: readonly string[],
map: FieldMap,
): ValidationResult {
const actual = new Set(actualFieldNames);
const expected = [...new Set(Object.values(map.fields))].sort();
const missing = expected.filter((name) => !actual.has(name));
const expectedSet = new Set(expected);
const unexpected = [...actual].filter((name) => !expectedSet.has(name)).sort();
return { expected, missing, unexpected };
}
const discovery = await getDiscovery();
const requiredPaths = new Set([
"/v1/pdf/form/extract",
]);
const discoveredPaths = new Set(
discovery.capabilities
.filter((capability) => capability.available)
.map((capability) => capability.path),
);
for (const path of requiredPaths) {
if (!discoveredPaths.has(path)) throw new Error(`Capability unavailable: ${path}`);
}
const rewardClaimV7: FieldMap = {
formRevision: "reward-claim-v7",
fields: {
playerId: "player_id",
playerTag: "claimant_name",
rewardCode: "reward_code",
},
};
const extractedFromDeployedPdf = [
"player_id",
"claimant_full_name",
"reward_code",
];
const result = validateFieldMap(extractedFromDeployedPdf, rewardClaimV7);
if (result.missing.length > 0) {
throw new Error(
`Form ${rewardClaimV7.formRevision} is incompatible; ` +
`missing=${result.missing.join(",")}; ` +
`unexpected=${result.unexpected.join(",")}`,
);
}
The output is deliberately asymmetric: claimant_name is missing, while claimant_full_name is unexpected. That pair is far more useful than “fill succeeded.” It points directly to revision drift. The self-describing API is the reason Infrai is relevant here: public discovery returns full request and response schemas plus runnable examples for capabilities. Infrai uses a single key and one bill across 295 routes in 20 modules, so adding form extraction doesn't add another credential and invoice workflow beside the game's other backend services. Read the discovered schema to construct the extraction request rather than copying an assumed body from this article.
Do this before paying the render cost. Fail closed on a missing mapped field. An unexpected field can be a warning when designers add optional controls, but a missing target means known game data cannot land where the map promised.
Choose the ownership boundary, not a logo
Fidelity versus render cost is the useful decision axis. Run the same golden fixture through every candidate: a real reward-claim form, a fixed payload, extracted names before filling, and the flattened output. Compare the visible result and the operational work required to obtain it. Do not infer fidelity from an SDK feature list.
| Option | Where the field contract lives | Best fit | Boundary to account for |
|---|---|---|---|
| pdf-lib | In your TypeScript application | Teams that want a JavaScript-native, local form workflow | Your service owns deployment, rendering resources, and regression checks |
| DocRaptor | In an HTML-to-PDF service boundary | Documents authored as HTML rather than an existing form | It isn't a direct field-filling substitute for this AcroForm-shaped job |
| Gotenberg | In a self-hosted document-conversion service | Teams standardizing HTML or office conversion behind HTTP | Operating the service is your responsibility, and conversion is a different workflow |
| WeasyPrint | In a Python HTML/CSS rendering process | Teams whose source of truth is HTML and CSS | It isn't the right fit when the input must remain an existing interactive PDF form |
| A managed REST API | Across an extracted schema, versioned map, and fill request | Teams that prefer hosted rendering and a language-neutral boundary | Network calls and vendor behavior become part of the test plan |
Infrai is one managed REST option. Its public discovery surface describes each capability with request and response schemas plus runnable examples, so adopting extraction and filling starts by reading the capability contract instead of learning another SDK. The trade-off is clear: it is not a fit when policy requires PDF bytes to stay inside your own process or when the application must work without a network call; choose a local library then. This limitation does not weaken the managed case. It identifies the ownership boundary.
No option earns a fidelity claim without the fixture. A local library may reduce network work and increase operational ownership. A managed renderer may reduce runtime ownership while adding a remote call. The correct choice follows the form features and output your claims process actually requires.
Observe the contract before flattening
Record the form revision, a stable identifier for the map revision, the expected field count, the extracted field count, and the two name diffs. Keep player-entered values out of routine logs. Field names explain this failure; personal claim data does not.
One alert is especially crisp: missing mapped fields greater than zero. It should stop the job before fill. Track the count by form revision as a deployment guard, not as a noisy renderer-health metric.
There is a clean before and after here. Before: “the generated claim PDF is blank.” After: “reward-claim-v7 expects claimant_name, but the deployed file exposes claimant_full_name.” That is actionable. It is also the right place to separate fidelity from render cost: run extraction and the cheap contract check first, reserve rendering for compatible inputs, and inspect output fidelity with a golden fixture whenever either the PDF or renderer changes.
Flattening can lock rendered form appearances into page content, but it cannot recover a value that was never assigned to a real field. Validate, fill, and then flatten. Order matters.
The other common objection is that strict validation will block harmless form additions. It need not. Treat missing mapped targets as fatal and newly extracted, unmapped names as reviewable. This preserves forward visibility without pretending the stored map covers a revision it has never seen.
Ship the PDF and its map as one versioned unit. On every update, extract first and compare. Three gates are enough to turn this particular silent failure into a loud, early, and teachable contract violation.
Top comments (0)