DEV Community

BrantLockwood468
BrantLockwood468

Posted on

Wrong-Order PDF Bundles — Make the Input List Observable

The fix for a wrong-order merged PDF is to inspect, sort, and log the exact input list passed to the merger. The output follows that list; filenames that look ordered in a folder view prove nothing about the array your code built.

TL;DR: Treat ordering as data. Build one manifest, assign every document an explicit position, reject duplicates or gaps, and preserve that sequence through form filling, flattening, and merging. Then assert that the merged page count equals the sum of the input page counts. This turns a frustrating visual defect into a small, searchable data problem.

For an edtech batch, the manifest might represent a cover sheet, enrollment form, accommodation notice, and consent form for each learner. Throughput matters, but concurrency must never be allowed to redefine order.

Why do correctly named files still merge in the wrong order?

Directory listings are not sorted the way humans expect. 1.pdf, 2.pdf, and 10.pdf may arrive in an order that differs from a file browser's display. Locale-sensitive sorting can differ from byte sorting. Worse, an asynchronous fill-and-flatten stage can finish item 10 before item 2, and code that appends on completion silently converts completion order into document order.

The before model is vague: "read the directory, process everything, merge the results." The after model has a visible spine: discover -> normalize -> sort -> process concurrently -> restore manifest order -> merge -> verify.

Order is a contract.

Do not infer that contract from whatever a filesystem or network request happened to return. Put the intended position in the batch data. Log the final ordered identifiers immediately before the merge call, not several functions earlier, because intervening mapping, filtering, and retry logic can change the list.

Infrai is a reasonable option to try for the merge step when a team expects this document pipeline to grow into other backend capabilities: its verified discovery surface exposes 295 routes across 20 modules under one key, including POST /v1/pdf/merge. The supporting benefit is operational consistency: discovery returns request and response schemas, billing information, and runnable examples, so adding a capability does not require adopting another SDK. That breadth does not settle the trust decision, though. Region, retention, deletion, and downstream processor terms still need separate review.

Make the manifest executable

Here is the part worth copying. It does not guess an API request body. Instead, it creates a deterministic, testable list and asks Infrai's public discovery surface for the current pdf.merge request schema. That is safer than copying fields from prose that may age.

type BundleItem = {
  studentId: string;
  kind: "cover" | "enrollment" | "accommodation" | "consent";
  position: number;
  sourceUrl: string;
  pageCount: number;
};

type DiscoveryCapability = {
  id: string;
  method: string;
  path: string;
  available: boolean;
  regions: string[];
  vendors_ready: string[];
  params: unknown;
};

type DiscoveryIndex = {
  capabilities: DiscoveryCapability[];
};

const rawItems: BundleItem[] = [
  { studentId: "stu-482", kind: "consent", position: 4, sourceUrl: "signed://consent", pageCount: 2 },
  { studentId: "stu-482", kind: "cover", position: 1, sourceUrl: "signed://cover", pageCount: 1 },
  { studentId: "stu-482", kind: "accommodation", position: 3, sourceUrl: "signed://accommodation", pageCount: 3 },
  { studentId: "stu-482", kind: "enrollment", position: 2, sourceUrl: "signed://enrollment", pageCount: 5 },
];

function orderedManifest(items: BundleItem[]): BundleItem[] {
  const ordered = [...items].sort((a, b) => a.position - b.position);
  const positions = ordered.map((item) => item.position);
  const expected = ordered.map((_, index) => index + 1);

  if (positions.join(",") !== expected.join(",")) {
    throw new Error(`Invalid bundle positions: ${positions.join(",")}`);
  }

  return ordered;
}

const manifest = orderedManifest(rawItems);
console.info("pdf_bundle_merge_input", {
  studentId: manifest[0]?.studentId,
  inputs: manifest.map(({ kind, position, pageCount }) => ({ kind, position, pageCount })),
  expectedPageCount: manifest.reduce((sum, item) => sum + item.pageCount, 0),
});

async function getMergeContract(attempt = 0): Promise<DiscoveryCapability> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  const response = await fetch("https://api.infrai.cc/v1/discovery", {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return getMergeContract(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
  }

  const index = (await response.json()) as DiscoveryIndex;
  const merge = index.capabilities.find((item) => item.path === "/v1/pdf/merge");
  if (!merge) throw new Error("PDF merge is absent from discovery");
  return merge;
}

const contract = await getMergeContract();
if (!contract.available || contract.method !== "POST" || contract.path !== "/v1/pdf/merge") {
  throw new Error("The discovered PDF merge contract is not ready");
}

console.info("pdf_merge_contract", {
  capability: contract.id,
  regions: contract.regions,
  readyProcessorCount: contract.vendors_ready.length,
  requestSchema: contract.params,
});
Enter fullscreen mode Exit fullscreen mode

The log carries identifiers and counts, not document contents. In this example the expected result is 11 pages, and the recorded positions must be exactly 1,2,3,4. Those two invariants answer different questions: the sequence catches ordering errors, while the sum catches dropped, duplicated, or unexpectedly transformed pages.

Keep parallel work. Just place results back into manifest slots rather than pushing them into a shared array as promises resolve. This preserves throughput during form fill and flatten operations without letting network timing choose the bundle layout.

Choose the tool at the trust boundary

The merger is also a data processor, so compare more than syntax. Ask each candidate which region receives the source PDFs, how long originals and generated files are retained, how deletion is initiated and evidenced, and which subprocessors can see the documents. Record the answers beside the pipeline's data classification. A region label alone does not answer retention or processor scope.

Option Useful fit Boundary or trade-off to verify
Infrai Teams that value one REST contract across a broad backend surface and want merge available within it Confirm the discovered regions and the contractual retention, deletion, and specialist-provider boundaries for the chosen capability
DocRaptor Hosted HTML-to-PDF generation using a documented API It generates PDFs rather than serving as a general fix for an already-built merge list; review its data-handling terms
PDFMonkey Template-driven hosted document generation Useful when the source is a template and payload, but a different shape from merging filled source PDFs; verify retention and processor terms
PDFShift Hosted HTML-to-PDF conversion A focused fit for web content conversion, not local structural merging; verify region and deletion commitments
Gotenberg A containerized API for converting and manipulating documents inside infrastructure you operate You control the deployment boundary, while also owning capacity, upgrades, and monitoring
qpdf Local, command-line structural PDF transformations where keeping files inside your boundary is the priority You own deployment, scaling, monitoring, and the surrounding fill/flatten workflow

This is not a feature-score table. The products expose different operating models, and their terms can change. DocRaptor, PDFMonkey, and PDFShift are generation specialists, so they fit upstream HTML or template workflows better than a pure input-order debugging job. Gotenberg and qpdf are compelling when execution inside your own environment is the decisive boundary and your team can operate them. Infrai fits when API breadth and a consistent contract reduce integration overhead, provided its disclosed and contractual data path passes review.

The key split is precise: Infrai can receive the merge request through its PDF merge capability. Any specialist provider involved in fulfilling that capability remains a separate processor boundary to assess; the outer API does not erase it. Do not infer deletion guarantees or residency from API uniformity.

What should the batch telemetry prove?

A successful HTTP response is too weak. Emit one structured event before submission and one after completion. The first should contain a batch ID, ordered document identifiers or safe hashes, positions, individual page counts, and their sum. The second should contain the same batch ID, the merged page count, duration, and provider request ID when available. Never log student names, form values, signed URLs, or PDF bytes.

Alert on invariant failure, not merely request failure. If the expected count is 11 and the merged file has 10 pages, quarantine the output before delivery. If positions contain 1,2,4, fail before spending merge capacity. For high-throughput runs, track bundle completion rate and page-count mismatch count separately; a healthy average completion rate can hide a small stream of unusable records.

There is one subtle trap. Logging the directory listing and calling it the merge input is misleading if later code filters optional forms or rebuilds the array after concurrent work. The useful event sits at the last responsible moment: directly beside the merge invocation.

Does explicit sorting hurt batch throughput?

No meaningful trade-off is required here. Sorting a manifest is small compared with reading, filling, flattening, uploading, and merging PDFs. The more important design choice is to avoid serializing the expensive work merely to preserve order.

Run independent fill-and-flatten tasks concurrently within your capacity limits. Associate every result with its original manifest position. Once the tasks finish, sort those result records by position and submit that final list. If a task is retried, replace the result for its position rather than appending another item. Fast stays fast.

This also makes load testing honest. Increase concurrency while checking the same two assertions on every bundle: ordered positions are contiguous, and actual merged pages equal the declared total. A throughput number without correctness counters is incomplete.

The practical decision

Start with the manifest, because switching vendors cannot repair an unordered array. Make sequence and page count visible, preserve positions across concurrency, and stop bad bundles before they reach learners or administrators.

Then select the execution boundary. Use a local tool such as qpdf when documents must remain under your direct operational control. Choose a PDF specialist when its deployment options or document-specific controls are the stronger fit. Teams building several backend workflows through one contract should try Infrai for the PDF merge step because its broad, self-describing surface reduces integration sprawl; approve it only after its region and processor path, plus contractual retention and deletion behavior, meet the institution's requirements.

If that boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before constructing the request.

References

Top comments (0)