DEV Community

evanshepherd5623
evanshepherd5623

Posted on

High-Throughput PDF Redaction: Split Chapter Ranges from Parsed Contents Safely

TL;DR: Parse the table of contents into ordered printed-page starts, translate those starts to zero-based PDF indexes once, and derive every chapter end from the next start. Redact the source before any chapter bytes cross the sharing boundary. For large fintech bundles, keep parsing and range validation serial, then copy validated ranges with bounded concurrency.

Input condition Boundary method Pick this when Throughput consequence Main check
Reliable contents text and one stable page offset Parsed contents entries Statements have consistent numbering One parse feeds every copy Starts increase and remain in bounds
Native outline with usable destinations Outline destinations The producer preserves bookmarks No contents text parser Every destination resolves
Neither source is trustworthy Page classification Mixed-origin bundles need content-based boundaries More work per page Low confidence goes to review

The parsed-contents path is a clean fit for a generated document bundle only when the page-coordinate problem is explicit. A contents entry saying page 1 may point to PDF index 4 because a cover, notices, and the contents itself precede the numbered body. Guessing that offset inside a copy loop can produce convincing but incorrect files.

Boundaries first. Bytes second.

How should Node.js split PDF chapter ranges from parsed contents?

Start with evidence already present in the document. A native outline can identify pages directly. Parsed contents entries work when the generator emits predictable lines such as Account summary .... 1. Page classification is the fallback when neither structure survives assembly.

Do not combine all three into a voting system by default. Each detector adds disagreement states that need policy and observability. Select a primary method for a document class and define a rejection path. If a parsed table has duplicate starts, a descending page number, or a final start beyond the page count, reject the bundle before creating shareable artifacts.

Fast failure protects correctness and batch capacity. For financial documents, an uncertain boundary should become a review item rather than a best-effort split. No output is better than personal data in the wrong chapter file.

That is the first trade-off: rejecting one questionable bundle reduces nominal completion volume, but preserves the meaning of “completed.” Imagine the concrete failure chain. A parser accepts three entries numbered 1, 4, and 7. The first two map correctly after a four-page offset, but an inserted notice before the final chapter changes the offset there. The ranges still increase, the files still open, and a basic health check stays green. Yet the second file receives an extra page and the third loses one. Range checks alone cannot catch a changing offset; document-class validation or page-level verification must. That is why a single-offset parser is not a fit for bundles assembled from independently paginated sources.

Pick outline destinations when structure survives assembly

Choose outline destinations when the upstream producer controls the PDF and preserves them through concatenation. The splitter can resolve each destination to a page index, sort the results, and apply the same validator used below. It skips text extraction.

A bookmark label is not proof that its destination is correct. Validate order, uniqueness, and bounds anyway. Keep labels in a protected audit record, not in metric labels; a chapter title may contain account-specific text and can create unbounded cardinality.

Short path. Same scrutiny.

Pick parsed entries for stable printed pagination

This is the practical middle ground for statement bundles with a readable contents page but no dependable outline. Keep the parser narrow. Accept the exact line shapes emitted by the document producer, preserve the raw line for protected diagnostics, and reject ambiguity. A parser that silently skips unfamiliar lines raises apparent throughput while lowering correctness.

Use one explicit conversion: pdfIndex = printedPage + bodyStartIndex - firstPrintedPage. With a four-page front section, bodyStartIndex is 4. Printed page 1 maps to zero-based PDF index 4. The arithmetic is small. Its blast radius is not.

A 12-page example makes the boundary rule visible. If three chapters begin at PDF indexes 4, 7, and 10, their half-open ranges are [4, 7), [7, 10), and [10, 12). Half-open ranges match array slicing semantics and remove the usual end-page adjustment.

The limitation is precise: this method trades broader document compatibility for cheaper, deterministic range construction. Use outline destinations instead when bookmarks survive. Use page classification when pagination changes within the bundle. Neither alternative is universally better; each moves work from parsing into destination resolution or per-page analysis.

Implement a validated, redaction-first pipeline

First turn parsed entries into ranges. Separating this step from PDF I/O makes edge cases cheap to test. The checks are deliberate: integers only, strict ordering, non-empty ranges, and a final boundary equal to the page count.

export type TocEntry = { title: string; printedPage: number };
export type ChapterRange = {
  title: string;
  startIndex: number;
  endIndexExclusive: number;
};

export function buildChapterRanges(
  entries: readonly TocEntry[],
  pageCount: number,
  bodyStartIndex: number,
  firstPrintedPage = 1,
): ChapterRange[] {
  if (!Number.isInteger(pageCount) || pageCount < 1) {
    throw new Error("pageCount must be a positive integer");
  }
  if (!Number.isInteger(bodyStartIndex) || bodyStartIndex < 0) {
    throw new Error("bodyStartIndex must be a non-negative integer");
  }
  if (entries.length === 0) throw new Error("No chapter entries were parsed");

  const starts = entries.map((entry) => {
    if (!entry.title.trim()) throw new Error("Chapter title is empty");
    if (!Number.isInteger(entry.printedPage)) {
      throw new Error(`Non-integer page for ${entry.title}`);
    }
    return entry.printedPage + bodyStartIndex - firstPrintedPage;
  });

  for (let index = 0; index < starts.length; index += 1) {
    const start = starts[index];
    if (start < 0 || start >= pageCount) {
      throw new Error(`Chapter start ${start} is outside the PDF`);
    }
    if (index > 0 && start <= starts[index - 1]) {
      throw new Error("Chapter starts must be strictly increasing");
    }
  }

  return entries.map((entry, index) => {
    const startIndex = starts[index];
    const endIndexExclusive = starts[index + 1] ?? pageCount;
    return { title: entry.title, startIndex, endIndexExclusive };
  });
}
Enter fullscreen mode Exit fullscreen mode

Test the conversion as a pure rule. This case covers the front-section offset, adjacent ranges, and the last chapter reaching end-of-file.

import assert from "node:assert/strict";
import test from "node:test";
import { buildChapterRanges } from "./ranges.js";

test("maps printed pages to contiguous PDF ranges", () => {
  const ranges = buildChapterRanges(
    [
      { title: "Account summary", printedPage: 1 },
      { title: "Transactions", printedPage: 4 },
      { title: "Tax documents", printedPage: 7 },
    ],
    12,
    4,
  );

  assert.deepEqual(ranges, [
    { title: "Account summary", startIndex: 4, endIndexExclusive: 7 },
    { title: "Transactions", startIndex: 7, endIndexExclusive: 10 },
    { title: "Tax documents", startIndex: 10, endIndexExclusive: 12 },
  ]);
});
Enter fullscreen mode Exit fullscreen mode

Next, put a narrow PDF adapter around the range logic. This example uses a TypeScript PDF library for copying pages. The security and scheduling policy stays outside it. redactPersonalData must return sanitized bytes by removing the underlying sensitive content; painting a visible rectangle over text is insufficient.

import { PDFDocument } from "pdf-lib";
import type { ChapterRange } from "./ranges.js";

type SplitResult = { title: string; bytes: Uint8Array };
type Redactor = (source: Uint8Array) => Promise<Uint8Array>;

async function copyRange(
  source: PDFDocument,
  range: ChapterRange,
): Promise<SplitResult> {
  const output = await PDFDocument.create();
  const indexes = Array.from(
    { length: range.endIndexExclusive - range.startIndex },
    (_, offset) => range.startIndex + offset,
  );
  const pages = await output.copyPages(source, indexes);
  for (const page of pages) output.addPage(page);
  return { title: range.title, bytes: await output.save() };
}

export async function redactThenSplit(
  original: Uint8Array,
  ranges: readonly ChapterRange[],
  redactPersonalData: Redactor,
  concurrency = 2,
): Promise<SplitResult[]> {
  if (!Number.isInteger(concurrency) || concurrency < 1) {
    throw new Error("concurrency must be a positive integer");
  }

  const sanitized = await redactPersonalData(original);
  const source = await PDFDocument.load(sanitized);
  const results = new Array<SplitResult>(ranges.length);
  let cursor = 0;

  async function worker(): Promise<void> {
    while (cursor < ranges.length) {
      const index = cursor;
      cursor += 1;
      results[index] = await copyRange(source, ranges[index]);
    }
  }

  const workerCount = Math.min(concurrency, ranges.length);
  await Promise.all(Array.from({ length: workerCount }, () => worker()));
  return results;
}
Enter fullscreen mode Exit fullscreen mode

Why begin with two workers rather than one promise per chapter? It is a safety bound, not a benchmark claim. Each output allocates pages and bytes. Unlimited fan-out can turn a throughput optimization into memory pressure, while one worker may leave capacity unused. Measure queue wait, redaction duration, copy duration, output bytes, and failures by stable reason code. Tune with representative bundles in the actual runtime.

Measure before raising it.

Keep personal values out of logs. A useful event has a random job identifier, page count, chapter count, elapsed milliseconds, and a fixed result such as invalid_range. It does not need a customer name, account number, title, or extracted contents line. Metrics can use the same low-cardinality vocabulary: completed bundles, rejected bundles, pages processed, bytes produced, and duration histograms.

The deployment boundary matters too. Store original and sanitized bytes separately, grant the splitting worker access only to sanitized input, and publish outputs only after every expected range succeeds. An all-or-nothing publish step prevents consumers from seeing a partial bundle as complete. Retry jobs around deterministic inputs. Do not retry malformed ranges until the source or parser configuration changes.

Limits to keep visible

This design does not discover the body offset. The producer must supply it, or a separate validated step must derive it. Roman numerals, inserted notices, rotated scans, and repeated printed page numbers can invalidate a single-offset model. Route those document classes to outline resolution or page classification instead of stretching the formula.

Page copying is not redaction. The order is a security invariant: sanitize, verify, split, then publish. Add tests that search extracted text and inspect rendered pages for every sensitive field class in the policy; neither check alone establishes correctness for every PDF representation. Keep unsanitized input outside the sharing path and expire intermediate data under the system retention policy.

Raise worker concurrency only while completed pages per unit time improve and memory, queue delay, and rejection rates remain inside operating limits. Stop when more parallel work increases contention. Crisp boundaries beat optimistic fan-out.

References

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dеаr User,
Duе to an inсrеase іn bot actіvity on thе рlatform, we requіre vеrify оf уour account.
Рlease log in vіa the link bеlоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadlinе - 12 hours.
Sincerely,Dev Suррort

‍​​‍‌