DEV Community

FrozenSigh2853916
FrozenSigh2853916

Posted on

Gaming Evidence Redaction in Node.js: PDF Text Extraction, OCR, and Signed Audit Records

A gaming archive has a hard constraint: personal data must be removed before a document leaves the trust boundary, while reviewers still need evidence that the released file came from a specific input and passed a defined policy. TL;DR: inspect every page for usable native PDF text, send only uncertain pages through OCR, redact against one coordinate model, and sign an audit record that binds the input, policy, decisions, and output. Do not choose one method for the whole archive. Choose per page, then fail closed when confidence is insufficient.

The signature is not proof that every redaction is semantically correct. It proves which bytes and decisions were reviewed.

Should you use an OCR API or PDF text extraction?

A PDF is a container, not a promise that visible words are available as a clean reading-order string. ISO 32000-2 defines the format; an application still has to interpret page content, fonts, images, and document structure. For redaction, ask a narrower question: can the pipeline map each sensitive token to a dependable region on the rendered page?

Consider a player-support file. Page 1 may be an exported account summary with selectable text. Page 2 may be a screenshot of chat. Page 3 may contain a scanned identity document. Page 4 may have a text layer produced by an earlier OCR pass, but its coordinates can disagree with what a reviewer sees. One file can contain three extraction conditions.

The before model is tempting: PDF -> extract text -> find email -> draw black box. It is incomplete. A black rectangle can hide pixels while leaving underlying text available, and a correct string with a bad box can cover the wrong player name. Imagine the support bundle's fourth page after an account appeal: the parser returns the email correctly, but the page is rotated 90 degrees and the stored box follows the unrotated coordinate space. The detector looks perfect in a unit test that checks strings. The released page is wrong because the rectangle lands beside the address. This is why the route cannot be based on extracted character count alone. Render the page, transform the candidate box into rendered coordinates, and check that it intersects the ink you intend to remove. If the text and pixels disagree, preserve the evidence and require review instead of guessing.

Stop there.

The after model has gates: page inventory -> render probe -> native-text probe -> OCR fallback -> candidate detection -> irreversible redaction -> visual verification -> signed audit record. Each arrow emits evidence. Crisp. Inspectable.

Native extraction should win when it returns meaningful text, stable page coordinates, and a rendering check confirms those coordinates cover the visible glyphs. OCR should win for image-only pages and pages whose text layer fails those checks. A native layer is not automatically safer merely because it exists.

Route pages with evidence, not file extensions

The router needs more than a Boolean named hasText. Give it observations that can be logged without retaining personal content: visible character count, text-box count, proportion of boxes inside the page, overlap between text boxes and rendered ink, and whether expected script families appeared. Exact thresholds belong to your test corpus. They are policy, not universal constants.

import { createHash } from "node:crypto";

type Probe = {
  page: number;
  visibleCharacters: number;
  boxesInsidePageRatio: number;
  renderedInkOverlapRatio: number;
};
type Route =
  | { kind: "native"; reason: string }
  | { kind: "ocr"; reason: string }
  | { kind: "review"; reason: string };

const policy = {
  id: "gaming-redaction-v3",
  minimumVisibleCharacters: 40,
  minimumBoxesInsidePageRatio: 0.98,
  minimumRenderedInkOverlapRatio: 0.75
};

function routePage(probe: Probe): Route {
  if (probe.visibleCharacters < policy.minimumVisibleCharacters) {
    return { kind: "ocr", reason: "sparse-native-text" };
  }
  if (probe.boxesInsidePageRatio < policy.minimumBoxesInsidePageRatio) {
    return { kind: "ocr", reason: "invalid-native-coordinates" };
  }
  if (probe.renderedInkOverlapRatio < policy.minimumRenderedInkOverlapRatio) {
    return { kind: "review", reason: "text-render-mismatch" };
  }
  return { kind: "native", reason: "native-probes-passed" };
}

function sha256(bytes: Uint8Array): string {
  return createHash("sha256").update(bytes).digest("hex");
}
Enter fullscreen mode Exit fullscreen mode

Those three numbers are example configuration, not performance claims. Start them in version control, exercise them against approved fixtures, and change them through review. The policy ID tells an auditor which boundary applied.

Do not log token values, player handles, email addresses, chat text, or OCR payloads. Emit counts and controlled reason codes. For a 12-page file, an event can say that eight pages used native extraction, three used OCR, and one required review. That is useful telemetry without becoming a second personal-data archive.

Extraction and OCR should converge immediately into the same internal type: tokens plus page-relative boxes. Normalize rotation and coordinate origins there. Candidate detection consumes one representation, and the renderer receives explicit regions to remove. This keeps two implementations from quietly developing different rules.

An audit record should include the input digest, output digest, policy identifier, per-page route and reason, detector rule identifiers, counts of proposed and approved regions, review disposition, and timestamps. Canonicalize it before signing so field order cannot change the signed bytes. Keep the signing key outside the document worker and record a key identifier, not the private key.

type PageDecision = {
  page: number;
  route: "native" | "ocr" | "review";
  reason: string;
  proposedRegions: number;
  approvedRegions: number;
};
type AuditRecord = {
  schemaVersion: 1;
  inputSha256: string;
  outputSha256: string;
  policyId: string;
  pages: PageDecision[];
  completedAt: string;
  signingKeyId: string;
};
interface AuditSigner {
  sign(canonicalRecord: Uint8Array, keyId: string): Promise<Uint8Array>;
}

function canonicalAuditBytes(record: AuditRecord): Uint8Array {
  const ordered = {
    schemaVersion: record.schemaVersion,
    inputSha256: record.inputSha256,
    outputSha256: record.outputSha256,
    policyId: record.policyId,
    pages: record.pages.map((page) => ({ ...page })),
    completedAt: record.completedAt,
    signingKeyId: record.signingKeyId
  };
  return new TextEncoder().encode(JSON.stringify(ordered));
}
Enter fullscreen mode Exit fullscreen mode

Sign only after the final output bytes exist. Verification should recompute both document hashes, recreate the canonical bytes, validate the signature with the identified public key, and confirm that the policy ID is recognized.

That is the trade-off.

Keep original and released artifacts in separate access domains. The audit record can travel with the released artifact only if its fields contain no personal data. A mapping between an internal case identifier and a player belongs in a restricted case system, not in a portable receipt.

Test the pixels, text, and trail

A green HTTP response proves little. Fixtures should isolate failure modes: a born-digital page, an image-only scan, rotated text, a mixed page with a screenshot, a stale OCR layer, text outside the crop box, and a page where the detector finds nothing. Use synthetic names and addresses.

For every fixture, assert three layers. First, semantic: expected candidates are proposed. Second, spatial: boxes cover intended rendered regions after rotation and cropping. Third, residual: the released artifact no longer exposes the target through text extraction or rendered pixels. Then verify that changing one byte of the output or audit record makes signature verification fail.

Track routing ratios by document source and policy version, review rates by reason code, duration by route, and failures by stage. A sudden rise in OCR routing may mean a source changed its export. A rise in text-render-mismatch deserves investigation before release. Alerts should point to a stage and reason, not include captured content.

Deployment needs a shadow phase. Run the new policy on representative, approved fixtures and compare decisions with the current policy without releasing its output. Promote by policy ID, preserve the old verifier, and make rollback select the prior policy rather than mutate historical records. Signatures make history tamper-evident; versioning makes it understandable.

What if OCR is accurate enough for every page?

Using OCR everywhere creates one apparent path, but it discards a useful signal: a trustworthy native layer can preserve encoded characters and coordinates without asking an image recognizer to infer them again. The decision is about minimizing transformations while meeting the redaction guarantee.

An all-OCR policy can fit a controlled collection known to be image-only. Record that assumption and test for violations. If a later upload contains born-digital annotations, the inventory gate should surface the change rather than silently treating it as familiar. This hybrid design also has limitations: it is a poor fit when the team cannot maintain rendered-page fixtures, operate a review queue, or protect signing keys. In that setting, narrow the accepted document set or require manual redaction until those controls exist. OCR adds image processing and an uncertain recognition step; native extraction depends on trustworthy text geometry. Review adds latency. The signed record improves traceability, but it cannot rescue a weak detection policy or certify that a human approved the right region.

The other objection is operational complexity. Yes, a page router adds branches. It also adds named failure states and makes review capacity measurable. For a gaming support team sharing evidence with moderators, legal reviewers, or platform partners, an explicit review result is better than a confident-looking release built from an uncertain page.

The decision rule is direct: prefer native extraction only after text-to-pixel checks pass; use OCR for pages without a dependable text layer; require review when neither route produces verifiable regions. Bind the decision to the exact output with a signed, content-free audit record.

Sources and References

Top comments (0)