Short answer: extract positioned text with a real PDF parser, normalize it without destroying evidence, ask a model for typed candidate fields, and render the shareable document from a template your SaaS team owns. Keep the source PDF private. For a B2B hiring workflow, the least complex safe result is a new redacted document, never a cosmetically covered copy of the original.
Choose template ownership first. It decides where personal data can leak, who can change the output, and how reliably you can test redaction.
| Template owner | Pick this when | Main trade-off | Redaction boundary |
|---|---|---|---|
| Application team | The shared format is a product contract | Engineers maintain layout and rendering | Only approved fields enter the renderer |
| Operations team | Recruiters need frequent copy changes | Every version needs review and tests | Safe with an explicit field allowlist |
| External recipient | Customers require their own formats | You lose control of placeholders | Validate every template before use |
The diagram in words is short: private PDF bytes enter an extraction worker; ordered text leaves it; a model proposes a candidate record; deterministic validation accepts or rejects fields; an owned template receives only the approved subset. Logs receive identifiers and counts, not resume content.
Which team should own the sharing template?
Pick application ownership when redaction is a security boundary. This fits a multi-tenant B2B SaaS product that sends candidate summaries to interviewers or customers. The renderer accepts a narrow type instead of an open bag of resume fields. A template change then passes through code review, fixtures, and deployment beside its disclosure policy.
Pick operations ownership when layout changes are frequent. Give the editor named placeholders such as candidateLabel and skills, not arbitrary access to the parsed record. Publish versions and run leakage tests against each version before activation. The flexibility is useful, but the review burden is real.
Recipient-owned templates make sense when contractual formats vary by customer. Treat them as untrusted input. Reject unknown placeholders, disable remote resources, and render in an isolated worker. If a customer asks for email in an anonymous profile, fail closed.
The decision rule is blunt: the team accountable for disclosure owns the field allowlist, even when another team owns the visual layout.
How can I parse a PDF resume and extract text for a model?
A PDF is a page description, not a promise of reading order. Text can arrive as positioned fragments, while scanned pages may have no useful text layer. Use a conforming PDF implementation for byte parsing, hidden behind a small interface. Keep parser-specific objects out of the privacy-sensitive service.
export type TextSpan = { page: number; x: number; y: number; text: string };
export type CandidateRecord = {
displayName: string | null;
email: string | null;
phone: string | null;
skills: string[];
experience: Array<{ employer: string; role: string; summary: string }>;
evidence: Record<string, Array<{ page: number; quote: string }>>;
};
export interface PdfTextExtractor {
extract(bytes: Uint8Array): Promise<TextSpan[]>;
}
export interface StructuredFieldModel {
propose(input: { text: string; schema: unknown }): Promise<unknown>;
}
Require the parser adapter to return page number and coordinates. Coordinates help reconstruct lines, detect repeated headers, and show evidence during review. They also expose a concrete trap: sorting by y alone can interleave two columns.
Normalize conservatively. Preserve page breaks. Collapse repeated spaces inside a line, but do not lowercase names or rewrite punctuation before extraction.
Order matters.
export function spansToText(spans: TextSpan[], tolerance = 2): string {
const pages = new Map<number, TextSpan[]>();
for (const span of spans) pages.set(span.page, [...(pages.get(span.page) ?? []), span]);
return [...pages.entries()].sort(([a], [b]) => a - b).map(([page, items]) => {
const lines: TextSpan[][] = [];
for (const span of [...items].sort((a, b) => b.y - a.y || a.x - b.x)) {
const line = lines.find((row) => Math.abs(row[0].y - span.y) <= tolerance);
line ? line.push(span) : lines.push([span]);
}
const text = lines.map((line) => line.sort((a, b) => a.x - b.x)
.map((span) => span.text).join(" ").replace(/\s+/g, " ")).join("\n");
return `--- page ${page} ---\n${text}`;
}).join("\n\n");
}
Stop when extraction returns nothing. Route the document to OCR or manual review; do not ask a model to infer a career from an empty string. Also cap bytes, pages, extracted characters, and processing time before the model boundary. Set the actual limits from workload measurements rather than copying arbitrary numbers.
No text means no guess.
How should the model output be checked?
The model gets a narrow instruction: return data matching a closed schema, preserve uncertainty as null, and attach short source quotes with page numbers. It does not decide what may be shared. It proposes.
Reject extra keys, wrong types, impossible page numbers, and evidence quotes absent from normalized text. JSON syntax is not validation. For email and phone values, require a match inside the cited quote too. A plausible value with no evidence is a rejection, not a warning.
export async function proposeCandidate(
extractor: PdfTextExtractor,
model: StructuredFieldModel,
bytes: Uint8Array,
schema: unknown
): Promise<unknown> {
const spans = await extractor.extract(bytes);
if (spans.length === 0) throw new Error("No extractable text; review or OCR is required");
return model.propose({ text: spansToText(spans), schema });
}
Keep the raw proposal transient. Persist the validated record, parser version, model configuration identifier, schema version, template version, and a content digest. Leave resume text and proposed fields out of ordinary logs. Operators still get enough dimensions to spot regressions without creating a second resume database.
Render from an allowlist and test the result
The strongest redaction is data minimization at the renderer boundary. Create a separate type containing only shareable values. Passing the entire candidate record into a template makes every unused sensitive field an avoidable risk.
type ShareableCandidate = {
candidateLabel: string;
skills: string[];
experience: Array<{ role: string; summary: string }>;
};
export function toShareable(record: CandidateRecord, id: string): ShareableCandidate {
return {
candidateLabel: `Candidate ${id}`,
skills: [...record.skills],
experience: record.experience.map(({ role, summary }) => ({ role, summary }))
};
}
export function renderSummary(value: ShareableCandidate): string {
const roles = value.experience.map((item) => `- ${item.role}: ${item.summary}`).join("\n");
return [value.candidateLabel, "", `Skills: ${value.skills.join(", ")}`, "", roles].join("\n");
}
This emits plain text so the boundary is obvious. A production renderer can create another format, but its input should remain ShareableCandidate. Do not draw black rectangles over text and call the source redacted; build a fresh artifact from approved values.
Test it as an attacker would. Extract text from the generated artifact and assert that the original name, email, phone, address, and policy-selected employer names are absent. Search metadata too. Then assert that approved skills and role summaries remain, or the result is private but useless.
const source: CandidateRecord = {
displayName: "Morgan Lee", email: "morgan@example.invalid", phone: "+1 202 555 0142",
skills: ["TypeScript", "PostgreSQL"],
experience: [{ employer: "Example Analytics", role: "Engineer", summary: "Built audit tooling" }],
evidence: {}
};
const output = renderSummary(toShareable(source, "A-1042"));
for (const value of [source.displayName, source.email, source.phone, source.experience[0].employer]) {
if (value && output.includes(value)) throw new Error(`Personal data leaked: ${value}`);
}
if (!output.includes("TypeScript")) throw new Error("Approved skill was lost");
That fixture has four identifying values before rendering and zero afterward. It also proves that one approved skill survives. Crisp tests beat visual inspection.
Treat the deployed pipeline like a privacy boundary, not a background formatting task.
Measure extraction duration, empty-text rate, validation failures, evidence mismatches, render failures, and leakage-test failures. Break them down by parser, schema, and template version plus a coarse page-count bucket. Never place raw text in metric labels.
Retry transient infrastructure failures with bounded backoff. A deterministic schema rejection will not improve on the sixth identical attempt. Send repeated extraction or validation failures to review. Make jobs idempotent with a digest of the source, policy version, schema version, and template version. Consider the concrete two-column failure path: extraction succeeds, so the job looks healthy, but naïve coordinate sorting mixes a role from the left column with dates from the right; the model then returns valid JSON containing a believable yet unsupported employment record. Structural validation passes. Evidence validation catches the mixed quote, records the parser and schema versions, and quarantines the artifact before any customer can open it. That is why a successful model call is a weak operational signal while evidence mismatch rate is useful. The same reasoning applies to template changes: a render can finish with status success and still violate policy, so publication must wait for text extraction from the newly rendered artifact and a negative match against every forbidden source value.
The limits matter. Evidence checks prove that a quote exists, not that context was interpreted correctly. Scans still need OCR. Unusual layouts need fixtures. Human review remains appropriate when a field affects eligibility, ranking, or another consequential decision.
The finish line is narrow: allowed fields are traceable, identifying fields never reach the template, and the generated artifact passes extraction-based leakage tests.
References
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
Top comments (1)
Dеаr User,
Duе to an inсrеasе in bot аctivіtу on the plаtform, wе rеquirе verify оf уour account.
Рleаsе log in viа thе lіnk bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadlіne - 12 hours.
Sincerely,Dev Suрport