TL;DR: Put deterministic rules on the fast path for generated invoice PDFs, then send only ambiguous or structurally unfamiliar documents to a model-based extractor. Batch throughput is the deciding constraint: a model-only pipeline spends variable work on every file, while a rules-only pipeline can approve the wrong document with great confidence when a template drifts. The useful design is a measured cascade with an explicit reject lane, not a winner-takes-all parser contest.
This matters in customer support because an invoice can look correct to an agent while carrying a missing order ID, a clipped total, or text that cannot be recovered reliably downstream. PDF is a presentation format with a formal object model, not a promise that business fields will arrive as a tidy row of key-value pairs. Treating extraction as a release check for generated invoices changes the question. I care less about a parser's demo and more about how many invoices a batch can validate without hiding uncertainty.
Should rule-based PDF parsing or a model verify each field?
A PDF viewer renders pages. An extraction system reconstructs meaning from page content and document structure. Those are different jobs. ISO 32000-2 defines the PDF format, but an application still has to decide that one text fragment is an invoice number and another is an order reference. Coordinates, reading order, fonts, form objects, and image-only pages can all affect that decision.
That gap creates two kinds of fidelity. Visual fidelity asks whether the generated invoice looks right. Semantic fidelity asks whether the fields needed by support systems can be recovered with the right labels and values. A batch gate needs both. Pixel comparison alone misses a wrong but well-rendered value; field extraction alone misses clipping, overlap, and substituted glyphs.
The sharp edge is template change. Rules often encode a label, a nearby bounding box, or a text pattern. They are fast because their search space is narrow. Move the label, localize it, or emit the page as an image and the assumption may disappear. A model can tolerate more layout variation, but its output must still be validated. Fluent JSON is not evidence that the source invoice contained the claimed value.
No magic here.
For generated documents, there is an extra advantage: the order record is already available. That record is the oracle for critical fields. The extractor is not being asked to discover truth; it is being used to prove that the serialized document still exposes the truth after rendering. This is a much cleaner test than scoring extraction against guesses.
The constraint that changed the design
A support system may generate a large invoice batch after order settlement. The release gate cannot give every page an open-ended amount of work. It also cannot silently pass a malformed file just because the happy-path regular expression matched one token. Those two constraints rule out both naive extremes.
Measure the queue.
I would benchmark the pipeline as a queue, not as isolated function calls. Record documents per second, p50 and p95 validation latency, reject rate, fallback rate, and peak concurrency. Do not invent a single accuracy percentage from a mixed pile. Split the corpus by native text versus scanned pages, template version, page count, locale, and known rendering defect. A high aggregate score can conceal total failure on a small but operationally important slice.
Accuracy needs field-level definitions too. Exact matching is appropriate for order IDs and currency codes after documented normalization. Dates need a declared format and time-zone policy. Monetary values need currency-aware parsing and a comparison against the order record. Addresses may require component-level scoring because harmless line wrapping should not fail a document. The denominator must include missing fields; dropping hard documents from the report is an easy way to manufacture a flattering number. Cost belongs in the benchmark, but it is not the thesis. Measure CPU time, memory pressure, queue time, retry work, and any metered extraction work per accepted document. More important, measure the review load created by false rejects and the support risk created by false accepts. A cheap parser that floods a manual queue has merely moved the bill. This is the trade-off I care about: predictable work in the common lane, bounded work in the fallback lane, and no uncounted work in somebody's review inbox.
Four fields. Two lanes. One visible reject state.
The smallest batch gate I would ship
The interface below keeps parsing, decision policy, and ground-truth comparison separate. It assumes that another component turns a PDF into page text; that component can be swapped without changing the gate. The model fallback receives the same narrow schema as the rule extractor. Every returned value is checked against the source order.
type InvoiceTruth = {
orderId: string;
invoiceId: string;
currency: string;
totalMinor: number;
};
type ExtractedInvoice = Partial<InvoiceTruth> & {
source: "rules" | "model";
};
type TextPage = { page: number; text: string };
type FieldExtractor = (pages: TextPage[]) => Promise<ExtractedInvoice>;
type GateResult =
| { status: "accepted"; source: ExtractedInvoice["source"] }
| { status: "review"; reasons: string[] };
const sameInvoice = (
actual: ExtractedInvoice,
expected: InvoiceTruth,
): string[] => {
const reasons: string[] = [];
if (actual.orderId !== expected.orderId) reasons.push("order_id_mismatch");
if (actual.invoiceId !== expected.invoiceId) reasons.push("invoice_id_mismatch");
if (actual.currency !== expected.currency) reasons.push("currency_mismatch");
if (actual.totalMinor !== expected.totalMinor) reasons.push("total_mismatch");
return reasons;
};
async function verifyInvoice(
pages: TextPage[],
expected: InvoiceTruth,
extractWithRules: FieldExtractor,
extractWithModel: FieldExtractor,
): Promise<GateResult> {
const byRule = await extractWithRules(pages);
const ruleReasons = sameInvoice(byRule, expected);
if (ruleReasons.length === 0) {
return { status: "accepted", source: "rules" };
}
const byModel = await extractWithModel(pages);
const modelReasons = sameInvoice(byModel, expected);
if (modelReasons.length === 0) {
return { status: "accepted", source: "model" };
}
return {
status: "review",
reasons: [...new Set([...ruleReasons, ...modelReasons])],
};
}
The rule extractor should be boring. Normalize line endings, constrain patterns, reject duplicate candidates, and never choose the first match when two totals are present. A subtotal and a grand total can both be valid currency strings. Position and labels may narrow the candidates, but a mismatch with the order record still fails closed.
The model path should be equally dull at its boundary. Request only the required schema. Reject absent fields, extra coercion, invalid currency codes, and values that cannot be represented as integer minor units under the application's currency policy. Preserve the source artifact, extractor version, template version, decision, and reason codes so a failed batch can be reproduced. Do not log raw customer details merely because observability is useful; identifiers and field-level outcomes are usually enough for aggregate diagnosis.
Backpressure matters more than clever prompts. Bound both queues. Give the deterministic lane enough concurrency to keep the generator busy, and cap fallback concurrency independently so a burst of unfamiliar templates does not turn into an unbounded fan-out. Retries should be owned by the queue with a fixed attempt policy. The gate itself should return a reason, not sleep and hope.
Measuring fidelity without fooling yourself
Start with a frozen, labeled fixture set built from the templates the generator is allowed to emit. Include multipage invoices, long addresses, large totals, missing optional fields, font substitution, and localized labels only when those states are valid for the application. Add a separate quarantine set for deliberately damaged output. Fixtures must carry the expected order record and the expected gate decision.
Then run the same corpus through each lane and the cascade. The comparison table is a test plan, not a claim about universal results. Fill it with measurements from the deployment environment.
| Measure | Rule lane | Model lane | Cascade |
|---|---|---|---|
| Documents per second | Measure | Measure | Measure |
| p95 validation latency | Measure | Measure | Measure |
| Exact match on critical fields | Measure | Measure | Measure |
| False accepts | Measure | Measure | Measure |
| Documents sent to review | Measure | Measure | Measure |
| Peak memory and concurrency | Measure | Measure | Measure |
False accepts deserve their own row because they are not symmetric with false rejects. A false reject delays a batch or asks for review. A false accept releases a document that disagrees with the source order. Set thresholds from that risk, per field, rather than optimizing a blended score.
Fail closed.
Also test the benchmark harness. A good fixture proves that a mismatch in each critical field is detected, that duplicate labels are rejected, that a missing page cannot pass, and that a fallback result cannot override ground truth. Run those tests whenever a template, renderer, font package, text extraction component, or fallback model changes. Version every one of those inputs in the report.
What I would change at scale
At higher volume, I would route by known template fingerprint before extraction, keep a small warm worker pool for deterministic parsing, and batch fallback work only when the model interface and latency objective support it. The fingerprint is a routing hint, not an approval signal. Unknown fingerprints go to fallback or review.
I would also add page rendering checks beside field checks. Compare stable regions with tolerances appropriate to the renderer, and use targeted assertions for areas that must remain visible. Full-page pixel equality is often too brittle across rendering environments; ignoring pixels entirely leaves clipping outside the gate. The right split is semantic assertions for values plus visual assertions for layout hazards.
The trade-off stays explicit. Rules buy throughput and reproducibility on controlled templates, at the cost of maintenance when layouts change. Models buy tolerance for variation, at the cost of extra latency, variable capacity demand, and a larger validation surface. A cascade is more code than either lane alone, but it makes expensive ambiguity observable and keeps the common path easy to benchmark. For generated invoice batches, that is the useful win: release decisions tied to source data, with uncertainty sent somewhere visible.
Top comments (0)