TL;DR: For marketplace invoice ingestion, choose rules when the document family is controlled and signatures must be validated against the original bytes. Choose model-based extraction when layouts and labels vary, but keep it away from signature validation and require field-level evidence. In a mixed marketplace, the practical design is usually a cascade: preserve the source PDF, verify its signature, attempt deterministic extraction, route uncertain fields to a model, and record every decision in an append-only audit event.
| Pick this path | Best fit | Main failure mode | Evidence to retain |
|---|---|---|---|
| Rules | A few stable seller templates | Silent layout drift | Rule version, page, bounding box, raw value |
| Model | Many unknown layouts or scans | Plausible but unsupported values | Model version, prompt/schema version, page image, confidence |
| Cascade | Mixed marketplace traffic | Bad routing or thresholds | Every stage result and routing reason |
Accuracy alone is the wrong finish line. An invoice total that looks right but cannot be traced to page 2, or one extracted from a PDF whose signed byte range was altered, is weak evidence. Fidelity means preserving what the document said, where it said it, and which transformation produced the normalized value.
1. Should rule-based PDF parsing or model field extraction run first?
Start with document diversity, not fashion. A rule-based parser can be exact when sellers emit PDFs from a stable template: locate the text token near Invoice total, constrain its region, parse the currency, and verify the arithmetic. Its behavior is inspectable. A changed font, shifted column, flattened form, or scanned page can invalidate those assumptions without making the file invalid.
A model-based extractor handles label and layout variation better because it can map semantically similar phrases into one schema. It also introduces a different risk: a syntactically valid answer can lack document support. Never let schema validation masquerade as evidence validation.
For a marketplace, make the first attempt with rules only when a template fingerprint is known and the required text layer is present. Route everything else to model extraction. This keeps deterministic work deterministic while giving unfamiliar invoices a path forward.
Rules are not suitable for an open-ended seller population unless the team accepts continuous template maintenance. A model-only path is a poor fit when every accepted field must be reproducible byte for byte. A cascade carries more moving parts, so a small marketplace with three controlled templates may reasonably avoid it. Those are real limitations, not edge-case trivia.
2. Preserve signature truth before touching content
PDF digital signatures cover byte ranges in the original file. Parsing, rendering, optimizing, or regenerating the PDF before verification can destroy the exact artifact needed to assess that signature. So the source bytes come first. Store a cryptographic digest, verification result, signer certificate details, and verification time alongside the immutable object identifier.
This separation matters. Signature verification answers whether covered bytes and credentials validate under the chosen policy. Field extraction answers what those bytes appear to contain. Neither answer proves the other.
Pause here.
The diagram in words is short: upload bytes -> hash and preserve -> verify signature -> classify document -> extract candidates -> validate invoice constraints -> emit audit event. Keep the arrows one-way.
3. Make every field carry its own receipt
One document-level confidence score hides the exact disagreement an operator needs to inspect. Represent each extracted field as a claim with provenance. Here is a compact TypeScript shape for both paths:
type Method = "rule" | "model";
type Evidence = {
page: number;
bounds: [number, number, number, number];
sourceText: string;
};
type FieldClaim<T> = {
name: "invoiceNumber" | "issuedAt" | "currency" | "total";
value: T;
method: Method;
extractorVersion: string;
evidence: Evidence[];
confidence?: number;
};
type AuditEvent = {
sourceSha256: string;
signatureStatus: "valid" | "invalid" | "unverified";
claims: FieldClaim<string | number>[];
recordedAt: string;
};
Do not overwrite sourceText with the normalized value. If the invoice prints 1.234,50 EUR, the audit record needs that literal text even if downstream accounting receives 1234.50 and EUR. Page numbers and bounding boxes let a reviewer return to the visual source instead of trusting a parser log.
Rules should identify themselves by a versioned rule set. Model runs need a model identifier plus the prompt or schema version. A retry is a new attempt, not an edit to the old event. That distinction makes reprocessing explainable months later.
Consider a two-page seller invoice that prints Order 7814 near the header, carries line items onto page 2, and places TOTAL EUR 842,10 below a repeated subtotal. A coordinate rule tied to page 1 may return the subtotal. A model may return the final total, yet omit which occurrence supported it. The useful result is neither bare value: it is the final total plus the literal source span, page, bounds, extractor version, and reconciliation outcome. If the evidence points at the subtotal, reject the claim even when another field happens to contain the right number. This is the trade-off in concrete form: semantic flexibility earns a wider input envelope, while explicit provenance supplies the control needed to trust that flexibility.
4. Test disagreement, drift, and abstention
Build the test set around failure classes. Include native-text PDFs, scans, rotated pages, duplicate labels, credit notes, multi-currency orders, line items split across pages, and signed revisions. Keep expected values and expected evidence locations. A correct value from the wrong label should fail.
Measure field-level exact match after explicit normalization, but also measure unsupported-claim rate, abstention rate, signature-verification coverage, and human-review rate. Totals deserve arithmetic checks: subtotal plus tax and adjustments should reconcile under the invoice's rounding convention. Invoice number and seller identity deserve cross-field checks against the order record.
Then shadow a candidate extractor on production-shaped traffic. Compare its claims with the active path without publishing them. Alert on shifts in missing-field rate, rule fallback rate, confidence distribution, page-count distribution, and reconciliation failures. Logs should carry the source digest and attempt ID, never the full invoice or personal data by default.
A crisp before/after helps. Before: total=842.10, confidence=0.97. After: total=842.10, page=2, bounds=[412,690,518,714], source="TOTAL EUR 842,10", method=model, extractor=v7. The second record is debuggable.
Tiny records win.
5. Set operational limits before rollout
Models do not remove rules; they move rules to validation and routing. Define which fields may auto-post, which require reconciliation, and which always require review. Define timeouts and bounded retries. If extraction fails, retain the source and emit a visible terminal status rather than an empty invoice object.
Cost belongs in capacity planning, not the headline decision. Track compute or request consumption per accepted invoice, plus review minutes and reprocessing volume. A cheap attempt that creates unsupported totals can be expensive operationally, while an elaborate path for a stable template wastes latency and complexity.
The limit is clear: no extractor can repair missing source information, prove a signature from regenerated bytes, or turn confidence into provenance. Use rules for known structure, models for layout variation, and evidence-bearing claims for both. That is the boundary that keeps invoice automation auditable.
Sources
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- NIST Digital Signature Standard (FIPS 186-5): https://csrc.nist.gov/pubs/fips/186-5/final
- W3C Trace Context: https://www.w3.org/TR/trace-context/
- OpenTelemetry logs data model: https://opentelemetry.io/docs/specs/otel/logs/data-model/
Top comments (0)