
Most production LLM extraction pipelines route every field through a frontier model, every time. It works, and for a while nobody questions the bill. Then volume grows, or someone runs the per-document math, and the real question shows up. How much of this spend is buying accuracy, and how much is just habit.
In practice, most of it is habit. The majority of fields in a typical extraction job do not need a frontier model at all. A large share of the model calls answer questions the document never asked. A surprising fraction of calls re-extract data that has not changed since the last run. None of that is a model problem. It is an architecture problem, and architecture problems are the good kind, because you can fix them without waiting on anyone's roadmap.
I build cost-optimized inference for document-extraction workloads with 100+ fields. This is the practitioner's version of what actually moves the number. The moves below, applied in sequence, take per-document cost down by 85 to 93 percent with negligible accuracy impact. The interesting part is not any single technique. It is why they compound instead of overlap.
TL;DR
- Set a trusted baseline with the expensive model first, then prove a cheaper model can match it.
- Skip fields a document type cannot contain (adaptive selection).
- Route each field to the cheapest model that can do the job (tiered routing + distillation).
- Stop re-extracting unchanged fields (field-level caching).
- Stop paying for idle (serverless scale-to-zero).
- The savings multiply, not add, because each move attacks a different kind of waste.
🧭 Set the baseline with the expensive model first
Before optimizing anything, run the workload on a frontier model and treat those results as the baseline. Early on the goal is not low cost. It is output you trust and a confidence signal you believe. Once that baseline exists, it becomes the yardstick for everything cheaper.
Then, field by field, test whether a smaller and cheaper model can match it. Run the cheap model over the same documents, compare against the frontier baseline, and keep the fields where agreement is high. The fields that match move to the cheaper tier. The fields that do not stay on the frontier model.
This is how the tier percentages below get decided in practice. You are not guessing which fields a lightweight model can handle. You are measuring it against output you already trust.
The same comparison runs later as the quality-floor monitor at the end of this post. It starts here, as the step that earns each downgrade.
✂️ Start by not doing unnecessary work: adaptive field selection
The cheapest extraction call is the one you never make.
A contract schema might define 100+ fields, but a given document, say a simple service agreement, only contains 40 to 50 of them. Extracting the rest produces "not found" responses that cost tokens and return nothing. Classify the document type first, then extract only the fields that can actually appear in it.
This one move eliminates 30 to 40 percent of calls per document. It is the right thing to do first, because every later optimization then operates on a smaller base.
🎯 Right-size the model to the field: tiered routing
Not every field needs the same horsepower. In a typical extraction job:
- ~65 to 70 percent of fields are structurally simple: dates, amounts, party names, reference numbers, categorical flags. Predictable formats, consistent locations. A lightweight model handles these at a fraction of frontier cost with equivalent accuracy.
- ~20 percent need moderate reasoning: interpreting indirect phrasing, resolving ambiguous references, synthesizing across sections. A mid-tier model (for example a Haiku or Nova Lite class model) is plenty.
- ~10 percent genuinely need a frontier model: obligation summaries, risk assessment, termination-condition analysis, cross-document context.
Route each field to the cheapest tier that can do the job. The routing decision is nearly free. It is a field-type to tier lookup table plus a confidence-based fallback, so a tier that returns low confidence escalates to the next one. No ML is needed for routing, because the mapping is stable for a given document corpus. Blended per-document cost drops 60 to 75 percent, and quality holds because the hard fields still get frontier attention. The split above is not assumed. It comes straight out of the baseline comparison from the previous section.
Distillation: earn your way to a cheaper Tier 1
Tiered routing gets better when Tier 1 is a model you distilled for your own documents. Run a frontier model (the teacher) over 200 to 500 representative documents to produce gold-standard extractions, then train a smaller student model to reproduce that behavior on your specific fields. Managed services such as Amazon Bedrock Model Distillation handle the training without ML expertise.
The student model learns your narrow task, the dates, amounts, party names, and term lengths you see every day, and it runs much cheaper and faster with a small accuracy gap (about 2 percent) on the tasks it was trained for. Fields outside its training distribution fall through to the mid-tier or frontier model.
The bootstrap problem is obvious. You do not have 200 documents on day one, so distill progressively. Months one and two run routing-only on general-purpose models (40 to 50 percent savings) while storing every validated result as future training data. In month three, once you have accumulated enough examples, distillation runs and Tier 1 becomes your student model (another 30 to 40 percent off). Each quarter you re-distill with more data, and the share of fields handled by the cheapest tier climbs from about 70 percent toward 85 percent. The cost curve bends down the longer the system runs.
♻️ Stop paying twice: field-level caching
Documents get reprocessed when schemas are updated, prompts improve, or a quality review flags systematic errors. Without caching, reprocessing 200 documents costs exactly what processing them the first time did.
Key each result by the document-section SHA-256, the field name, the prompt version, and the model version. If all four match, return the cached value with zero inference. When you update a prompt for one field type, only that field's cache entries invalidate, and everything else serves from cache. A prompt update typically touches three to five field types, so 85 to 90 percent of fields serve from cache on reprocessing. For a 200-document corpus reprocessed after a four-field prompt change, that is 800 extractions instead of 7,000, an 88 percent reduction on that cycle.
🛑 Stop paying for idle: serverless scale-to-zero
A 200-document-per-month workload is idle most of the time. Documents land in object storage, a lightweight function writes metadata that triggers a stream event, and a containerized extraction agent runs only when work exists. During idle periods the system costs nothing beyond storage. For bursty, low-volume workloads this removes the fixed infrastructure floor that usually makes small deployments uneconomical. KEDA is a clean way to get scale-to-zero and event-driven scaling for this on Kubernetes.
🧮 Why these compound instead of add
A tempting mistake is to add the savings together. Distillation saves 75 percent, adaptive selection 35 percent, caching 30 percent, so the total must be 140 percent. That is impossible. The techniques do not each operate on the original cost. Each operates on the residual the previous stage left behind. They address different sources of waste, so they multiply:
| Stage | Technique | Operates on | Residual |
|---|---|---|---|
| 1 | Adaptive field selection | Total field count | 60 to 70 percent of baseline |
| 2 | Distillation + tiered routing | Remaining fields | 20 to 25 percent of baseline |
| 3 | Field-level caching | Reprocessed documents | 10 to 15 percent on reprocessing |
| 4 | Serverless scale-to-zero | Infrastructure | Zero idle cost |
| Result | Compound effect | End to end | 5 to 15 percent of naive baseline |
Roughly (1 - 0.35) x (1 - 0.65) x (1 - 0.25), which lands around 0.07 to 0.15 of the original, an 85 to 93 percent reduction. Each technique matters most at a different stage. Adaptive selection eliminates unnecessary work, distillation reduces the cost of necessary work, caching eliminates repeated work, and serverless eliminates idle overhead. No one of them gets you there alone.
🛡️ Making it safe: the quality-floor monitor
Aggressive cost optimization earns an obvious objection. If 70 percent of fields go to a cheaper model, you need a way to know it is not silently getting worse, or drifting as document patterns change.
Build the check in. Each week, randomly re-run 5 to 10 percent of Tier-1 fields through the frontier model as a shadow extraction and compare. This is the same baseline comparison from the start of the post, now running continuously. Agreement at or above 98 percent means the floor is holding. Below 95 percent raises an alert. Below 90 percent auto-triggers re-distillation on the last three months of data. Shadow validation adds only about 3 to 5 percent to total cost, and it means the system monitors itself, with no dedicated QA team required.
The same monitor also spots promotions. If the mid-tier and frontier models consistently agree on a field currently routed to Tier 2, that field is a candidate to fold into the next distillation cycle, so the system keeps shifting work to the cheapest tier on its own.
💡 The payoff beyond the bill
The reason this is worth doing well is not only the savings on an existing workload. Every percentage point off per-unit cost widens the set of use cases where automation is worth building at all. A workload that only penciled out for one of your five document types at naive pricing might clear the bar for three or four once the architecture is right. Cost optimization is not a late-stage cleanup. Treat it as a first-class design requirement alongside accuracy, latency, and reliability, and more of what you want to automate becomes viable.
Five moves, five kinds of waste:
- Expensive-model baseline, then prove the cheap model turns every downgrade into a measured decision.
- Adaptive field selection removes unnecessary work.
- Tiered routing + distillation reduces the cost of necessary work.
- Field-level caching removes repeated work.
- Serverless scale-to-zero removes idle overhead.
Apply them in sequence, guard them with a self-checking quality floor, and you get frontier-grade extraction at a small fraction of the naive cost, without betting accuracy on it.
Two patterns I open-sourced
Two of the ideas above are available as small, dependency-light libraries under AWS Samples (MIT-0), if you want a reference to build from:
aws-samples
/
sample-textract-field-memory
Spatial field location memory for document processing pipelines. Learns field positions, validates extractions, identifies document types by layout, detects drift, and monitors template health. Zero dependencies, pure Python.
textract-field-memory
Disclaimer
This is sample code, for non-production usage. You should work with your security and legal teams to meet your organizational security, regulatory, and compliance requirements before deployment.
Note: When processing documents containing PII (e.g. SSNs), PHI, or payment data, ensure your implementation meets applicable compliance requirements (HIPAA, PCI-DSS, GDPR, etc.). This library stores only field names and bounding-box coordinates — never field values or document content — but you remain responsible for securing the underlying documents.
The Problem
Every system that processes recurring structured data starts from scratch each time. It has no memory of what it saw yesterday.
WITHOUT spatial memory
┌──────────────┐
│ Document 1 │──► OCR/Extract ──► "Employee Name" at (0.05, 0.10) ──► ✓ stored nowhere
│ Document 2 │──► OCR/Extract ──► "Employee Name" at (0.05, 0.10) ──► ✓ re-discovered
│ Document 3 │──► OCR/Extract ──► "Employee Name" at (0.05, 0.10) ──► ✓ re-discovered again
│…
aws-samples
/
sample-prompt-correction-memory
Self-improving LLM document extraction on AWS. Each human QA correction improves future extractions via few-shot self-healing and deterministic rule graduation — no retraining, no redeployment. Costs decrease over time.
sample-prompt-correction-memory
A sample project showing automated prompt self-correction for LLM document extraction on AWS. Each human correction improves future extractions — no retraining, no redeployment.
Disclaimer: This is sample code, for non-production usage. You should work with your security and legal teams to meet your organizational security, regulatory and compliance requirements before deployment. Do not use this sample with personal data, health data or other regulated data.
The Problem
Every LLM-based extraction pipeline makes the same mistakes repeatedly. Your QA team corrects an error today, but the model makes the identical error on tomorrow's document. The correction evaporates.
WITHOUT correction memory
┌──────────────┐
│ Document 1 │──► LLM Extract ──► "quarterly" for payment_terms ──► QA corrects to "30"
│ Document 2 │──► LLM Extract ──► "quarterly" for payment_terms ──► QA corrects AGAIN
│ Document 3 │──► LLM Extract ──► "quarterly" for payment_terms ──► QA corrects AGAIN
│ ... │
└──────────────┘
Same…If you are optimizing an extraction pipeline, I would be curious where your biggest cost sink turned out to be: unnecessary calls, over-powered models, reprocessing, or idle. Drop a comment.
Top comments (0)