DEV Community

Avneet bansal
Avneet bansal

Posted on

Your LLM Keeps Making the Same Extraction Mistake. Here's How to Make It Learn.

 > TL;DR — I open-sourced an AWS sample that makes LLM document extraction learn from human corrections. Each QA fix is captured once and reused, so the pipeline gets more accurate and cheaper over time — with no retraining and no redeployment. Repo: github.com/aws-samples/sample-prompt-correction-memory. You can run the whole self-healing loop locally in 30 seconds, no AWS account needed.

The problem: corrections that evaporate

If you've built an LLM extraction pipeline, you know this pain.

Your QA analyst opens today's invoice. The model extracted payment_terms as "quarterly." That's wrong — "quarterly" is the billing frequency; the actual payment term is Net 30. The analyst corrects it. Done.

Tomorrow, an identical invoice from the same vendor arrives. The model extracts... "quarterly" again. Same mistake. The analyst corrects it again.

The correction evaporated. It fixed one document and taught the system nothing.

Multiply that across thousands of documents and dozens of recurring error patterns, and you get the two options most teams settle for:

  1. Pay analysts to re-fix the same patterns forever (QA cost never goes down), or
  2. Fine-tune a model (expensive, slow, needs ML ops, and you redo it every time the patterns shift).

There's a third way that needs neither.

The idea: treat corrections as durable memory

What if every human correction became a permanent, reusable asset?

That's the pattern behind sample-prompt-correction-memory. Instead of throwing corrections away, it stores them and feeds them back into future extractions — so the same mistake is never made twice.

It works as three tiers, and the magic is that traffic gradually shifts from the expensive tier to the free one:

🟢 Tier 1 — Graduated rules ($0, no LLM call)

When the same correction pattern recurs enough times, the system synthesizes a deterministic rule (a regex, a lookup, a normalization) that handles that field with zero LLM cost and sub-millisecond latency. "Net 30" → 30 becomes a rule; it never needs a model again.

🔵 Tier 2 — Cloud LLM extraction (confidence-gated)

Fields with no matching rule go to Amazon Bedrock (Claude Haiku 4.5). Every field comes back with a confidence score. If confidence clears the threshold, accept it and move on.

🟠 Tier 3 — Self-healing (only when needed)

If confidence is below threshold, the system retrieves the most relevant past corrections, formats them as few-shot examples, and re-extracts with a stronger model (Claude Sonnet 4.5). The model learns from the specific mistakes it made before — in-context, no training.

Every QA correction flows into a correction log (DynamoDB) that powers both the self-healing retrieval and the rule graduation. So the more the system is used, the smarter and cheaper it gets.

Why this matters (the practical payoff)

  • No ML ops. No retraining, no GPUs, no deployment cycles. A correction takes effect on the very next document.
  • Cost drops as quality rises. Normally you trade one for the other. Here, recurring fixes graduate to $0 rules — so accuracy climbs while per-document cost falls.
  • QA effort compounds instead of repeating. Fix it once → fixed forever.
  • It's a QA action, not an ML project. Improving the pipeline no longer requires the ML team. An analyst uploads a correction; the system improves.

The clearest fit is high-repetition, back-office document processing — accounts payable, contracts, compliance filings, supply-chain docs — where the same field errors recur and analyst time is the real expense.

The architecture

Fully serverless, deployed with AWS SAM:

Document ──► S3 ──► EventBridge ──► SQS ──► Lambda (extract)
                                              │
                          ┌───────────────────┼───────────────────┐
                          ▼                    ▼                   ▼
                   Tier 1: Rules       Tier 2: Bedrock      Tier 3: Self-heal
                   (deterministic)     (Claude Haiku 4.5)   (Claude Sonnet 4.5)
                          ▲                                        │
                          │                                        ▼
                          └──────── Correction Log (DynamoDB) ◄────┘
                                    (grows from QA feedback)
Enter fullscreen mode Exit fullscreen mode

QA corrections are uploaded as JSON to an S3 corrections/ prefix; a second Lambda validates and ingests them into the correction log. Everything is encrypted with a customer-managed KMS key, and IAM is scoped to the specific Bedrock model ARNs the sample uses.

See it in 30 seconds (no AWS account)

The quickstart runs the entire self-healing loop with a mocked Bedrock client, so you can watch the mechanism without deploying anything:

git clone https://github.com/aws-samples/sample-prompt-correction-memory
cd sample-prompt-correction-memory
pip install -e ".[dev]"
python examples/quickstart.py
Enter fullscreen mode Exit fullscreen mode

You'll see output like this:

Step 1: Initial extraction (no correction memory)
  Field:      effective_date
  Value:      March 2024        Confidence: 0.55   ❌ NO

Step 2: Self-healing triggers (0.55 < threshold 0.70)
  Retrieving corrections for 'effective_date'...
  Found: "March 2024" → "2024-03-01"

Step 3: Re-extraction with correction memory
  Value:       2024-03-01       Confidence: 0.95   ✓ YES
  Self-Healed: True
Enter fullscreen mode Exit fullscreen mode

Low-confidence extraction → retrieve past correction → re-extract → correct answer. No retraining. No redeployment.

Deploying to AWS

With the AWS CLI, SAM CLI, and Bedrock model access for Claude Haiku 4.5 + Sonnet 4.5:

make deploy    # S3, DynamoDB, Lambda, EventBridge, SQS, KMS
make seed      # load sample corrections into the correction log
make trigger   # upload a sample document → triggers real extraction
make verify    # read back results (fields, confidence, self_healed)
make destroy   # empty buckets + delete the stack
Enter fullscreen mode Exit fullscreen mode

make verify prints each extracted field with its confidence and whether self-healing kicked in — your proof it works end-to-end.

A companion: spatial memory

This sample provides semantic correction memory (what was extracted wrong, and why). Its companion, sample-textract-field-memory, provides spatial memory (where fields appear on a document layout). Together they form a dual memory for document pipelines: use the cheap spatial lookup when you're confident where a field is, and fall back to self-healing extraction when you're not.

Honest limitations

Because it's more useful when it's used well:

  • It's sample/reference code, not a production library. Harden IAM, add monitoring, and put a human in the loop for high-stakes fields before real use.
  • It helps most where the same field errors recur. For wildly heterogeneous, one-off documents, there's less pattern to learn.
  • Graduated rules currently take effect without a human approval step — in production you'd add a review gate before a synthesized rule starts answering on its own.
  • Don't use it with regulated data (PII/PHI/PCI) as-is; that path needs a proper compliance review.

Try it and tell me

The repo is open source under aws-samples, MIT-0 licensed, with a benchmark suite, an interactive dashboard, and an offline test harness.

👉 github.com/aws-samples/sample-prompt-correction-memory

What repetitive extraction error is your team still fixing by hand? I'd love to hear which patterns you'd want a system like this to learn.


This is an open-source AWS sample intended for demonstration and non-production use.

Top comments (1)

Collapse
 
hannune profile image
Tae Kim •

The rule graduation idea solved a real headache for me on entity canonicalization, but we ran into trouble when the same string mapped to different correct values depending on context. "quarterly" meaning Net 30 for one vendor but Q3 billing for another meant naive pattern promotion kept breaking other records. We ended up tagging corrections with a context key before graduation, which wasn't in any writeup I'd found. Worth thinking about before the rule count grows.