Fine-tuning is usually discussed as a way to improve LLM accuracy.
But there is another benefit that deserves more attention:
Fine-tuning can reduce the number of tokens you send with every request.
I wanted to measure this rather than assume it.
So I built an invoice-extraction experiment comparing four approaches:
- Instructions
- Few-shot prompting
- Retrieval
- LoRA fine-tuning
The result:
1,014 tokens → 345 tokens per invoice
That's approximately 66% fewer tokens.
And the fine-tuned model was slightly more accurate.
There was, however, an important catch.
The experiment
The task was simple:
Take invoices from Germany, France and the UK and extract:
{
"invoice_number": "...",
"date": "YYYY-MM-DD",
"vendor_name": "...",
"line_items": [],
"total": 0.0
}
The invoices deliberately used different number formats such as:
Germany: 1.234,56
UK: 1,234.56
France: 1 234,56
I used Qwen2.5-1.5B-Instruct locally and generated a reproducible dataset of 357 invoices.
The evaluation set contained 93 invoices:
- 48 using layouts represented during training
- 45 using deliberately unseen layouts
All four approaches used the same evaluation set and exact-match scoring.
Four approaches
1. Instructions
The model received detailed extraction instructions followed by the invoice.
551 input tokens
73.1% accuracy
2. Few-shot prompting
I added one worked invoice example and its expected JSON.
Accuracy improved significantly:
73.1% → 82.8%
But input size increased:
551 → 781 tokens
3. Retrieval
Instead of using the same example, I retrieved a similar training invoice and included it in the prompt.
Accuracy:
83.9%
Input:
851 tokens
So the best prompting approach required:
851 input + 163 output = 1,014 total tokens
Then came fine-tuning.
Fine-tuning
I used LoRA to teach the model the extraction behavior.
The inference prompt became:
Extract the invoice data.
<invoice>
The training pipeline used the expected JSON as the answer and masked the prompt tokens so that the loss was calculated only on the answer.
The training data was constructed using the same prompt format that the model would receive during evaluation.
The LoRA configuration used rank 16, alpha 32 and targeted the Q/K/V/O projection layers.
The result
| Approach | Input | Output | Total | Accuracy |
|---|---|---|---|---|
| Instructions | 551 | 165 | 716 | 73.1% |
| Few-shot | 781 | 163 | 944 | 82.8% |
| Retrieval | 851 | 163 | 1,014 | 83.9% |
| Fine-tuning | 215 | 130 | 345 | 86.0% |
The fine-tuned model used:
669 fewer tokens per invoice
or approximately:
66% fewer tokens
The invoice itself didn't become shorter.
The saving came mainly from removing the repeated instructions and example context.
Where did the tokens go?
With prompting, every request might contain instructions such as:
Extract these fields.
Use YYYY-MM-DD for dates.
Handle German number formatting.
Handle French number formatting.
Use the line-item total, not the unit rate.
Return JSON only.
Here is an example...
Now process this invoice.
After fine-tuning:
Extract the invoice data.
<invoice>
The behavior has partly moved from input context into the fine-tuned model.
That's the key idea.
Fine-tuning can act as a form of prompt compression.
What does that mean at scale?
The experiment saved approximately 669 tokens per invoice.
That becomes:
| Volume | Approx. tokens saved |
|---|---|
| 1,000 invoices | 669K |
| 30,000 invoices | 20.07M |
| 100,000 invoices | 66.9M |
| 1 million invoices | 669M |
These aren't direct cost savings—the actual amount depends on model pricing, input/output rates, caching and infrastructure.
But the principle is important:
Repeated prompt context becomes a recurring cost at scale.
And accuracy improved too
The token reduction wasn't achieved by sacrificing overall accuracy.
The best prompting approach achieved:
83.9%
The fine-tuned model achieved:
86.0%
It also produced clean JSON for all 93 evaluation invoices.
So in this experiment, fine-tuning gave us:
- 66% fewer total tokens
- higher overall accuracy
- more consistent structured output
Sounds like an easy decision, right?
Not quite.
The catch: unseen layouts
The training data contained eight invoice layouts.
I deliberately held three layouts out of training.
The fine-tuned model achieved:
95.8% on seen layouts
but only:
75.6% on unseen layouts
The few-shot approach was much more stable:
83.3% seen
82.2% unseen
So fine-tuning was excellent when the invoice looked familiar, but less robust when the structure changed.
One unseen invoice contained running subtotals.
The correct line-item amount was:
5743.80
But the model returned:
11021.02
because that was a running subtotal created from multiple line items.
Another unseen layout used a horizontal two-column structure, causing the model to associate a line-item description with the invoice-number field.
These failures suggest that fine-tuning had learned useful structural patterns from the training layouts—but some of those assumptions didn't transfer when the layout changed.
So when does fine-tuning make sense?
This experiment changed how I think about fine-tuning.
The question isn't only:
"Will fine-tuning improve accuracy?"
I'd also ask:
"How many tokens am I repeatedly sending to describe a task that rarely changes?"
Fine-tuning becomes particularly attractive when you have:
High volume
Thousands or millions of repeated requests.
Stable behavior
The extraction rules don't change frequently.
Stable input formats
Your documents mostly resemble the training distribution.
If those conditions hold, moving repeated instructions into the model can significantly reduce inference tokens.
If your document formats change constantly, prompting may be safer.
The architecture I'd consider
For a production invoice system, I wouldn't necessarily choose one approach for everything.
A hybrid design could look like:
Incoming Invoice
│
▼
Familiar format?
/ \
Yes No
│ │
▼ ▼
Fine-tuned Prompted
model model
│ │
└──────┬──────┘
▼
Structured JSON
The fine-tuned model handles high-volume, familiar cases with a small prompt.
The prompted path handles unusual or newly discovered formats.
This is a proposed architecture based on the experiment, not something measured in the experiment itself.
Final takeaway
Fine-tuning isn't only an accuracy technique.
It can also be a token optimization technique.
In this experiment:
1,014 → 345 total tokens
83.9% → 86.0% accuracy
But unseen-layout accuracy dropped to 75.6%.
So the lesson isn't:
"Fine-tuning is always better."
It's:
Fine-tune when the task and input distribution are stable. Prompt when flexibility matters more.
And if you're repeatedly sending the same instructions to an LLM, ask yourself:
How much am I paying to tell the model something I've already told it thousands of times?
Sometimes the best way to shorten your prompt isn't to rewrite it.
It's to teach the model what you keep telling it.
Top comments (0)