DEV Community

Cover image for Fine-Tuning Reduced My LLM Token Usage by 66% — Here's What I Learned
Ritesh Totlani
Ritesh Totlani

Posted on AI-assisted

Fine-Tuning Reduced My LLM Token Usage by 66% — Here's What I Learned

Fine-tuning is usually discussed as a way to improve LLM accuracy.

But there is another benefit that deserves more attention:

Fine-tuning can reduce the number of tokens you send with every request.

I wanted to measure this rather than assume it.

So I built an invoice-extraction experiment comparing four approaches:

  • Instructions
  • Few-shot prompting
  • Retrieval
  • LoRA fine-tuning

The result:

1,014 tokens → 345 tokens per invoice

That's approximately 66% fewer tokens.

And the fine-tuned model was slightly more accurate.

There was, however, an important catch.


The experiment

The task was simple:

Take invoices from Germany, France and the UK and extract:

{
  "invoice_number": "...",
  "date": "YYYY-MM-DD",
  "vendor_name": "...",
  "line_items": [],
  "total": 0.0
}
Enter fullscreen mode Exit fullscreen mode

The invoices deliberately used different number formats such as:

Germany:  1.234,56
UK:       1,234.56
France:   1 234,56
Enter fullscreen mode Exit fullscreen mode

I used Qwen2.5-1.5B-Instruct locally and generated a reproducible dataset of 357 invoices.

The evaluation set contained 93 invoices:

  • 48 using layouts represented during training
  • 45 using deliberately unseen layouts

All four approaches used the same evaluation set and exact-match scoring.


Four approaches

1. Instructions

The model received detailed extraction instructions followed by the invoice.

551 input tokens

73.1% accuracy

2. Few-shot prompting

I added one worked invoice example and its expected JSON.

Accuracy improved significantly:

73.1% → 82.8%

But input size increased:

551 → 781 tokens

3. Retrieval

Instead of using the same example, I retrieved a similar training invoice and included it in the prompt.

Accuracy:

83.9%

Input:

851 tokens

So the best prompting approach required:

851 input + 163 output = 1,014 total tokens

Then came fine-tuning.


Fine-tuning

I used LoRA to teach the model the extraction behavior.

The inference prompt became:

Extract the invoice data.

<invoice>
Enter fullscreen mode Exit fullscreen mode

The training pipeline used the expected JSON as the answer and masked the prompt tokens so that the loss was calculated only on the answer.

The training data was constructed using the same prompt format that the model would receive during evaluation.

The LoRA configuration used rank 16, alpha 32 and targeted the Q/K/V/O projection layers.


The result

Approach Input Output Total Accuracy
Instructions 551 165 716 73.1%
Few-shot 781 163 944 82.8%
Retrieval 851 163 1,014 83.9%
Fine-tuning 215 130 345 86.0%

The fine-tuned model used:

669 fewer tokens per invoice

or approximately:

66% fewer tokens

The invoice itself didn't become shorter.

The saving came mainly from removing the repeated instructions and example context.


Where did the tokens go?

With prompting, every request might contain instructions such as:

Extract these fields.

Use YYYY-MM-DD for dates.

Handle German number formatting.

Handle French number formatting.

Use the line-item total, not the unit rate.

Return JSON only.

Here is an example...

Now process this invoice.
Enter fullscreen mode Exit fullscreen mode

After fine-tuning:

Extract the invoice data.

<invoice>
Enter fullscreen mode Exit fullscreen mode

The behavior has partly moved from input context into the fine-tuned model.

That's the key idea.

Fine-tuning can act as a form of prompt compression.


What does that mean at scale?

The experiment saved approximately 669 tokens per invoice.

That becomes:

Volume Approx. tokens saved
1,000 invoices 669K
30,000 invoices 20.07M
100,000 invoices 66.9M
1 million invoices 669M

These aren't direct cost savings—the actual amount depends on model pricing, input/output rates, caching and infrastructure.

But the principle is important:

Repeated prompt context becomes a recurring cost at scale.


And accuracy improved too

The token reduction wasn't achieved by sacrificing overall accuracy.

The best prompting approach achieved:

83.9%

The fine-tuned model achieved:

86.0%

It also produced clean JSON for all 93 evaluation invoices.

So in this experiment, fine-tuning gave us:

  • 66% fewer total tokens
  • higher overall accuracy
  • more consistent structured output

Sounds like an easy decision, right?

Not quite.


The catch: unseen layouts

The training data contained eight invoice layouts.

I deliberately held three layouts out of training.

The fine-tuned model achieved:

95.8% on seen layouts

but only:

75.6% on unseen layouts

The few-shot approach was much more stable:

83.3% seen

82.2% unseen

So fine-tuning was excellent when the invoice looked familiar, but less robust when the structure changed.

One unseen invoice contained running subtotals.

The correct line-item amount was:

5743.80
Enter fullscreen mode Exit fullscreen mode

But the model returned:

11021.02
Enter fullscreen mode Exit fullscreen mode

because that was a running subtotal created from multiple line items.

Another unseen layout used a horizontal two-column structure, causing the model to associate a line-item description with the invoice-number field.

These failures suggest that fine-tuning had learned useful structural patterns from the training layouts—but some of those assumptions didn't transfer when the layout changed.


So when does fine-tuning make sense?

This experiment changed how I think about fine-tuning.

The question isn't only:

"Will fine-tuning improve accuracy?"

I'd also ask:

"How many tokens am I repeatedly sending to describe a task that rarely changes?"

Fine-tuning becomes particularly attractive when you have:

High volume
Thousands or millions of repeated requests.

Stable behavior
The extraction rules don't change frequently.

Stable input formats
Your documents mostly resemble the training distribution.

If those conditions hold, moving repeated instructions into the model can significantly reduce inference tokens.

If your document formats change constantly, prompting may be safer.


The architecture I'd consider

For a production invoice system, I wouldn't necessarily choose one approach for everything.

A hybrid design could look like:

                 Incoming Invoice
                        │
                        ▼
              Familiar format?
                 /          \
               Yes           No
                │             │
                ▼             ▼
         Fine-tuned       Prompted
            model           model
                │             │
                └──────┬──────┘
                       ▼
                 Structured JSON
Enter fullscreen mode Exit fullscreen mode

The fine-tuned model handles high-volume, familiar cases with a small prompt.

The prompted path handles unusual or newly discovered formats.

This is a proposed architecture based on the experiment, not something measured in the experiment itself.


Final takeaway

Fine-tuning isn't only an accuracy technique.

It can also be a token optimization technique.

In this experiment:

1,014 → 345 total tokens

83.9% → 86.0% accuracy

But unseen-layout accuracy dropped to 75.6%.

So the lesson isn't:

"Fine-tuning is always better."

It's:

Fine-tune when the task and input distribution are stable. Prompt when flexibility matters more.

And if you're repeatedly sending the same instructions to an LLM, ask yourself:

How much am I paying to tell the model something I've already told it thousands of times?

Sometimes the best way to shorten your prompt isn't to rewrite it.

It's to teach the model what you keep telling it.

Top comments (0)