"Should we fine-tune a model on our documents?"
It's one of the most common questions I hear from teams starting with AI, and the answer is usually no.
The short answer
Use RAG (retrieval-augmented generation) when your AI needs to answer from your own, changing information: documents, policies, product data.
Use fine-tuning when you need the model to behave differently: follow a strict format, tone or narrow task very consistently.
For most business assistants, start with RAG.
What is RAG?
RAG connects a model to a search system over your content. When someone asks a question, the system first finds the most relevant passages from your documents, then gives them to the model along with the question. The model answers from those passages, and can cite them.
Think of it as an open-book exam: the model doesn't memorise your handbook, it looks up the right page every time. Update the handbook and answers change immediately, with no retraining.
What is fine-tuning?
Fine-tuning continues training a model on your own examples so it learns a pattern: a writing style, a classification scheme, a structured output format.
It changes behaviour, but it's not a reliable way to teach facts that change, and it can't show where an answer came from.
Think of it as training a new employee on how your team writes. Useful, but you'd still hand them the current policy document.
Side by side
| RAG | Fine-tuning | |
|---|---|---|
| Best for | Answering from your documents and data | Consistent style, format or narrow task |
| Keeping info current | Update the documents, done | New training data and retraining |
| Shows sources | Yes, can cite passages | No |
| Access control | Can filter by each user's permissions | Anything in training data may surface to anyone |
| What you need | Clean, organised content + good search | Hundreds to thousands of quality examples |
4 questions to choose
- Does the answer depend on information that changes? Prices, policies, specs, case files → RAG.
- Do users need to see where an answer came from? Compliance, support, legal → RAG.
- Is the problem how the model responds, not what it knows? Strict format, brand voice, specialised classification → consider fine-tuning.
- Have you tried good prompts and examples first? Clear instructions and a few worked examples solve many "behaviour" problems with no training at all.
⚠️ The most common mistake
Fine-tuning a model on company documents so it "knows the business."
The model may pick up the style but still get facts wrong, can't cite its sources, and needs retraining every time the documents change. For knowledge, use retrieval.
What makes RAG actually work
- Good content: remove outdated and duplicate documents. Answers are only as good as the sources.
- Smart search: split documents into meaningful sections and combine keyword + semantic search, so exact terms like product codes are found.
- Permissions: filter results by what each user is allowed to see.
- Evaluation: test with real questions, checking both retrieval and the final answer.
- Honest fallbacks: when nothing relevant is found, the assistant should say so instead of guessing.
When to use both
Some systems combine them: RAG supplies the facts, while a fine-tuned model formats answers exactly as required or handles a narrow task more cheaply at high volume. Add fine-tuning only when you can measure the improvement against a test set.
Which approach are you using today, and what's been the hardest part? I'd love to hear in the comments. 👇
Originally published on ITACC Insights.
@ITACC we build production AI, RAG assistants and data pipelines.
Top comments (1)
The distinction between changing information and changing behaviour is useful, especially when the two are evaluated separately. For the strict-format case, I would first establish whether schema-constrained decoding and application validation already solve the failure, before investing in training examples.
A valid JSON object and a correct task result are separate outcomes. Fine-tuning can improve one while leaving the other unchanged, so the test set should track both, plus refusal on inputs outside the intended task. When retrieval is added later, keep some examples where the supplied evidence conflicts with familiar training patterns; that tests whether the adapted model still follows current evidence.