A platform team is evaluating an LLM-based assistant, still in testing, meant to answer developers' architecture questions using the company's ADRs (Architecture Decision Records) and internal documentation. During an evaluation round, the team itself notices that the assistant recommended a service-to-service authentication pattern the company abandoned last year and, on top of that, suggested calling an endpoint on an internal API already marked as deprecated. The answer was plausible, well written, and technically defensible in general terms. It just contradicted a standing decision.
The failure was caught before it reached any user, and that is exactly why the conversation that follows matters. Two proposals land on the table. The first: "We need to fine-tune on our ADRs, so the model learns our decisions." The second: "We need RAG over the ADRs, so it can look up the standing decision." Both are reasonable, and each starts from a theory about what the model is missing: one assumes it needs to learn the content, the other that it needs to consult the content. But there is a prior question that neither one answers: what, exactly, went wrong in that answer?
That question matters more than it seems, because the same outdated recommendation can have very different origins. The ADR that superseded the old pattern may never have been indexed. The old ADR, marked "superseded," may have been retrieved instead of the new one. Both may have been retrieved and the model ignored the context, answering from what it learned about the pattern in public documentation. Or the question may have been generic and the system instruction never asked it to prioritize the standing decision. Each cause calls for a different fix, and none of them is resolved, on its own, by whichever technique won the argument. Fine-tuning on today's ADRs, for example, does not help when the decision is superseded tomorrow. This is not unique to our scenario: an engineering report on RAG systems identified seven distinct failure points when building this kind of system.
To understand how we got to this debate, it helps to recall the context. Since ChatGPT arrived in late 2022, organizations have been trying to put LLMs in front of what used to live in wikis, shared folders, and operational documents, with the promise of making that knowledge more accessible to whoever needs it. Along the way, two techniques that were research topics became everyday vocabulary: retrieval-augmented generation (RAG), proposed by Lewis et al. in 2020, and fine-tuning, which parameter-efficient methods like LoRA made far cheaper to run. When two powerful tools arrive together and address the same apparent "pain," it is natural for the market to put them head to head. That is how the standard framing was born: RAG or fine-tuning?
The problem with that framing is that it asks you to pick the answer before agreeing on the question. There is research comparing the two approaches side by side, and it deserves a careful reading, not a slogan. OpenAI's guide on optimizing accuracy, for instance, treats them as answers to different problems and as approaches that are "additive, not exclusive." The question, therefore, should not start from the technique. It should start from the failure.

Figure 1. The difference between the two conversations is not the technique chosen, but the step that was skipped.
This article argues for that inversion and proposes a practical way to make it, separating what a system needs to know, how it needs to behave, and what it needs to do. It is the first step in a series about the decisions that happen beyond the model, where an AI system either becomes reliable or stops being so.
1. The problem
Back to the architecture assistant. The two proposals on the table have something in common: both start from a technique and only then look for the problem it would solve. It is a subtle inversion, because the conversation sounds technical and mature. But engineering's usual path is the opposite: observe the symptom, form a hypothesis about the cause, test it, and only then intervene. In LLM-based systems, diagnosis tends to be the easiest step to skip.
Why is it skipped? From my own observation, we can think of several reasons, but three seem the most plausible. First, techniques are tangible: they have tools, tutorials, budget, and a name that fits on a roadmap. Second, diagnosing is work: it requires opening the traces, separating what was retrieved from what was generated, reproducing the failure, and understanding at which stage it originated. Third, "RAG or fine-tuning?" is a question you can answer with an opinion, while "what caused this failure?" can only be answered with evidence.
The cost of choosing the technique before the diagnosis shows up in at least three ways.
The original failure is still there. If the problem was the old ADR being retrieved instead of the new one, training the model on the ADRs changes nothing: the wrong content still reaches the context. If the problem was the model ignoring the retrieved context, retrieving more context does not solve it either. The intervention ships, the team considers the problem solved, and the outdated recommendation shows up again in the next evaluation round.
The intervention can make the system worse. OpenAI's own guide on optimizing accuracy includes an example where RAG "confused" the model by adding noise and lowered the score by four points. On the fine-tuning side, a controlled study of question answering without external lookup showed that examples containing new knowledge are learned more slowly and, once learned, increase the model's tendency to hallucinate. That result holds for that specific setting and should not be extended to all fine-tuning. Still, it shows that "one more technique" is not a risk-free bet.
The complexity becomes yours. Each technique brings a new maintenance surface. RAG requires indexing, updating, ranking, and permission control. Fine-tuning requires training data, versioning, re-evaluation on every base-model change, and a plan for when the content changes. If the real cause was an ambiguous instruction, the team paid that cost to solve something a prompt adjustment would have fixed. Not by coincidence, that same OpenAI guide recommends starting with prompt engineering.
None of this means either technique is the villain. The evidence is mixed, and that is part of the problem. In direct comparisons of knowledge injection, RAG outperformed unsupervised fine-tuning, both for previously seen knowledge and for new knowledge (Ovadia et al., 2023). A more recent study, on multi-hop questions over novel knowledge, found the highest overall accuracy with supervised fine-tuning. Both are preprints, with small or synthetic benchmarks, and neither answers our team's question. They show that each technique has its own territory, and that finding out which territory the failure sits in is precisely the work the question "RAG or fine-tuning?" skips. We return to these studies, in more detail, in section 5.
2. The engineering hypothesis
If the problem is starting from the technique, the alternative needs to be something testable, not just a preference for a method. I propose the following working hypothesis:
H1. A large share of the observable failures in an assistant like ours can be attributed to one of three categories, knowledge, behavior, or execution, and each category responds better to a different intervention.
This is a hypothesis, not a result. It was not tested for this article. To test it, a team could collect the failures from an evaluation round, ask independent reviewers to classify each one, and measure the agreement between them. Next, they would apply the intervention indicated for each category and compare the outcome against the alternative. The hypothesis weakens if inter-reviewer agreement is low, if most failures do not fit the three categories, or if all of them call for the same intervention. Upcoming articles in the series return to this experimental design.
There is, however, a refinement our scenario suggests, and one that usually stays out of the discussion. Back to the superseded ADR. In his original text on the format, Michael Nygard recommends that reversed decisions not be deleted: they stay on record and are marked "superseded," because "it is still relevant to know that it was the decision, but is no longer the decision" (Documenting Architecture Decisions). In other words, a healthy ADR repository contains, by design, obsolete knowledge. What distinguishes the standing from the obsolete is the status, the date, and the relationship between records, not the text itself. A model that only sees the text has no way to know.
That leads to a second hypothesis, complementary to the first:
H2. Part of the failures classified as "knowledge" are not a lack of knowledge. They are curation failures: the right information exists, but the way the corpus is organized does not allow telling the standing from the obsolete, and no technique applied to the model fixes that.
If H2 is true, the first question in the face of a knowledge failure stops being "RAG or fine-tuning?" and becomes "is the corpus in a condition to be consulted?". Curation, in this sense, is a set of engineering decisions about knowledge, prior to any decision about the model. Four forms, in increasing order of cost:
| Form of curation | What it addresses in our scenario | Cost and risk |
|---|---|---|
| Metadata and lifecycle | Status (proposed, accepted, deprecated, superseded), date, owner, and a reference to the record that supersedes it; filtering on that information at retrieval time | Low. Requires discipline in filling and updating |
| Canonical source and deduplication | Guarantees a single source of truth per decision, instead of diverging copies across wikis and folders | Medium. Requires an owner and a review process |
| Explicit relationships | Records that ADR 31 supersedes ADR 12 and that a pattern applies to a set of services | Medium to high. Requires modelling the domain |
| Ontology and knowledge graph | Represents concepts (decision, pattern, API, service) and relationships in a queryable, verifiable way | High. Requires modelling, construction, and continuous maintenance |
The last two items deserve a comment, because they enter ontology territory. Gruber's classic definition describes an ontology as a formal specification of a conceptualization, that is, an explicit agreement about which concepts exist in a domain and how they relate. In our case, that would mean declaring that "Decision," "Pattern," "API," and "Service" are kinds of thing, and that "supersedes," "depends on," and "applies to" are relationships between them. With that, the question "what is the standing decision for service-to-service authentication?" stops being a text-similarity search and becomes a query with a verifiable answer.
The literature on combining LLMs with knowledge graphs points to both the potential and the cost. Pan et al. argue that knowledge graphs store factual knowledge explicitly and can help LLMs with external information and interpretability, and that the two are complementary. The same authors note that these graphs are hard to build and are constantly evolving. On the practical side, GraphRAG proposes using an LLM to derive an entity graph from the documents and, with it, answer questions about the entire corpus, which conventional RAG struggles with. The authors report improvements in comprehensiveness and diversity of answers on datasets of around one million tokens. One caveat: GraphRAG was designed for global summarization questions, not for the problem of resolving which decision is standing. It shows that giving structure to the corpus can help, but it does not demonstrate that it solves our case. That remains a hypothesis to test.
In short, the article's question gains two layers. The first is H1's: which category is the failure in? The second is H2's: if it is knowledge, is the corpus curated well enough to be used? Only after those two answers does it make sense to discuss RAG, fine-tuning, or tools.
3. Three axes: knowledge, behavior, task execution
If H1 is to be useful, it needs categories a team can apply to a real failure without relying on interpretation. I propose three axes, each defined by a question you can ask of the system.

Figure 2. Each kind of failure points to a different capability.
Knowledge: does the model have access to the right information?
This is the axis of the outdated recommendation. The typical symptom is an answer citing a decision that does not exist, has expired, or has been superseded. The diagnostic test is simple: if we manually place the standing ADR in the context, does the model get it right? If it does, the problem is not the model's capability but access to information, and the conversation shifts to retrieval and curation (H2). This is the territory OpenAI's guide assigns to context optimization: the case where the model lacks knowledge because it was not in the training data. The natural intervention is RAG, with the caveat that it has several failure points of its own, from indexing to generation.
Behavior: given the information, does the model use it the expected way?
Here the model has the right ADR in context and still answers poorly. It does not follow the format the team requires, does not cite the decision's status, mixes recommendation with opinion, or loses the expected structure in long answers. The diagnostic test changes: with correct context and a clear instruction, does the formatting error persist consistently? If it does, and only after exhausting instructions and examples, fine-tuning becomes a candidate. That same OpenAI guide associates it with inconsistent or incorrectly formatted results. There is conceptual support for this split: Gekhman et al. conclude that models acquire most factual knowledge during pre-training, and that fine-tuning mainly teaches them to use it better. RAFT shows a concrete use: training the model to ignore irrelevant retrieved documents and cite the relevant passages, that is, fine-tuning shaping how the model uses retrieval, not what it knows.
Execution: does the task require querying a system or acting?
Some questions should not be answered by text, neither from the model nor from a document. If the question is "is this API deprecated today?", the source of truth is the API catalog, not an ADR written months ago. The symptom is an answer that mixes model memory or outdated text with a fact that lives in a queryable system. The intervention is a tool: the model calls the catalog, receives the current status, and answers based on it. Work like ReAct, which combines reasoning and acting, and Toolformer, where the model learns when and how to call APIs such as a calculator and search, shows that tools address what the model cannot maintain on its own. There is a porous boundary with the knowledge axis, because querying a catalog is also access to information. The practical difference is the source: a document corpus that needs curating, on one side, and an authoritative, queryable, current system on the other. Tools also carry their own risks. The MCP specification, for example, treats tools as arbitrary code execution and advises treating tool behavior descriptions as untrusted unless they come from a trusted source.
What falls outside the three axes
The three axes do not cover everything, and it is important to say so. An ambiguous instruction and insufficient reasoning also produce wrong answers, and none of the three capabilities fixes that. That is why the decision tree in section 7 begins with a step zero, before any technique: improve the task specification, decompose it, or use a more capable model. This is also consistent with OpenAI's advice to start with prompt engineering. This limitation is one of the reasons H1 speaks of "a large share of failures," not all of them.
There is also a point about combinations. The axes describe the cause of a failure, not an exclusive choice of solution. RAFT is an example of fine-tuning (behavior) working in service of retrieval (knowledge), and that same OpenAI guide describes the approaches as additive, not exclusive. The gain of thinking in axes is knowing which lever to pull first and how to tell whether it worked.
4. The architecture

Figure 3. The model is one component. Knowledge, behavior, and actions have different homes in the system.
The three axes from the previous section take on practical meaning once we place them in an architecture. Figure 3 shows the ADR assistant from our scenario and indicates where each capability lives. The central idea is that the model is one component among several, and that each axis corresponds to a different place in the system where you can intervene.
Follow the path of a question like "which authentication pattern should we use between services?".
- Retrieval. The question passes through a block that searches the documents and applies filters, for example discarding ADRs with superseded status. This is where the knowledge axis manifests at runtime.
- Knowledge base. Behind retrieval sits the corpus, and that is where section 2's curation happens: status, relationships, canonical source. Optimal retrieval over a poorly curated corpus still returns outdated answers.
- Model. Receives the question and the context and generates the answer. The behavior axis lives here: instructions, examples, and, when justified, a model adapted by fine-tuning.
- Tools. Instead of trusting text, the model can query an authoritative system, such as the API catalog, and come back with the current fact. This is the execution axis.
- Answer. The result returns to the user, ideally with the source and the status of the decisions used.
This reading is consistent with a trend described in the AI engineering literature. The Berkeley group, in a post describing "compound AI systems", argues that state-of-the-art results increasingly come from systems with multiple components, such as model calls, retrievers, and tools, rather than from an isolated model, and that iterating on the system tends to be faster than retraining the model. Anthropic's guide on agents starts from the same idea and recommends seeking the simplest possible solution, noting that many applications need only a single well-optimized model call. The architecture in Figure 3 is therefore a ceiling of complexity, not a goal. Not every system needs every block, and starting without them is a legitimate choice.
Two observations about the diagram. The first is that the blocks have different owners and lifecycles. The corpus changes when a decision changes, the model changes when the provider changes or when there is retraining, and tools change with the systems they query. Mixing these layers into one, for example by training the model on content that should live in the corpus, is the kind of decision section 1 called "complexity that becomes yours." The second is that the figure deliberately omits the layer that cuts across every block: evaluation and observability. Without recording what was retrieved, what the model received, and which tool was called, the diagnosis from the previous sections is not possible, and we go back to choosing techniques by opinion. That topic deserves its own article in the series.
5. Evidence
Before proposing a decision tree, it is worth asking what research has already shown. The short answer is that useful evidence exists, but it is more limited than the debates make it seem. The studies below compare the two approaches for injecting knowledge. This section is based on the studies' abstracts, not the full papers, so I only cite results the abstracts themselves state.
| Study | Setting | What the authors report |
|---|---|---|
| Ovadia et al. (2023) | Knowledge-intensive tasks, with previously seen and new knowledge | Unsupervised fine-tuning improves things somewhat, but RAG is better in both cases. Models struggle to learn new facts this way |
| Soudani et al. (2024) | Synthetic question answering over low-popularity knowledge, across twelve models | Fine-tuning improves at every popularity level, but RAG beats it by a wide margin, especially on the least popular facts |
| Balaguer et al. (2024), Microsoft | Agriculture case study, with Llama2-13B, GPT-3.5, and GPT-4 | Fine-tuning raised accuracy by more than 6 percentage points, and RAG added about 5 points on top. The effects accumulated |
| Zhang et al. (2024), RAFT | Domain-specific adaptation with RAG | Training the model to ignore irrelevant documents and cite the relevant ones brought consistent gains across three datasets |
| Mecklenburg et al. (2024), Microsoft | Recent sports events | Supervised fine-tuning can inject facts, provided the training data is well constructed. Coverage was uneven with token-based scaling and more uniform with fact-based scaling |
| Yang et al. (2026) | Multi-hop questions over novel knowledge, across three 7B models | RAG gave consistent gains, but supervised fine-tuning had the highest overall accuracy |
| Abonizio et al. (2025) | Knowledge injection in a low-resource regime | RAG-based injection tended to degrade performance on control sets more than parametric methods did |
What can be stated carefully. For new or rare facts, unsupervised fine-tuning over raw text is a weak path, and RAG was superior in every direct comparison I read. At the same time, supervised fine-tuning with well-constructed data can add knowledge and, in a 2026 study, had the best overall accuracy. And the effects can add up, as in the agriculture case. The most honest reading is that the techniques act in different places, which supports H1, not that one beats the other.
What the evidence does not say. There are four limits the reader should keep in mind:
- Quality of the evidence. Almost all the studies above are preprints, and most use small or synthetic benchmarks, such as agriculture, sports events, and invented entities. There is no reason to assume the results transfer to a corporate corpus.
- None of them tests our problem. The studies measure whether the model learns or retrieves facts. None measures the distinction between a standing and an obsolete decision, which is the center of this article's scenario, and therefore none tests H2.
- RAG is not always better. Besides the Abonizio study, the example from OpenAI's guide shows a case where RAG lowered the score through noise. The gain depends on retrieval bringing in the right content.
- Fine-tuning effects depend on the setting. The increase in hallucination reported by Gekhman et al. was measured on question answering without external lookup, with new knowledge. It does not extend to fine-tuning for format or style.
In short, the research supports the thesis that the two approaches do not compete on the same ground, but it does not replace a test on your own corpus and with your real failures. That is why one of the upcoming articles in the series covers the experiment.
6. The trade-offs
Choosing a capability means accepting a set of costs that do not appear on the proposal slide. The table below compares the three capabilities across six dimensions, applied to the ADR assistant. It is a qualitative comparison: this article did not measure latency or cost, and I leave those measurements for a future article in the series rather than repeating numbers of dubious origin.
| Dimension | Retrieval (RAG) | Fine-tuning | Tools |
|---|---|---|---|
| Updating | Reindex what changed. When an ADR is superseded, updating the corpus is enough | Retrain to reflect the change, and re-evaluate | Immediate: the tool queries the current system |
| Recurring cost | Indexing infrastructure and more context text on every query | Training data, training, evaluation, and a new round on every base-model change | Integration and maintenance of each tool and its contract |
| Latency | Adds at least one search step before generation | Generally adds no steps at inference | Adds round trips on every call |
| Auditability | Lets you point to the source, but the citation must be verified | Knowledge diffused in the weights, hard to trace or remove | Calls are loggable and reproducible |
| Security | Permissions and cross-context leakage must be handled at retrieval | Training data becomes part of the model | Tools are code execution and require trust and limits |
| Maintenance | Continuous curation of the corpus (H2) | Coupling to the base model and the provider's lifecycle | Versioning and compatibility of the integrations |
Some points in the table deserve an explanation.
Updating and cost follow the speed of change. In our scenario, architecture decisions change at a reasonable rate. An ADR superseded today requires, under RAG, updating the corpus, and under fine-tuning, retraining and re-evaluating the model so it forgets the previous decision. The more volatile the content, the heavier fine-tuning's recurring cost. On the other hand, parameter-efficient methods reduce the cost of training itself: LoRA reports roughly 10,000 times fewer trainable parameters and about 3 times less GPU memory, compared to full fine-tuning of GPT-3 175B with Adam. That lowers the barrier to entry, but it does not eliminate the cost of data, evaluation, and maintenance.
Auditability is not the same as verifiability. RAG lets you show where the answer came from, and that is a real advantage over knowledge diffused in weights. But citing a source does not guarantee the source supports the claim. In a human audit of generative search engines, only 51.5% of sentences were fully supported by their citations, and 74.5% of citations supported the corresponding sentence. These were 2023 systems and not corporate assistants, so the number does not transfer, but the warning stands: a citation must be verified, not merely displayed.
Security enters through retrieval. OWASP treats vectors and embeddings as a risk of their own in its 2025 Top 10 (LLM08), including context leakage between users in shared vector databases, and recommends permission-aware storage and immutable logs of retrievals. In an ADR assistant, not every ADR should be visible to every team. This topic deserves its own article in the series.
Fine-tuning ties you to the base model's lifecycle. Providers retire models. Anthropic states a notice of at least 60 days before retiring public models, and OpenAI, of at least 6 months for generally available models and around 2 weeks for previews, according to the pages consulted in October 2026. A model fine-tuned on a base that is discontinued has to be redone. This is an inference by the author, not a documented rule for every case, but it is a cost worth putting in the budget from the start.
The balance point between these dimensions depends on the case. In general, and this is a rule of thumb from the author rather than a research result, a stable corpus with tight latency and a rigid output format leans toward fine-tuning, while a corpus that changes every week and demands traceability leans toward retrieval and tools. What the table delivers is the list of questions to ask before deciding, and it is with that list that the next section's decision tree is built.
7. The decision framework

Figure 4. From failure to intervention. A synthesis by the author, to validate with your own data.
Figure 4 turns the previous sections into a procedure. Before using it, a caveat: it is a synthesis by the author, built from the axes in section 3 and from the split OpenAI makes between context and behavior. It is not a published taxonomy and has not been empirically validated. Treat it as a starting point for your team, not as a rule.
How to use it. The tree starts from a concrete failure, already observed and recorded, never from a preference. The tree has a step zero and three steps. For each question, the diagnostic test from section 3 says how to answer with evidence instead of opinion:
- Is the instruction ambiguous or the reasoning insufficient? Adjust the task specification (instruction, success criteria, and output format), decompose the task, or consider a more capable model. It is the cheapest step to check, and OpenAI's guide recommends starting with prompt engineering.
- Does the answer depend on private, recent, or mutable information? Run the manual context test: with the right information at hand, does the model get it right? If so, the path is retrieval, and the first check is the corpus's curation (H2) before touching retrieval.
- Is the content right, but form, style, or format fail? If the error persists with correct context and instructions, fine-tuning is a candidate, but only after proving that instructions and examples are not enough. That "proving" is the subject of one of the upcoming articles in the series.
- Does the task require computation, a system query, or an action? Prefer a tool that queries the authoritative source over an answer based on text.
Step zero comes first because fixing the task specification is the cheapest and because no capability compensates for a poorly specified task. After each fix, re-evaluate the original failure before moving on. And if no question gets a "yes," the conclusion is to go back to diagnosis: the failure may have more than one cause, or the cause may not have been identified yet.
Applying it to our scenario. The table shows how four hypothetical failures of the ADR assistant would travel the tree. These are illustrative examples, not results.
| Observed failure | Diagnostic test | Path through the tree | Intervention |
|---|---|---|---|
| The ADR that superseded the old pattern was never indexed | With the new ADR in context, the model gets it right | Step 1: yes | Index it and fix the ingestion process |
| The old, "superseded" ADR is retrieved instead of the new one | The model gets it right when only the standing ADR is provided | Step 1: yes | Curation: status and supersession relationship, filter at retrieval |
| The model receives the right ADR and still recommends the old pattern | With clear context and instructions, does the error persist, or does it disappear with a better instruction? | Step 0 first; then step 2 | Reinforce the instruction to prioritize the standing decision; fine-tuning only if it persists |
| The question is whether an API is still active today | The source of truth is the catalog, not a document | Step 3: yes | A tool that queries the API catalog |
Notice what the four rows have in common: none of them calls, as a first step, for choosing between RAG and fine-tuning. Three of them are resolved without touching the model's weights. This is exactly what the article's thesis anticipates, but the limit is worth repeating: these are four scenarios chosen to illustrate the procedure, not an estimate of how often each cause appears in practice.
Checklist before acting. On any team, before approving an intervention, it is worth being able to answer:
- What is the failure, and is it recorded with the input, the expected output, and the output obtained?
- Which diagnostic test confirmed the category of the cause?
- What is the cheapest alternative that has already been tried?
- How will we know the intervention worked, with which evaluation set and which criterion defined before running?
8. The takeaway
Let us return, one last time, to the team that saw the assistant recommend an abandoned architectural pattern. In the version of the conversation that starts from the technique, they leave with a decision: fine-tuning or RAG. In the version that starts from the failure, they leave with an answered question: what, exactly, went wrong in that answer? The second version is slower on day one and cheaper on every other day.
The principle this article proposes fits in one sentence: diagnose the failure before choosing the technique. To apply it, three ideas:
- Separate knowing, behaving, and doing. Knowledge, behavior, and execution are different problems (H1), and each points to a different capability: retrieval, model adaptation, or tools. Before them, confirm that the task is well specified.
- Curate the corpus before blaming the model. Part of knowledge failures are curation failures (H2): the corpus contains, by design, both the standing and the obsolete, and no technique applied to the model tells one from the other without status, dates, and explicit relationships.
- Demand evidence before intervening. Record the failure, confirm the category with a diagnostic test, try the cheapest alternative, and define, before running, how you will know it worked.
One final caveat, for consistency with what the article asks for. H1 and H2 are working hypotheses, not results: the research we reviewed shows that the techniques act in different places, but none of the studies tests the distinction between a standing and an obsolete decision, and the decision tree is a synthesis by the author. Its value lies in offering a procedure your team can test and correct with your own data, not in replacing that test.
The question "RAG or fine-tuning?" is not wrong because it is useless. It is wrong because it arrives too early. When diagnosis comes first, it comes back as a better question: which of these capabilities is my failure asking for?
About the series
This is the first article in the Beyond the Model series, about the decisions that turn language models into reliable software systems. In the next ones, the series goes deeper into the two points left open here:
- Stop Adding RAG to Everything details the failure taxonomy: how a wrong answer is classified into knowledge, retrieval, instruction, reasoning, tools, or behavior, and why each cause points to a different intervention.
- Before You Fine-Tune an LLM, Prove That You Need To shows how to turn the decision into an experiment: baseline, evaluation set, metrics, and decision criteria defined before training.
If you have been through a conversation like the one at the start of this article, tell me in the comments which failure sparked the discussion and which intervention the team chose. Those cases are exactly the material the next articles need.
References
- Barnett, S. et al. (2024). Seven Failure Points When Engineering a Retrieval Augmented Generation System. arXiv:2401.05856
- Lewis, P. et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020
- Hu, E. J. et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685
- Ovadia, O. et al. (2023). Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs. arXiv:2312.05934
- OpenAI. Optimizing LLM Accuracy
- Gekhman, Z. et al. (2024). Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?. EMNLP 2024
- Yang et al. (2026). Fine-Tuning vs. RAG for Multi-Hop Question Answering with Novel Knowledge. arXiv:2601.07054 (preprint)
- Nygard, M. (2011). Documenting Architecture Decisions. Cognitect blog, 15 Nov 2011
- Gruber, T. R. (1993). A Translation Approach to Portable Ontology Specifications. Knowledge Acquisition, 5(2), 199–220
- Pan, S. et al. (2024). Unifying Large Language Models and Knowledge Graphs: A Roadmap. IEEE TKDE (arXiv:2306.08302)
- Edge, D. et al. (2024). From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130
- Zhang, T. et al. (2024). RAFT: Adapting Language Model to Domain Specific RAG. arXiv:2403.10131 (preprint)
- Yao, S. et al. (2022). ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629 (ICLR 2023)
- Schick, T. et al. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv:2302.04761
- Model Context Protocol. Specification (revision 2026-07-28)
- Zaharia, M. et al. (2024). The Shift from Models to Compound AI Systems. BAIR Blog, 18 Feb 2024
- Anthropic (2024). Building Effective Agents. 19 Dec 2024
- Soudani, H., Kanoulas, E., Hasibi, F. (2024). Fine Tuning vs. Retrieval Augmented Generation for Less Popular Knowledge. arXiv:2403.01432
- Balaguer, A. et al. (2024). RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture. arXiv:2401.08406 (preprint)
- Abonizio, H., Almeida, T. S., Lotufo, R., Nogueira, R. (2025). Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime. arXiv:2508.06178 (preprint)
- Mecklenburg, N. et al. (2024). Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning. arXiv:2404.00213 (preprint)
- Liu, N. F., Zhang, T., Liang, P. (2023). Evaluating Verifiability in Generative Search Engines. Findings of EMNLP 2023 (arXiv:2304.09848)
- OWASP GenAI Security Project. LLM08:2025 Vector and Embedding Weaknesses
- Anthropic. Model deprecations (consulted 2026-10-10)
- OpenAI. Deprecations (consulted 2026-10-10)
Top comments (0)