<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Victor Lopes</title>
    <description>The latest articles on DEV Community by Victor Lopes (@theguitarvity).</description>
    <link>https://dev.to/theguitarvity</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4175558%2Fe11865f1-9655-423f-87e0-a80fbef65caf.jpg</url>
      <title>DEV Community: Victor Lopes</title>
      <link>https://dev.to/theguitarvity</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/theguitarvity"/>
    <language>en</language>
    <item>
      <title>RAG vs. Fine-Tuning: The Wrong Question to Ask When Building AI Systems</title>
      <dc:creator>Victor Lopes</dc:creator>
      <pubDate>Sat, 10 Oct 2026 16:52:32 +0000</pubDate>
      <link>https://dev.to/theguitarvity/rag-vs-fine-tuning-the-wrong-question-to-ask-when-building-ai-systems-4p3a</link>
      <guid>https://dev.to/theguitarvity/rag-vs-fine-tuning-the-wrong-question-to-ask-when-building-ai-systems-4p3a</guid>
      <description>&lt;p&gt;A platform team is evaluating an LLM-based assistant, still in testing, meant to answer developers' architecture questions using the company's ADRs (Architecture Decision Records) and internal documentation. During an evaluation round, the team itself notices that the assistant recommended a service-to-service authentication pattern the company abandoned last year and, on top of that, suggested calling an endpoint on an internal API already marked as deprecated. The answer was plausible, well written, and technically defensible in general terms. It just contradicted a standing decision.&lt;/p&gt;

&lt;p&gt;The failure was caught before it reached any user, and that is exactly why the conversation that follows matters. Two proposals land on the table. The first: "We need to fine-tune on our ADRs, so the model learns our decisions." The second: "We need RAG over the ADRs, so it can look up the standing decision." Both are reasonable, and each starts from a theory about what the model is missing: one assumes it needs to &lt;em&gt;learn&lt;/em&gt; the content, the other that it needs to &lt;em&gt;consult&lt;/em&gt; the content. But there is a prior question that neither one answers: what, exactly, went wrong in that answer?&lt;/p&gt;

&lt;p&gt;That question matters more than it seems, because the same outdated recommendation can have very different origins. The ADR that superseded the old pattern may never have been indexed. The old ADR, marked "superseded," may have been retrieved instead of the new one. Both may have been retrieved and the model ignored the context, answering from what it learned about the pattern in public documentation. Or the question may have been generic and the system instruction never asked it to prioritize the standing decision. Each cause calls for a different fix, and none of them is resolved, on its own, by whichever technique won the argument. Fine-tuning on today's ADRs, for example, does not help when the decision is superseded tomorrow. This is not unique to our scenario: an engineering report on RAG systems identified &lt;a href="https://arxiv.org/abs/2401.05856" rel="noopener noreferrer"&gt;seven distinct failure points&lt;/a&gt; when building this kind of system.&lt;/p&gt;

&lt;p&gt;To understand how we got to this debate, it helps to recall the context. Since ChatGPT arrived in late 2022, organizations have been trying to put LLMs in front of what used to live in wikis, shared folders, and operational documents, with the promise of making that knowledge more accessible to whoever needs it. Along the way, two techniques that were research topics became everyday vocabulary: retrieval-augmented generation (RAG), proposed by &lt;a href="https://arxiv.org/abs/2005.11401" rel="noopener noreferrer"&gt;Lewis et al. in 2020&lt;/a&gt;, and fine-tuning, which parameter-efficient methods like &lt;a href="https://arxiv.org/abs/2106.09685" rel="noopener noreferrer"&gt;LoRA&lt;/a&gt; made far cheaper to run. When two powerful tools arrive together and address the same apparent "pain," it is natural for the market to put them head to head. That is how the standard framing was born: &lt;em&gt;RAG or fine-tuning?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The problem with that framing is that it asks you to pick the answer before agreeing on the question. There is &lt;a href="https://arxiv.org/abs/2312.05934" rel="noopener noreferrer"&gt;research comparing the two approaches side by side&lt;/a&gt;, and it deserves a careful reading, not a slogan. &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;OpenAI's guide on optimizing accuracy&lt;/a&gt;, for instance, treats them as answers to different problems and as approaches that are "additive, not exclusive." The question, therefore, should not start from the technique. It should start from the failure.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9fglx205tqzi67wwg55a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9fglx205tqzi67wwg55a.png" alt="Two ways to start the same conversation: starting from the technique, the team picks RAG or fine-tuning and the original failure remains undiagnosed; starting from the failure, the cause is identified before the intervention." width="800" height="427"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 1. The difference between the two conversations is not the technique chosen, but the step that was skipped.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This article argues for that inversion and proposes a practical way to make it, separating what a system needs to &lt;em&gt;know&lt;/em&gt;, how it needs to &lt;em&gt;behave&lt;/em&gt;, and what it needs to &lt;em&gt;do&lt;/em&gt;. It is the first step in a series about the decisions that happen &lt;em&gt;beyond the model&lt;/em&gt;, where an AI system either becomes reliable or stops being so.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The problem
&lt;/h2&gt;

&lt;p&gt;Back to the architecture assistant. The two proposals on the table have something in common: both start from a technique and only then look for the problem it would solve. It is a subtle inversion, because the conversation sounds technical and mature. But engineering's usual path is the opposite: observe the symptom, form a hypothesis about the cause, test it, and only then intervene. In LLM-based systems, diagnosis tends to be the easiest step to skip.&lt;/p&gt;

&lt;p&gt;Why is it skipped? From my own observation, we can think of several reasons, but three seem the most plausible. First, techniques are tangible: they have tools, tutorials, budget, and a name that fits on a roadmap. Second, diagnosing is work: it requires opening the traces, separating what was retrieved from what was generated, reproducing the failure, and understanding at which stage it originated. Third, "RAG or fine-tuning?" is a question you can answer with an opinion, while "what caused this failure?" can only be answered with evidence.&lt;/p&gt;

&lt;p&gt;The cost of choosing the technique before the diagnosis shows up in at least three ways.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The original failure is still there.&lt;/strong&gt; If the problem was the old ADR being retrieved instead of the new one, training the model on the ADRs changes nothing: the wrong content still reaches the context. If the problem was the model ignoring the retrieved context, retrieving more context does not solve it either. The intervention ships, the team considers the problem solved, and the outdated recommendation shows up again in the next evaluation round.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The intervention can make the system worse.&lt;/strong&gt; &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;OpenAI's own guide on optimizing accuracy&lt;/a&gt; includes an example where RAG "confused" the model by adding noise and lowered the score by four points. On the fine-tuning side, a &lt;a href="https://arxiv.org/abs/2405.05904" rel="noopener noreferrer"&gt;controlled study of question answering without external lookup&lt;/a&gt; showed that examples containing new knowledge are learned more slowly and, once learned, increase the model's tendency to hallucinate. That result holds for that specific setting and should not be extended to all fine-tuning. Still, it shows that "one more technique" is not a risk-free bet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The complexity becomes yours.&lt;/strong&gt; Each technique brings a new maintenance surface. RAG requires indexing, updating, ranking, and permission control. Fine-tuning requires training data, versioning, re-evaluation on every base-model change, and a plan for when the content changes. If the real cause was an ambiguous instruction, the team paid that cost to solve something a prompt adjustment would have fixed. Not by coincidence, that same &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;OpenAI guide&lt;/a&gt; recommends starting with prompt engineering.&lt;/p&gt;

&lt;p&gt;None of this means either technique is the villain. The evidence is mixed, and that is part of the problem. In direct comparisons of knowledge injection, RAG outperformed unsupervised fine-tuning, both for previously seen knowledge and for new knowledge (&lt;a href="https://arxiv.org/abs/2312.05934" rel="noopener noreferrer"&gt;Ovadia et al., 2023&lt;/a&gt;). A &lt;a href="https://arxiv.org/abs/2601.07054" rel="noopener noreferrer"&gt;more recent study, on multi-hop questions over novel knowledge&lt;/a&gt;, found the highest overall accuracy with supervised fine-tuning. Both are preprints, with small or synthetic benchmarks, and neither answers our team's question. They show that each technique has its own territory, and that finding out which territory the failure sits in is precisely the work the question "RAG or fine-tuning?" skips. We return to these studies, in more detail, in section 5.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The engineering hypothesis
&lt;/h2&gt;

&lt;p&gt;If the problem is starting from the technique, the alternative needs to be something testable, not just a preference for a method. I propose the following working hypothesis:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;H1.&lt;/strong&gt; A large share of the observable failures in an assistant like ours can be attributed to one of three categories, &lt;strong&gt;knowledge&lt;/strong&gt;, &lt;strong&gt;behavior&lt;/strong&gt;, or &lt;strong&gt;execution&lt;/strong&gt;, and each category responds better to a different intervention.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is a &lt;strong&gt;hypothesis, not a result&lt;/strong&gt;. It was not tested for this article. To test it, a team could collect the failures from an evaluation round, ask independent reviewers to classify each one, and measure the agreement between them. Next, they would apply the intervention indicated for each category and compare the outcome against the alternative. The hypothesis weakens if inter-reviewer agreement is low, if most failures do not fit the three categories, or if all of them call for the same intervention. Upcoming articles in the series return to this experimental design.&lt;/p&gt;

&lt;p&gt;There is, however, a refinement our scenario suggests, and one that usually stays out of the discussion. Back to the superseded ADR. In his original text on the format, Michael Nygard recommends that reversed decisions not be deleted: they stay on record and are marked "superseded," because "it is still relevant to know that it &lt;em&gt;was&lt;/em&gt; the decision, but is &lt;em&gt;no longer&lt;/em&gt; the decision" (&lt;a href="https://www.cognitect.com/blog/2011/11/15/documenting-architecture-decisions" rel="noopener noreferrer"&gt;Documenting Architecture Decisions&lt;/a&gt;). In other words, a healthy ADR repository &lt;strong&gt;contains, by design, obsolete knowledge&lt;/strong&gt;. What distinguishes the standing from the obsolete is the status, the date, and the relationship between records, not the text itself. A model that only sees the text has no way to know.&lt;/p&gt;

&lt;p&gt;That leads to a second hypothesis, complementary to the first:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;H2.&lt;/strong&gt; Part of the failures classified as "knowledge" are not a lack of knowledge. They are &lt;strong&gt;curation&lt;/strong&gt; failures: the right information exists, but the way the corpus is organized does not allow telling the standing from the obsolete, and no technique applied to the model fixes that.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If H2 is true, the first question in the face of a knowledge failure stops being "RAG or fine-tuning?" and becomes "is the corpus in a condition to be consulted?". Curation, in this sense, is a set of engineering decisions about knowledge, prior to any decision about the model. Four forms, in increasing order of cost:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Form of curation&lt;/th&gt;
&lt;th&gt;What it addresses in our scenario&lt;/th&gt;
&lt;th&gt;Cost and risk&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Metadata and lifecycle&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Status (proposed, accepted, deprecated, superseded), date, owner, and a reference to the record that supersedes it; filtering on that information at retrieval time&lt;/td&gt;
&lt;td&gt;Low. Requires discipline in filling and updating&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Canonical source and deduplication&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Guarantees a single source of truth per decision, instead of diverging copies across wikis and folders&lt;/td&gt;
&lt;td&gt;Medium. Requires an owner and a review process&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Explicit relationships&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Records that ADR 31 &lt;em&gt;supersedes&lt;/em&gt; ADR 12 and that a pattern &lt;em&gt;applies to&lt;/em&gt; a set of services&lt;/td&gt;
&lt;td&gt;Medium to high. Requires modelling the domain&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Ontology and knowledge graph&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Represents concepts (decision, pattern, API, service) and relationships in a queryable, verifiable way&lt;/td&gt;
&lt;td&gt;High. Requires modelling, construction, and continuous maintenance&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The last two items deserve a comment, because they enter ontology territory. &lt;a href="https://tomgruber.org/writing/ontolingua-kaj-1993/" rel="noopener noreferrer"&gt;Gruber's classic definition&lt;/a&gt; describes an ontology as a formal specification of a conceptualization, that is, an explicit agreement about which concepts exist in a domain and how they relate. In our case, that would mean declaring that "Decision," "Pattern," "API," and "Service" are kinds of thing, and that "supersedes," "depends on," and "applies to" are relationships between them. With that, the question "what is the standing decision for service-to-service authentication?" stops being a text-similarity search and becomes a query with a verifiable answer.&lt;/p&gt;

&lt;p&gt;The literature on combining LLMs with knowledge graphs points to both the potential and the cost. &lt;a href="https://arxiv.org/abs/2306.08302" rel="noopener noreferrer"&gt;Pan et al.&lt;/a&gt; argue that knowledge graphs store factual knowledge explicitly and can help LLMs with external information and interpretability, and that the two are complementary. The same authors note that these graphs are hard to build and are constantly evolving. On the practical side, &lt;a href="https://arxiv.org/abs/2404.16130" rel="noopener noreferrer"&gt;GraphRAG&lt;/a&gt; proposes using an LLM to derive an entity graph from the documents and, with it, answer questions about the entire corpus, which conventional RAG struggles with. The authors report improvements in comprehensiveness and diversity of answers on datasets of around one million tokens. One caveat: GraphRAG was designed for global summarization questions, not for the problem of resolving which decision is standing. It shows that giving structure to the corpus can help, but it does not demonstrate that it solves our case. That remains a hypothesis to test.&lt;/p&gt;

&lt;p&gt;In short, the article's question gains two layers. The first is H1's: which category is the failure in? The second is H2's: if it is knowledge, is the corpus curated well enough to be used? Only after those two answers does it make sense to discuss RAG, fine-tuning, or tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Three axes: knowledge, behavior, task execution
&lt;/h2&gt;

&lt;p&gt;If H1 is to be useful, it needs categories a team can apply to a real failure without relying on interpretation. I propose three axes, each defined by a question you can ask of the system.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkicp82h2xvclvw7kb2cw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkicp82h2xvclvw7kb2cw.png" alt="Three kinds of problem, three capabilities: knowledge (RAG), behavior (fine-tuning), and execution (tools)." width="800" height="440"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 2. Each kind of failure points to a different capability.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Knowledge: does the model have access to the right information?
&lt;/h3&gt;

&lt;p&gt;This is the axis of the outdated recommendation. The typical symptom is an answer citing a decision that does not exist, has expired, or has been superseded. The diagnostic test is simple: if we manually place the standing ADR in the context, does the model get it right? If it does, the problem is not the model's capability but access to information, and the conversation shifts to retrieval and curation (H2). This is the territory &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;OpenAI's guide&lt;/a&gt; assigns to context optimization: the case where the model lacks knowledge because it was not in the training data. The natural intervention is RAG, with the caveat that it has &lt;a href="https://arxiv.org/abs/2401.05856" rel="noopener noreferrer"&gt;several failure points of its own, from indexing to generation&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Behavior: given the information, does the model use it the expected way?
&lt;/h3&gt;

&lt;p&gt;Here the model has the right ADR in context and still answers poorly. It does not follow the format the team requires, does not cite the decision's status, mixes recommendation with opinion, or loses the expected structure in long answers. The diagnostic test changes: with correct context and a clear instruction, does the formatting error persist consistently? If it does, and only after exhausting instructions and examples, fine-tuning becomes a candidate. That same &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;OpenAI guide&lt;/a&gt; associates it with inconsistent or incorrectly formatted results. There is conceptual support for this split: &lt;a href="https://arxiv.org/abs/2405.05904" rel="noopener noreferrer"&gt;Gekhman et al.&lt;/a&gt; conclude that models acquire most factual knowledge during pre-training, and that fine-tuning mainly teaches them to use it better. &lt;a href="https://arxiv.org/abs/2403.10131" rel="noopener noreferrer"&gt;RAFT&lt;/a&gt; shows a concrete use: training the model to ignore irrelevant retrieved documents and cite the relevant passages, that is, fine-tuning shaping how the model uses retrieval, not what it knows.&lt;/p&gt;

&lt;h3&gt;
  
  
  Execution: does the task require querying a system or acting?
&lt;/h3&gt;

&lt;p&gt;Some questions should not be answered by text, neither from the model nor from a document. If the question is "is this API deprecated today?", the source of truth is the API catalog, not an ADR written months ago. The symptom is an answer that mixes model memory or outdated text with a fact that lives in a queryable system. The intervention is a tool: the model calls the catalog, receives the current status, and answers based on it. Work like &lt;a href="https://arxiv.org/abs/2210.03629" rel="noopener noreferrer"&gt;ReAct&lt;/a&gt;, which combines reasoning and acting, and &lt;a href="https://arxiv.org/abs/2302.04761" rel="noopener noreferrer"&gt;Toolformer&lt;/a&gt;, where the model learns when and how to call APIs such as a calculator and search, shows that tools address what the model cannot maintain on its own. There is a porous boundary with the knowledge axis, because querying a catalog is also access to information. The practical difference is the source: a document corpus that needs curating, on one side, and an authoritative, queryable, current system on the other. Tools also carry their own risks. The &lt;a href="https://modelcontextprotocol.io/specification/latest" rel="noopener noreferrer"&gt;MCP specification&lt;/a&gt;, for example, treats tools as arbitrary code execution and advises treating tool behavior descriptions as untrusted unless they come from a trusted source.&lt;/p&gt;

&lt;h3&gt;
  
  
  What falls outside the three axes
&lt;/h3&gt;

&lt;p&gt;The three axes do not cover everything, and it is important to say so. An ambiguous instruction and insufficient reasoning also produce wrong answers, and none of the three capabilities fixes that. That is why the decision tree in section 7 begins with a step zero, before any technique: improve the task specification, decompose it, or use a more capable model. This is also consistent with &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;OpenAI's advice&lt;/a&gt; to start with prompt engineering. This limitation is one of the reasons H1 speaks of "a large share of failures," not all of them.&lt;/p&gt;

&lt;p&gt;There is also a point about combinations. The axes describe the cause of a failure, not an exclusive choice of solution. RAFT is an example of fine-tuning (behavior) working in service of retrieval (knowledge), and that same &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;OpenAI guide&lt;/a&gt; describes the approaches as additive, not exclusive. The gain of thinking in axes is knowing which lever to pull first and how to tell whether it worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. The architecture
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe68gnht0atdizyipszjc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe68gnht0atdizyipszjc.png" alt="Architecture: question, retrieval with knowledge base, base or adapted model, tools, and answer." width="800" height="454"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 3. The model is one component. Knowledge, behavior, and actions have different homes in the system.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The three axes from the previous section take on practical meaning once we place them in an architecture. Figure 3 shows the ADR assistant from our scenario and indicates where each capability lives. The central idea is that the model is one component among several, and that each axis corresponds to a different place in the system where you can intervene.&lt;/p&gt;

&lt;p&gt;Follow the path of a question like "which authentication pattern should we use between services?".&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval.&lt;/strong&gt; The question passes through a block that searches the documents and applies filters, for example discarding ADRs with superseded status. This is where the &lt;strong&gt;knowledge&lt;/strong&gt; axis manifests at runtime.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Knowledge base.&lt;/strong&gt; Behind retrieval sits the corpus, and that is where section 2's curation happens: status, relationships, canonical source. Optimal retrieval over a poorly curated corpus still returns outdated answers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model.&lt;/strong&gt; Receives the question and the context and generates the answer. The &lt;strong&gt;behavior&lt;/strong&gt; axis lives here: instructions, examples, and, when justified, a model adapted by fine-tuning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tools.&lt;/strong&gt; Instead of trusting text, the model can query an authoritative system, such as the API catalog, and come back with the current fact. This is the &lt;strong&gt;execution&lt;/strong&gt; axis.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answer.&lt;/strong&gt; The result returns to the user, ideally with the source and the status of the decisions used.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This reading is consistent with a trend described in the AI engineering literature. The Berkeley group, in a &lt;a href="https://bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/" rel="noopener noreferrer"&gt;post describing "compound AI systems"&lt;/a&gt;, argues that state-of-the-art results increasingly come from systems with multiple components, such as model calls, retrievers, and tools, rather than from an isolated model, and that iterating on the system tends to be faster than retraining the model. &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;Anthropic's guide on agents&lt;/a&gt; starts from the same idea and recommends seeking the simplest possible solution, noting that many applications need only a single well-optimized model call. The architecture in Figure 3 is therefore a ceiling of complexity, not a goal. Not every system needs every block, and starting without them is a legitimate choice.&lt;/p&gt;

&lt;p&gt;Two observations about the diagram. The first is that the blocks have different owners and lifecycles. The corpus changes when a decision changes, the model changes when the provider changes or when there is retraining, and tools change with the systems they query. Mixing these layers into one, for example by training the model on content that should live in the corpus, is the kind of decision section 1 called "complexity that becomes yours." The second is that the figure deliberately omits the layer that cuts across every block: evaluation and observability. Without recording what was retrieved, what the model received, and which tool was called, the diagnosis from the previous sections is not possible, and we go back to choosing techniques by opinion. That topic deserves its own article in the series.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Evidence
&lt;/h2&gt;

&lt;p&gt;Before proposing a decision tree, it is worth asking what research has already shown. The short answer is that useful evidence exists, but it is more limited than the debates make it seem. The studies below compare the two approaches for injecting knowledge. This section is based on the studies' abstracts, not the full papers, so I only cite results the abstracts themselves state.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Study&lt;/th&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;What the authors report&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2312.05934" rel="noopener noreferrer"&gt;Ovadia et al. (2023)&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Knowledge-intensive tasks, with previously seen and new knowledge&lt;/td&gt;
&lt;td&gt;Unsupervised fine-tuning improves things somewhat, but RAG is better in both cases. Models struggle to learn new facts this way&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2403.01432" rel="noopener noreferrer"&gt;Soudani et al. (2024)&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Synthetic question answering over low-popularity knowledge, across twelve models&lt;/td&gt;
&lt;td&gt;Fine-tuning improves at every popularity level, but RAG beats it by a wide margin, especially on the least popular facts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2401.08406" rel="noopener noreferrer"&gt;Balaguer et al. (2024), Microsoft&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Agriculture case study, with Llama2-13B, GPT-3.5, and GPT-4&lt;/td&gt;
&lt;td&gt;Fine-tuning raised accuracy by more than 6 percentage points, and RAG added about 5 points on top. The effects accumulated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2403.10131" rel="noopener noreferrer"&gt;Zhang et al. (2024), RAFT&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Domain-specific adaptation with RAG&lt;/td&gt;
&lt;td&gt;Training the model to ignore irrelevant documents and cite the relevant ones brought consistent gains across three datasets&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2404.00213" rel="noopener noreferrer"&gt;Mecklenburg et al. (2024), Microsoft&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Recent sports events&lt;/td&gt;
&lt;td&gt;Supervised fine-tuning can inject facts, provided the training data is well constructed. Coverage was uneven with token-based scaling and more uniform with fact-based scaling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2601.07054" rel="noopener noreferrer"&gt;Yang et al. (2026)&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Multi-hop questions over novel knowledge, across three 7B models&lt;/td&gt;
&lt;td&gt;RAG gave consistent gains, but supervised fine-tuning had the highest overall accuracy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2508.06178" rel="noopener noreferrer"&gt;Abonizio et al. (2025)&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Knowledge injection in a low-resource regime&lt;/td&gt;
&lt;td&gt;RAG-based injection tended to degrade performance on control sets more than parametric methods did&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;What can be stated carefully.&lt;/strong&gt; For new or rare facts, unsupervised fine-tuning over raw text is a weak path, and RAG was superior in every direct comparison I read. At the same time, supervised fine-tuning with &lt;a href="https://arxiv.org/abs/2404.00213" rel="noopener noreferrer"&gt;well-constructed data&lt;/a&gt; can add knowledge and, in a &lt;a href="https://arxiv.org/abs/2601.07054" rel="noopener noreferrer"&gt;2026 study&lt;/a&gt;, had the best overall accuracy. And the effects can add up, as in the &lt;a href="https://arxiv.org/abs/2401.08406" rel="noopener noreferrer"&gt;agriculture case&lt;/a&gt;. The most honest reading is that the techniques act in different places, which supports H1, not that one beats the other.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the evidence does not say.&lt;/strong&gt; There are four limits the reader should keep in mind:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Quality of the evidence.&lt;/strong&gt; Almost all the studies above are preprints, and most use small or synthetic benchmarks, such as agriculture, sports events, and invented entities. There is no reason to assume the results transfer to a corporate corpus.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;None of them tests our problem.&lt;/strong&gt; The studies measure whether the model learns or retrieves facts. None measures the distinction between a standing and an obsolete decision, which is the center of this article's scenario, and therefore none tests H2.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RAG is not always better.&lt;/strong&gt; Besides the &lt;a href="https://arxiv.org/abs/2508.06178" rel="noopener noreferrer"&gt;Abonizio study&lt;/a&gt;, the example from &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;OpenAI's guide&lt;/a&gt; shows a case where RAG lowered the score through noise. The gain depends on retrieval bringing in the right content.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fine-tuning effects depend on the setting.&lt;/strong&gt; The increase in hallucination reported by &lt;a href="https://arxiv.org/abs/2405.05904" rel="noopener noreferrer"&gt;Gekhman et al.&lt;/a&gt; was measured on question answering without external lookup, with new knowledge. It does not extend to fine-tuning for format or style.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In short, the research supports the thesis that the two approaches do not compete on the same ground, but it does not replace a test on your own corpus and with your real failures. That is why one of the upcoming articles in the series covers the experiment.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. The trade-offs
&lt;/h2&gt;

&lt;p&gt;Choosing a capability means accepting a set of costs that do not appear on the proposal slide. The table below compares the three capabilities across six dimensions, applied to the ADR assistant. It is a &lt;strong&gt;qualitative&lt;/strong&gt; comparison: this article did not measure latency or cost, and I leave those measurements for a future article in the series rather than repeating numbers of dubious origin.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Retrieval (RAG)&lt;/th&gt;
&lt;th&gt;Fine-tuning&lt;/th&gt;
&lt;th&gt;Tools&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Updating&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Reindex what changed. When an ADR is superseded, updating the corpus is enough&lt;/td&gt;
&lt;td&gt;Retrain to reflect the change, and re-evaluate&lt;/td&gt;
&lt;td&gt;Immediate: the tool queries the current system&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Recurring cost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Indexing infrastructure and more context text on every query&lt;/td&gt;
&lt;td&gt;Training data, training, evaluation, and a new round on every base-model change&lt;/td&gt;
&lt;td&gt;Integration and maintenance of each tool and its contract&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Latency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Adds at least one search step before generation&lt;/td&gt;
&lt;td&gt;Generally adds no steps at inference&lt;/td&gt;
&lt;td&gt;Adds round trips on every call&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Auditability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Lets you point to the source, but the citation must be verified&lt;/td&gt;
&lt;td&gt;Knowledge diffused in the weights, hard to trace or remove&lt;/td&gt;
&lt;td&gt;Calls are loggable and reproducible&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Security&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Permissions and cross-context leakage must be handled at retrieval&lt;/td&gt;
&lt;td&gt;Training data becomes part of the model&lt;/td&gt;
&lt;td&gt;Tools are code execution and require trust and limits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Maintenance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Continuous curation of the corpus (H2)&lt;/td&gt;
&lt;td&gt;Coupling to the base model and the provider's lifecycle&lt;/td&gt;
&lt;td&gt;Versioning and compatibility of the integrations&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Some points in the table deserve an explanation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Updating and cost follow the speed of change.&lt;/strong&gt; In our scenario, architecture decisions change at a reasonable rate. An ADR superseded today requires, under RAG, updating the corpus, and under fine-tuning, retraining and re-evaluating the model so it forgets the previous decision. The more volatile the content, the heavier fine-tuning's recurring cost. On the other hand, parameter-efficient methods reduce the cost of training itself: &lt;a href="https://arxiv.org/abs/2106.09685" rel="noopener noreferrer"&gt;LoRA&lt;/a&gt; reports roughly 10,000 times fewer trainable parameters and about 3 times less GPU memory, compared to full fine-tuning of GPT-3 175B with Adam. That lowers the barrier to entry, but it does not eliminate the cost of data, evaluation, and maintenance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Auditability is not the same as verifiability.&lt;/strong&gt; RAG lets you show where the answer came from, and that is a real advantage over knowledge diffused in weights. But citing a source does not guarantee the source supports the claim. In a &lt;a href="https://arxiv.org/abs/2304.09848" rel="noopener noreferrer"&gt;human audit of generative search engines&lt;/a&gt;, only 51.5% of sentences were fully supported by their citations, and 74.5% of citations supported the corresponding sentence. These were 2023 systems and not corporate assistants, so the number does not transfer, but the warning stands: a citation must be verified, not merely displayed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security enters through retrieval.&lt;/strong&gt; OWASP treats vectors and embeddings as &lt;a href="https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/" rel="noopener noreferrer"&gt;a risk of their own in its 2025 Top 10 (LLM08)&lt;/a&gt;, including context leakage between users in shared vector databases, and recommends permission-aware storage and immutable logs of retrievals. In an ADR assistant, not every ADR should be visible to every team. This topic deserves its own article in the series.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fine-tuning ties you to the base model's lifecycle.&lt;/strong&gt; Providers retire models. Anthropic states a notice of &lt;a href="https://platform.claude.com/docs/en/about-claude/model-deprecations" rel="noopener noreferrer"&gt;at least 60 days&lt;/a&gt; before retiring public models, and OpenAI, of &lt;a href="https://developers.openai.com/api/docs/deprecations" rel="noopener noreferrer"&gt;at least 6 months&lt;/a&gt; for generally available models and around 2 weeks for previews, according to the pages consulted in October 2026. A model fine-tuned on a base that is discontinued has to be redone. This is an inference by the author, not a documented rule for every case, but it is a cost worth putting in the budget from the start.&lt;/p&gt;

&lt;p&gt;The balance point between these dimensions depends on the case. In general, and this is a rule of thumb from the author rather than a research result, a stable corpus with tight latency and a rigid output format leans toward fine-tuning, while a corpus that changes every week and demands traceability leans toward retrieval and tools. What the table delivers is the list of questions to ask before deciding, and it is with that list that the next section's decision tree is built.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. The decision framework
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnmmhrw99pvap6bepb8as.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnmmhrw99pvap6bepb8as.png" alt="Decision tree: from failure to intervention. Step 0: ambiguous instruction or weak reasoning lead to task specification, decomposition, or a more capable model. Step 1: private or mutable information leads to RAG and curation. Step 2: form and style lead to fine-tuning. Step 3: computation or action leads to tools." width="800" height="554"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 4. From failure to intervention. A synthesis by the author, to validate with your own data.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Figure 4 turns the previous sections into a procedure. Before using it, a caveat: it is a &lt;strong&gt;synthesis by the author&lt;/strong&gt;, built from the axes in section 3 and from the split &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;OpenAI makes between context and behavior&lt;/a&gt;. It is not a published taxonomy and has not been empirically validated. Treat it as a starting point for your team, not as a rule.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to use it.&lt;/strong&gt; The tree starts from a concrete failure, already observed and recorded, never from a preference. The tree has a step zero and three steps. For each question, the diagnostic test from section 3 says how to answer with evidence instead of opinion:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;em&gt;Is the instruction ambiguous or the reasoning insufficient?&lt;/em&gt; Adjust the task specification (instruction, success criteria, and output format), decompose the task, or consider a more capable model. It is the cheapest step to check, and &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;OpenAI's guide&lt;/a&gt; recommends starting with prompt engineering.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Does the answer depend on private, recent, or mutable information?&lt;/em&gt; Run the manual context test: with the right information at hand, does the model get it right? If so, the path is retrieval, and the first check is the corpus's curation (H2) before touching retrieval.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Is the content right, but form, style, or format fail?&lt;/em&gt; If the error persists with correct context and instructions, fine-tuning is a candidate, but only after proving that instructions and examples are not enough. That "proving" is the subject of one of the upcoming articles in the series.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Does the task require computation, a system query, or an action?&lt;/em&gt; Prefer a tool that queries the authoritative source over an answer based on text.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Step zero comes first because fixing the task specification is the cheapest and because no capability compensates for a poorly specified task. After each fix, re-evaluate the original failure before moving on. And if no question gets a "yes," the conclusion is to go back to diagnosis: the failure may have more than one cause, or the cause may not have been identified yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Applying it to our scenario.&lt;/strong&gt; The table shows how four hypothetical failures of the ADR assistant would travel the tree. These are illustrative examples, not results.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Observed failure&lt;/th&gt;
&lt;th&gt;Diagnostic test&lt;/th&gt;
&lt;th&gt;Path through the tree&lt;/th&gt;
&lt;th&gt;Intervention&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;The ADR that superseded the old pattern was never indexed&lt;/td&gt;
&lt;td&gt;With the new ADR in context, the model gets it right&lt;/td&gt;
&lt;td&gt;Step 1: yes&lt;/td&gt;
&lt;td&gt;Index it and fix the ingestion process&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The old, "superseded" ADR is retrieved instead of the new one&lt;/td&gt;
&lt;td&gt;The model gets it right when only the standing ADR is provided&lt;/td&gt;
&lt;td&gt;Step 1: yes&lt;/td&gt;
&lt;td&gt;Curation: status and supersession relationship, filter at retrieval&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The model receives the right ADR and still recommends the old pattern&lt;/td&gt;
&lt;td&gt;With clear context and instructions, does the error persist, or does it disappear with a better instruction?&lt;/td&gt;
&lt;td&gt;Step 0 first; then step 2&lt;/td&gt;
&lt;td&gt;Reinforce the instruction to prioritize the standing decision; fine-tuning only if it persists&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The question is whether an API is still active today&lt;/td&gt;
&lt;td&gt;The source of truth is the catalog, not a document&lt;/td&gt;
&lt;td&gt;Step 3: yes&lt;/td&gt;
&lt;td&gt;A tool that queries the API catalog&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Notice what the four rows have in common: none of them calls, as a first step, for choosing between RAG and fine-tuning. Three of them are resolved without touching the model's weights. This is exactly what the article's thesis anticipates, but the limit is worth repeating: these are four scenarios chosen to illustrate the procedure, not an estimate of how often each cause appears in practice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Checklist before acting.&lt;/strong&gt; On any team, before approving an intervention, it is worth being able to answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What is the failure, and is it recorded with the input, the expected output, and the output obtained?&lt;/li&gt;
&lt;li&gt;Which diagnostic test confirmed the category of the cause?&lt;/li&gt;
&lt;li&gt;What is the cheapest alternative that has already been tried?&lt;/li&gt;
&lt;li&gt;How will we know the intervention worked, with which evaluation set and which criterion defined before running?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  8. The takeaway
&lt;/h2&gt;

&lt;p&gt;Let us return, one last time, to the team that saw the assistant recommend an abandoned architectural pattern. In the version of the conversation that starts from the technique, they leave with a decision: fine-tuning or RAG. In the version that starts from the failure, they leave with an answered question: what, exactly, went wrong in that answer? The second version is slower on day one and cheaper on every other day.&lt;/p&gt;

&lt;p&gt;The principle this article proposes fits in one sentence: &lt;strong&gt;diagnose the failure before choosing the technique.&lt;/strong&gt; To apply it, three ideas:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Separate knowing, behaving, and doing.&lt;/strong&gt; Knowledge, behavior, and execution are different problems (H1), and each points to a different capability: retrieval, model adaptation, or tools. Before them, confirm that the task is well specified.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Curate the corpus before blaming the model.&lt;/strong&gt; Part of knowledge failures are curation failures (H2): the corpus contains, by design, both the standing and the obsolete, and no technique applied to the model tells one from the other without status, dates, and explicit relationships.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Demand evidence before intervening.&lt;/strong&gt; Record the failure, confirm the category with a diagnostic test, try the cheapest alternative, and define, before running, how you will know it worked.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One final caveat, for consistency with what the article asks for. H1 and H2 are &lt;strong&gt;working hypotheses, not results&lt;/strong&gt;: the research we reviewed shows that the techniques act in different places, but none of the studies tests the distinction between a standing and an obsolete decision, and the decision tree is a synthesis by the author. Its value lies in offering a procedure your team can test and correct with your own data, not in replacing that test.&lt;/p&gt;

&lt;p&gt;The question "RAG or fine-tuning?" is not wrong because it is useless. It is wrong because it arrives too early. When diagnosis comes first, it comes back as a better question: &lt;em&gt;which of these capabilities is my failure asking for?&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  About the series
&lt;/h2&gt;

&lt;p&gt;This is the first article in the &lt;strong&gt;Beyond the Model&lt;/strong&gt; series, about the decisions that turn language models into reliable software systems. In the next ones, the series goes deeper into the two points left open here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stop Adding RAG to Everything&lt;/strong&gt; details the failure taxonomy: how a wrong answer is classified into knowledge, retrieval, instruction, reasoning, tools, or behavior, and why each cause points to a different intervention.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Before You Fine-Tune an LLM, Prove That You Need To&lt;/strong&gt; shows how to turn the decision into an experiment: baseline, evaluation set, metrics, and decision criteria defined before training.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have been through a conversation like the one at the start of this article, tell me in the comments which failure sparked the discussion and which intervention the team chose. Those cases are exactly the material the next articles need.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Barnett, S. et al. (2024). &lt;a href="https://arxiv.org/abs/2401.05856" rel="noopener noreferrer"&gt;&lt;em&gt;Seven Failure Points When Engineering a Retrieval Augmented Generation System&lt;/em&gt;&lt;/a&gt;. arXiv:2401.05856&lt;/li&gt;
&lt;li&gt;Lewis, P. et al. (2020). &lt;a href="https://arxiv.org/abs/2005.11401" rel="noopener noreferrer"&gt;&lt;em&gt;Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks&lt;/em&gt;&lt;/a&gt;. NeurIPS 2020&lt;/li&gt;
&lt;li&gt;Hu, E. J. et al. (2021). &lt;a href="https://arxiv.org/abs/2106.09685" rel="noopener noreferrer"&gt;&lt;em&gt;LoRA: Low-Rank Adaptation of Large Language Models&lt;/em&gt;&lt;/a&gt;. arXiv:2106.09685&lt;/li&gt;
&lt;li&gt;Ovadia, O. et al. (2023). &lt;a href="https://arxiv.org/abs/2312.05934" rel="noopener noreferrer"&gt;&lt;em&gt;Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs&lt;/em&gt;&lt;/a&gt;. arXiv:2312.05934&lt;/li&gt;
&lt;li&gt;OpenAI. &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;&lt;em&gt;Optimizing LLM Accuracy&lt;/em&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gekhman, Z. et al. (2024). &lt;a href="https://arxiv.org/abs/2405.05904" rel="noopener noreferrer"&gt;&lt;em&gt;Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?&lt;/em&gt;&lt;/a&gt;. EMNLP 2024&lt;/li&gt;
&lt;li&gt;Yang et al. (2026). &lt;a href="https://arxiv.org/abs/2601.07054" rel="noopener noreferrer"&gt;&lt;em&gt;Fine-Tuning vs. RAG for Multi-Hop Question Answering with Novel Knowledge&lt;/em&gt;&lt;/a&gt;. arXiv:2601.07054 (preprint)&lt;/li&gt;
&lt;li&gt;Nygard, M. (2011). &lt;a href="https://www.cognitect.com/blog/2011/11/15/documenting-architecture-decisions" rel="noopener noreferrer"&gt;&lt;em&gt;Documenting Architecture Decisions&lt;/em&gt;&lt;/a&gt;. Cognitect blog, 15 Nov 2011&lt;/li&gt;
&lt;li&gt;Gruber, T. R. (1993). &lt;a href="https://tomgruber.org/writing/ontolingua-kaj-1993/" rel="noopener noreferrer"&gt;&lt;em&gt;A Translation Approach to Portable Ontology Specifications&lt;/em&gt;&lt;/a&gt;. Knowledge Acquisition, 5(2), 199–220&lt;/li&gt;
&lt;li&gt;Pan, S. et al. (2024). &lt;a href="https://arxiv.org/abs/2306.08302" rel="noopener noreferrer"&gt;&lt;em&gt;Unifying Large Language Models and Knowledge Graphs: A Roadmap&lt;/em&gt;&lt;/a&gt;. IEEE TKDE (arXiv:2306.08302)&lt;/li&gt;
&lt;li&gt;Edge, D. et al. (2024). &lt;a href="https://arxiv.org/abs/2404.16130" rel="noopener noreferrer"&gt;&lt;em&gt;From Local to Global: A Graph RAG Approach to Query-Focused Summarization&lt;/em&gt;&lt;/a&gt;. arXiv:2404.16130&lt;/li&gt;
&lt;li&gt;Zhang, T. et al. (2024). &lt;a href="https://arxiv.org/abs/2403.10131" rel="noopener noreferrer"&gt;&lt;em&gt;RAFT: Adapting Language Model to Domain Specific RAG&lt;/em&gt;&lt;/a&gt;. arXiv:2403.10131 (preprint)&lt;/li&gt;
&lt;li&gt;Yao, S. et al. (2022). &lt;a href="https://arxiv.org/abs/2210.03629" rel="noopener noreferrer"&gt;&lt;em&gt;ReAct: Synergizing Reasoning and Acting in Language Models&lt;/em&gt;&lt;/a&gt;. arXiv:2210.03629 (ICLR 2023)&lt;/li&gt;
&lt;li&gt;Schick, T. et al. (2023). &lt;a href="https://arxiv.org/abs/2302.04761" rel="noopener noreferrer"&gt;&lt;em&gt;Toolformer: Language Models Can Teach Themselves to Use Tools&lt;/em&gt;&lt;/a&gt;. arXiv:2302.04761&lt;/li&gt;
&lt;li&gt;Model Context Protocol. &lt;a href="https://modelcontextprotocol.io/specification/latest" rel="noopener noreferrer"&gt;&lt;em&gt;Specification&lt;/em&gt;&lt;/a&gt; (revision 2026-07-28)&lt;/li&gt;
&lt;li&gt;Zaharia, M. et al. (2024). &lt;a href="https://bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/" rel="noopener noreferrer"&gt;&lt;em&gt;The Shift from Models to Compound AI Systems&lt;/em&gt;&lt;/a&gt;. BAIR Blog, 18 Feb 2024&lt;/li&gt;
&lt;li&gt;Anthropic (2024). &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;&lt;em&gt;Building Effective Agents&lt;/em&gt;&lt;/a&gt;. 19 Dec 2024&lt;/li&gt;
&lt;li&gt;Soudani, H., Kanoulas, E., Hasibi, F. (2024). &lt;a href="https://arxiv.org/abs/2403.01432" rel="noopener noreferrer"&gt;&lt;em&gt;Fine Tuning vs. Retrieval Augmented Generation for Less Popular Knowledge&lt;/em&gt;&lt;/a&gt;. arXiv:2403.01432&lt;/li&gt;
&lt;li&gt;Balaguer, A. et al. (2024). &lt;a href="https://arxiv.org/abs/2401.08406" rel="noopener noreferrer"&gt;&lt;em&gt;RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture&lt;/em&gt;&lt;/a&gt;. arXiv:2401.08406 (preprint)&lt;/li&gt;
&lt;li&gt;Abonizio, H., Almeida, T. S., Lotufo, R., Nogueira, R. (2025). &lt;a href="https://arxiv.org/abs/2508.06178" rel="noopener noreferrer"&gt;&lt;em&gt;Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime&lt;/em&gt;&lt;/a&gt;. arXiv:2508.06178 (preprint)&lt;/li&gt;
&lt;li&gt;Mecklenburg, N. et al. (2024). &lt;a href="https://arxiv.org/abs/2404.00213" rel="noopener noreferrer"&gt;&lt;em&gt;Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning&lt;/em&gt;&lt;/a&gt;. arXiv:2404.00213 (preprint)&lt;/li&gt;
&lt;li&gt;Liu, N. F., Zhang, T., Liang, P. (2023). &lt;a href="https://arxiv.org/abs/2304.09848" rel="noopener noreferrer"&gt;&lt;em&gt;Evaluating Verifiability in Generative Search Engines&lt;/em&gt;&lt;/a&gt;. Findings of EMNLP 2023 (arXiv:2304.09848)&lt;/li&gt;
&lt;li&gt;OWASP GenAI Security Project. &lt;a href="https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/" rel="noopener noreferrer"&gt;&lt;em&gt;LLM08:2025 Vector and Embedding Weaknesses&lt;/em&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Anthropic. &lt;a href="https://platform.claude.com/docs/en/about-claude/model-deprecations" rel="noopener noreferrer"&gt;&lt;em&gt;Model deprecations&lt;/em&gt;&lt;/a&gt; (consulted 2026-10-10)&lt;/li&gt;
&lt;li&gt;OpenAI. &lt;a href="https://developers.openai.com/api/docs/deprecations" rel="noopener noreferrer"&gt;&lt;em&gt;Deprecations&lt;/em&gt;&lt;/a&gt; (consulted 2026-10-10)&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>llm</category>
      <category>architecture</category>
    </item>
    <item>
      <title>RAG ou fine-tuning: a pergunta errada para quem constrói sistemas de IA</title>
      <dc:creator>Victor Lopes</dc:creator>
      <pubDate>Sat, 10 Oct 2026 16:21:02 +0000</pubDate>
      <link>https://dev.to/theguitarvity/rag-ou-fine-tuning-a-pergunta-errada-para-quem-constroi-sistemas-de-ia-38ld</link>
      <guid>https://dev.to/theguitarvity/rag-ou-fine-tuning-a-pergunta-errada-para-quem-constroi-sistemas-de-ia-38ld</guid>
      <description>&lt;p&gt;Um time de plataforma está validando um assistente baseado em LLM, ainda em fase de testes, para responder dúvidas de arquitetura dos desenvolvedores, com acesso aos ADRs (Architecture Decision Records) e à documentação interna. Numa rodada de avaliação, o próprio time percebe que o assistente recomendou um padrão de autenticação entre serviços que a empresa abandonou no ano passado e, de quebra, sugeriu chamar um endpoint de uma API interna já marcada como obsoleta. A resposta era plausível, bem escrita e tecnicamente defensável em termos gerais. Só contrariava uma decisão vigente.&lt;/p&gt;

&lt;p&gt;A falha foi pega antes de chegar a qualquer usuário, e é justamente aí que a conversa seguinte importa. Duas propostas aparecem na mesa. A primeira: "Precisamos fazer fine-tuning com os nossos ADRs, para o modelo aprender as nossas decisões." A segunda: "Precisamos de RAG sobre os ADRs, para ele consultar a decisão vigente." As duas são razoáveis, e cada uma parte de uma teoria sobre o que falta ao modelo: uma supõe que falta &lt;em&gt;aprender&lt;/em&gt; o conteúdo, a outra que falta &lt;em&gt;consultar&lt;/em&gt; o conteúdo. Mas existe uma pergunta anterior que nenhuma das duas responde: o que, exatamente, deu errado nessa resposta?&lt;/p&gt;

&lt;p&gt;Essa pergunta importa mais do que parece, porque a mesma recomendação desatualizada pode ter origens bem diferentes. O ADR que substituiu o padrão antigo pode nunca ter sido indexado. O ADR antigo, com status "superseded", pode ter sido recuperado no lugar do novo. Os dois podem ter sido recuperados e o modelo ter ignorado o contexto, respondendo a partir do que aprendeu sobre o padrão em documentação pública. Ou a pergunta pode ter sido genérica e a instrução do sistema nunca ter pedido para priorizar a decisão vigente. Cada causa pede uma correção diferente, e nenhuma delas se resolve, por si só, com a técnica que ganhou a discussão. Um fine-tuning com os ADRs de hoje, por exemplo, não ajuda quando a decisão for substituída amanhã. Isso não é exclusivo do nosso cenário: um relato de engenharia sobre sistemas RAG identificou &lt;a href="https://arxiv.org/abs/2401.05856" rel="noopener noreferrer"&gt;sete pontos de falha distintos&lt;/a&gt; ao construir esse tipo de sistema.&lt;/p&gt;

&lt;p&gt;Para entender como chegamos a esse debate, vale lembrar o contexto. Desde a chegada do ChatGPT, no fim de 2022, as organizações tentam colocar LLMs na frente do que antes vivia em wikis, pastas compartilhadas e documentos operacionais, com a promessa de tornar esse conhecimento mais acessível a quem precisa dele. No caminho, duas técnicas que eram tema de pesquisa viraram vocabulário do dia a dia: a geração aumentada por recuperação (RAG), proposta por &lt;a href="https://arxiv.org/abs/2005.11401" rel="noopener noreferrer"&gt;Lewis et al. em 2020&lt;/a&gt;, e o fine-tuning, que métodos eficientes em parâmetros como o &lt;a href="https://arxiv.org/abs/2106.09685" rel="noopener noreferrer"&gt;LoRA&lt;/a&gt; tornaram bem mais barato de executar. Quando duas ferramentas poderosas chegam juntas e resolvem o mesmo tipo de "dor" aparente, é natural que o mercado as coloque frente a frente. Assim nasceu o enquadramento padrão: &lt;em&gt;RAG ou fine-tuning?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;O problema desse enquadramento é que ele pede para escolher a resposta antes de combinar a pergunta. Existe &lt;a href="https://arxiv.org/abs/2312.05934" rel="noopener noreferrer"&gt;pesquisa comparando as duas abordagens lado a lado&lt;/a&gt;, e ela merece uma leitura cuidadosa, não um slogan. O &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;guia da OpenAI sobre otimização de acurácia&lt;/a&gt;, por exemplo, trata as duas como respostas a problemas diferentes e como abordagens "aditivas, não exclusivas". A pergunta, portanto, não deveria partir da técnica. Deveria partir da falha.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F260kmzfgvz8x0xsnsljh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F260kmzfgvz8x0xsnsljh.png" alt="Duas formas de começar a mesma conversa: começando pela técnica, o time escolhe RAG ou fine-tuning e a falha original continua sem diagnóstico; começando pela falha, a causa é identificada antes da intervenção." width="800" height="427"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figura 1. A diferença entre as duas conversas não está na técnica escolhida, e sim na etapa que foi pulada.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Este artigo defende essa inversão e propõe uma forma prática de fazê-la, separando o que um sistema precisa &lt;em&gt;saber&lt;/em&gt;, como precisa &lt;em&gt;se comportar&lt;/em&gt; e o que precisa &lt;em&gt;fazer&lt;/em&gt;. É o primeiro passo de uma série sobre as decisões que acontecem &lt;em&gt;além do modelo&lt;/em&gt;, onde um sistema de IA se torna confiável, ou deixa de ser.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. O problema
&lt;/h2&gt;

&lt;p&gt;Voltemos ao assistente de arquitetura. As duas propostas que apareceram na mesa têm algo em comum: ambas começam por uma técnica e só depois procuram o problema que ela resolveria. É uma inversão sutil, porque a conversa soa técnica e madura. Mas o caminho habitual da engenharia é o contrário: observar o sintoma, formular uma hipótese sobre a causa, testar e só então intervir. Em sistemas baseados em LLMs, o diagnóstico costuma ser a etapa mais fácil de pular.&lt;/p&gt;

&lt;p&gt;Por que ele é pulado? Pela minha observação, podemos pensar em várias razões, mas três parecem as mais plausíveis. Primeiro, técnicas são tangíveis: têm ferramentas, tutoriais, orçamento e um nome que cabe num roadmap. Segundo, diagnosticar dá trabalho: exige abrir os traces, separar o que foi recuperado do que foi gerado, reproduzir a falha e entender em qual etapa ela nasceu. Terceiro, "RAG ou fine-tuning?" é uma pergunta que se responde com opinião, enquanto "o que causou esta falha?" só se responde com evidência.&lt;/p&gt;

&lt;p&gt;O custo de escolher a técnica antes do diagnóstico aparece de pelo menos três formas.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A falha original continua lá.&lt;/strong&gt; Se o problema era o ADR antigo sendo recuperado no lugar do novo, treinar o modelo com os ADRs não muda nada: o conteúdo errado continua chegando ao contexto. Se o problema era o modelo ignorar o contexto recuperado, recuperar mais contexto também não resolve. A intervenção é entregue, o time dá o problema por resolvido, e a recomendação desatualizada volta a aparecer na rodada de testes seguinte.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A intervenção pode piorar o sistema.&lt;/strong&gt; O próprio &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;guia da OpenAI sobre otimização de acurácia&lt;/a&gt; traz um exemplo em que o RAG "confundiu" o modelo ao adicionar ruído e reduziu a nota em quatro pontos. Do lado do fine-tuning, um &lt;a href="https://arxiv.org/abs/2405.05904" rel="noopener noreferrer"&gt;estudo controlado de perguntas e respostas sem consulta externa&lt;/a&gt; mostrou que exemplos com conhecimento novo são aprendidos mais devagar e, depois de aprendidos, aumentam a tendência do modelo a alucinar. Esse resultado vale para aquele cenário específico e não deve ser estendido a todo fine-tuning. Ainda assim, mostra que "mais uma técnica" não é uma aposta sem risco.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A complexidade passa a ser sua.&lt;/strong&gt; Cada técnica traz uma superfície nova de manutenção. O RAG exige indexação, atualização, ranqueamento e controle de permissões. O fine-tuning exige dados de treinamento, versionamento, reavaliação a cada mudança de modelo base e um plano para quando o conteúdo mudar. Se a causa real era uma instrução ambígua, o time pagou esse custo para resolver algo que um ajuste na instrução resolveria. Não por acaso, o mesmo &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;guia da OpenAI&lt;/a&gt; recomenda começar pela engenharia de prompt.&lt;/p&gt;

&lt;p&gt;Nada disso significa que uma das técnicas seja a vilã. A evidência é mista, e isso é parte do problema. Em comparações diretas de injeção de conhecimento, o RAG superou o fine-tuning não supervisionado, tanto para conhecimento já visto quanto para conhecimento novo (&lt;a href="https://arxiv.org/abs/2312.05934" rel="noopener noreferrer"&gt;Ovadia et al., 2023&lt;/a&gt;). Já um &lt;a href="https://arxiv.org/abs/2601.07054" rel="noopener noreferrer"&gt;estudo mais recente, com perguntas de múltiplos saltos sobre conhecimento novo&lt;/a&gt;, encontrou a maior acurácia geral no fine-tuning supervisionado. Ambos são preprints, com benchmarks pequenos ou sintéticos, e nenhum deles responde à pergunta do nosso time. Eles mostram que cada técnica tem o seu território, e que descobrir em qual território a falha está é o trabalho que a pergunta "RAG ou fine-tuning?" dispensa. Voltamos a esses estudos, com mais detalhe, na seção 5.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. A hipótese de engenharia
&lt;/h2&gt;

&lt;p&gt;Se o problema é começar pela técnica, a alternativa precisa ser algo que possa ser testado, e não apenas uma preferência de método. Proponho a seguinte hipótese de trabalho:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;H1.&lt;/strong&gt; Grande parte das falhas observáveis em um assistente como o nosso pode ser atribuída a uma de três categorias, &lt;strong&gt;conhecimento&lt;/strong&gt;, &lt;strong&gt;comportamento&lt;/strong&gt; ou &lt;strong&gt;execução&lt;/strong&gt;, e cada categoria responde melhor a uma intervenção diferente.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Isto é uma &lt;strong&gt;hipótese, não um resultado&lt;/strong&gt;. Ela não foi testada para este artigo. Para testá-la, um time poderia coletar as falhas de uma rodada de avaliação, pedir a revisores independentes que classifiquem cada uma e medir a concordância entre eles. Em seguida, aplicaria a intervenção indicada para cada categoria e compararia o resultado com a alternativa. A hipótese enfraquece se a concordância entre revisores for baixa, se a maioria das falhas não couber nas três categorias ou se todas pedirem a mesma intervenção. Os próximos artigos da série voltam a esse desenho experimental.&lt;/p&gt;

&lt;p&gt;Há, porém, um refinamento que o nosso cenário sugere, e que costuma ficar de fora da discussão. Voltemos ao ADR substituído. Em seu &lt;a href="https://www.cognitect.com/blog/2011/11/15/documenting-architecture-decisions" rel="noopener noreferrer"&gt;texto original sobre o formato&lt;/a&gt;, Michael Nygard recomenda que decisões revertidas não sejam apagadas: elas ficam registradas e marcadas como "superseded", porque "ainda é relevante saber que aquela &lt;em&gt;foi&lt;/em&gt; a decisão, mas que &lt;em&gt;não é mais&lt;/em&gt; a decisão". Ou seja, um repositório de ADRs saudável &lt;strong&gt;contém, por desenho, conhecimento obsoleto&lt;/strong&gt;. O que distingue o vigente do obsoleto é o status, a data e a relação entre os registros, e não o texto em si. Um modelo que só vê o texto não tem como saber.&lt;/p&gt;

&lt;p&gt;Isso leva a uma segunda hipótese, complementar à primeira:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;H2.&lt;/strong&gt; Parte das falhas classificadas como "conhecimento" não é falta de conhecimento. É falha de &lt;strong&gt;curadoria&lt;/strong&gt;: a informação certa existe, mas a organização do acervo não permite distinguir o vigente do obsoleto, e nenhuma técnica sobre o modelo corrige isso.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Se H2 for verdadeira, a primeira pergunta diante de uma falha de conhecimento deixa de ser "RAG ou fine-tuning?" e passa a ser "o acervo está em condições de ser consultado?". Curadoria, nesse sentido, é um conjunto de decisões de engenharia sobre o conhecimento, antes de qualquer decisão sobre o modelo. Quatro formas, em ordem crescente de custo:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Forma de curadoria&lt;/th&gt;
&lt;th&gt;O que ela trata no nosso cenário&lt;/th&gt;
&lt;th&gt;Custo e risco&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Metadados e ciclo de vida&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Status (proposto, aceito, depreciado, substituído), data, responsável e referência ao registro que o substitui; filtro dessas informações na hora da recuperação&lt;/td&gt;
&lt;td&gt;Baixo. Exige disciplina de preenchimento e atualização&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Fonte canônica e deduplicação&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Garante uma única fonte de verdade por decisão, em vez de cópias divergentes em wikis e pastas&lt;/td&gt;
&lt;td&gt;Médio. Exige dono e processo de revisão&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Relações explícitas&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Registra que o ADR 31 &lt;em&gt;substitui&lt;/em&gt; o ADR 12 e que um padrão &lt;em&gt;se aplica a&lt;/em&gt; um conjunto de serviços&lt;/td&gt;
&lt;td&gt;Médio a alto. Exige modelar o domínio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Ontologia e grafo de conhecimento&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Representa conceitos (decisão, padrão, API, serviço) e relações de forma consultável e verificável&lt;/td&gt;
&lt;td&gt;Alto. Exige modelagem, construção e manutenção contínuas&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Os dois últimos itens merecem um comentário, porque entram no território de ontologias. A &lt;a href="https://tomgruber.org/writing/ontolingua-kaj-1993/" rel="noopener noreferrer"&gt;definição clássica de Gruber&lt;/a&gt; descreve uma ontologia como uma especificação formal de uma conceitualização, ou seja, um acordo explícito sobre quais conceitos existem em um domínio e como se relacionam. No nosso caso, isso significaria declarar que "Decisão", "Padrão", "API" e "Serviço" são tipos de coisa, e que "substitui", "depende de" e "se aplica a" são relações entre elas. Com isso, a pergunta "qual é a decisão vigente para autenticação entre serviços?" deixa de ser uma busca por similaridade de texto e passa a ser uma consulta com resposta verificável.&lt;/p&gt;

&lt;p&gt;A literatura sobre a combinação de LLMs com grafos de conhecimento aponta tanto o potencial quanto o custo. &lt;a href="https://arxiv.org/abs/2306.08302" rel="noopener noreferrer"&gt;Pan et al.&lt;/a&gt; defendem que grafos de conhecimento guardam conhecimento factual de forma explícita e podem ajudar os LLMs com informação externa e interpretabilidade, e que os dois são complementares. Os mesmos autores lembram que esses grafos são difíceis de construir e estão em constante evolução. Na linha prática, o &lt;a href="https://arxiv.org/abs/2404.16130" rel="noopener noreferrer"&gt;GraphRAG&lt;/a&gt; propõe usar um LLM para derivar um grafo de entidades a partir dos documentos e, com ele, responder a perguntas sobre o corpus inteiro, para as quais o RAG convencional tem dificuldade. Os autores relatam melhorias em abrangência e diversidade das respostas em conjuntos de dados de cerca de um milhão de tokens. Vale a ressalva: o GraphRAG foi desenhado para perguntas globais de síntese, e não para o problema de resolver qual decisão está vigente. Ele mostra que dar estrutura ao acervo pode ajudar, mas não demonstra que resolve o nosso caso. Essa continua sendo uma hipótese a testar.&lt;/p&gt;

&lt;p&gt;Em resumo, a pergunta do artigo ganha duas camadas. A primeira é a de H1: em qual categoria a falha está? A segunda é a de H2: se for conhecimento, o acervo está curado o bastante para ser usado? Só depois dessas duas respostas faz sentido discutir RAG, fine-tuning ou ferramentas.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Três eixos: conhecimento, comportamento e execução
&lt;/h2&gt;

&lt;p&gt;Se H1 for útil, ela precisa de categorias que um time consiga aplicar a uma falha real, sem depender de interpretação. Proponho três eixos, cada um definido por uma pergunta que se pode fazer ao sistema.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw0h17at19s6e24d9rqy2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw0h17at19s6e24d9rqy2.png" alt="Três tipos de problema, três capacidades: conhecimento (RAG), comportamento (fine-tuning) e execução (ferramentas)." width="800" height="440"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figura 2. Cada tipo de falha aponta para uma capacidade diferente.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Conhecimento: o modelo tem acesso à informação certa?
&lt;/h3&gt;

&lt;p&gt;É o eixo da recomendação desatualizada. O sintoma típico é uma resposta que cita uma decisão inexistente, vencida ou substituída. O teste diagnóstico é simples: se colocarmos manualmente o ADR vigente no contexto, o modelo acerta? Se acerta, o problema não é de capacidade do modelo, e sim de acesso à informação, e a conversa passa a ser sobre recuperação e curadoria (H2). Esse é o território que o &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;guia da OpenAI&lt;/a&gt; atribui à otimização de contexto: o caso em que o modelo não tem conhecimento por ele não estar nos dados de treinamento. A intervenção natural é o RAG, com a ressalva de que ele próprio tem &lt;a href="https://arxiv.org/abs/2401.05856" rel="noopener noreferrer"&gt;vários pontos de falha, da indexação à geração&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Comportamento: tendo a informação, o modelo a usa do jeito esperado?
&lt;/h3&gt;

&lt;p&gt;Aqui o modelo tem o ADR certo no contexto e ainda assim responde mal. Ele não segue o formato exigido pelo time, não cita o status da decisão, mistura recomendação com opinião ou perde a estrutura esperada em respostas longas. O teste diagnóstico muda: com contexto correto e instrução clara, o erro de forma persiste de modo consistente? Se persistir, e só depois de esgotar instrução e exemplos, o fine-tuning passa a ser um candidato. O mesmo &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;guia da OpenAI&lt;/a&gt; o associa a resultados inconsistentes ou com formato incorreto. Há uma sustentação conceitual para essa divisão: &lt;a href="https://arxiv.org/abs/2405.05904" rel="noopener noreferrer"&gt;Gekhman et al.&lt;/a&gt; concluem que os modelos adquirem a maior parte do conhecimento factual no pré-treino, e que o fine-tuning ensina principalmente a usá-lo melhor. O &lt;a href="https://arxiv.org/abs/2403.10131" rel="noopener noreferrer"&gt;RAFT&lt;/a&gt; mostra um uso concreto: treinar o modelo para ignorar documentos irrelevantes recuperados e citar os trechos relevantes, ou seja, o fine-tuning moldando como o modelo usa a recuperação, e não o que ele sabe.&lt;/p&gt;

&lt;h3&gt;
  
  
  Execução: a tarefa exige consultar um sistema ou agir?
&lt;/h3&gt;

&lt;p&gt;Algumas perguntas não deveriam ser respondidas por texto, nem do modelo nem de um documento. Se a dúvida é "esta API está obsoleta hoje?", a fonte de verdade é o catálogo de APIs, não um ADR escrito meses atrás. O sintoma é uma resposta que mistura memória do modelo ou texto desatualizado com um dado que existe em um sistema consultável. A intervenção é uma ferramenta: o modelo chama o catálogo, recebe o status atual e responde com base nele. Trabalhos como o &lt;a href="https://arxiv.org/abs/2210.03629" rel="noopener noreferrer"&gt;ReAct&lt;/a&gt;, que combina raciocínio e ação, e o &lt;a href="https://arxiv.org/abs/2302.04761" rel="noopener noreferrer"&gt;Toolformer&lt;/a&gt;, em que o modelo aprende quando e como chamar APIs como calculadora e busca, mostram que ferramentas atacam o que o modelo não consegue manter por conta própria. Há uma fronteira porosa com o eixo do conhecimento, porque uma consulta a um catálogo também é acesso à informação. A diferença prática está na fonte: um acervo de documentos que precisa ser curado, de um lado, e um sistema autoritativo, consultável e atual, do outro. Ferramentas também trazem riscos próprios. A &lt;a href="https://modelcontextprotocol.io/specification/latest" rel="noopener noreferrer"&gt;especificação do MCP&lt;/a&gt;, por exemplo, trata ferramentas como execução arbitrária de código e orienta tratar descrições de comportamento de ferramentas como não confiáveis, a menos que venham de uma fonte confiável.&lt;/p&gt;

&lt;h3&gt;
  
  
  O que fica fora dos três eixos
&lt;/h3&gt;

&lt;p&gt;Os três eixos não cobrem tudo, e é importante dizer isso. Uma instrução ambígua e um raciocínio insuficiente também geram respostas erradas, e nenhuma das três capacidades resolve isso. Por isso a árvore de decisão da seção 7 começa por um passo zero, antes de qualquer técnica: melhorar a especificação da tarefa, decompô-la ou usar um modelo mais capaz. Isso também é coerente com o &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;conselho da OpenAI&lt;/a&gt; de começar por engenharia de prompt. Essa limitação é uma das razões pelas quais H1 fala em "grande parte das falhas", e não em todas.&lt;/p&gt;

&lt;p&gt;Há ainda um ponto sobre combinações. Os eixos descrevem a causa de uma falha, não uma escolha exclusiva de solução. O RAFT é um exemplo de fine-tuning (comportamento) trabalhando a serviço da recuperação (conhecimento), e o mesmo &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;guia da OpenAI&lt;/a&gt; descreve as abordagens como aditivas, não exclusivas. O ganho de pensar em eixos é saber qual alavanca puxar primeiro e como saber se ela funcionou.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. A arquitetura
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F12ucjkklqnhpsxxrk810.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F12ucjkklqnhpsxxrk810.png" alt="Arquitetura: pergunta, recuperação com base de conhecimento, modelo base ou adaptado, ferramentas e resposta." width="799" height="453"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figura 3. O modelo é um componente. Conhecimento, comportamento e ações têm lugares diferentes no sistema.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Os três eixos da seção anterior ganham sentido prático quando os colocamos em uma arquitetura. A Figura 3 mostra o assistente de ADRs do nosso cenário e indica onde cada capacidade vive. A ideia central é que o modelo é um componente entre vários, e que cada eixo corresponde a um lugar diferente do sistema onde se pode intervir.&lt;/p&gt;

&lt;p&gt;Siga o caminho de uma pergunta como "qual padrão de autenticação devemos usar entre serviços?".&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Recuperação.&lt;/strong&gt; A pergunta passa por um bloco que busca nos documentos e aplica filtros, por exemplo descartar ADRs com status substituído. É aqui que o eixo do &lt;strong&gt;conhecimento&lt;/strong&gt; se manifesta no tempo de execução.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Base de conhecimento.&lt;/strong&gt; Por trás da recuperação está o acervo, e é nele que a curadoria da seção 2 acontece: status, relações, fonte canônica. Uma recuperação ótima sobre um acervo mal curado continua devolvendo respostas desatualizadas.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Modelo.&lt;/strong&gt; Recebe a pergunta e o contexto e gera a resposta. O eixo do &lt;strong&gt;comportamento&lt;/strong&gt; vive aqui: instrução, exemplos e, quando justificado, um modelo adaptado por fine-tuning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ferramentas.&lt;/strong&gt; Em vez de confiar em texto, o modelo pode consultar um sistema autoritativo, como o catálogo de APIs, e voltar com o dado atual. É o eixo da &lt;strong&gt;execução&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resposta.&lt;/strong&gt; O resultado volta ao usuário, idealmente com a fonte e o status das decisões usadas.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Essa leitura é coerente com uma tendência descrita na literatura de engenharia de IA. O grupo de Berkeley, em um &lt;a href="https://bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/" rel="noopener noreferrer"&gt;post que descreve os "compound AI systems"&lt;/a&gt;, argumenta que resultados de ponta vêm cada vez mais de sistemas com vários componentes, como chamadas ao modelo, recuperadores e ferramentas, e não de um modelo isolado, e que iterar no sistema costuma ser mais rápido do que retreinar o modelo. O &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;guia da Anthropic sobre agentes&lt;/a&gt; parte da mesma ideia e recomenda buscar a solução mais simples possível, lembrando que muitas aplicações precisam apenas de uma única chamada bem otimizada ao modelo. A arquitetura da Figura 3 é, portanto, um teto de complexidade, não uma meta. Nem todo sistema precisa de todos os blocos, e começar sem eles é uma escolha legítima.&lt;/p&gt;

&lt;p&gt;Duas observações sobre o diagrama. A primeira é que os blocos têm donos e ciclos de vida diferentes. O acervo muda quando uma decisão muda, o modelo muda quando o provedor muda ou quando há retreino, e as ferramentas mudam com os sistemas que elas consultam. Misturar essas camadas em uma só, por exemplo treinando o modelo com conteúdo que deveria viver no acervo, é o tipo de decisão que a seção 1 chamou de "complexidade que passa a ser sua". A segunda é que a figura omite, de propósito, a camada que atravessa todos os blocos: avaliação e observabilidade. Sem registrar o que foi recuperado, o que o modelo recebeu e qual ferramenta foi chamada, o diagnóstico das seções anteriores não é possível, e voltamos a escolher técnica por opinião. Esse tema merece artigo próprio dentro da série.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. O que diz a evidência
&lt;/h2&gt;

&lt;p&gt;Antes de propor uma árvore de decisão, vale perguntar o que a pesquisa já mostrou. A resposta curta é que existe evidência útil, mas ela é mais limitada do que os debates fazem parecer. Os estudos abaixo comparam as duas abordagens para injetar conhecimento. Esta seção se baseia nos resumos (abstracts) dos estudos, não nos artigos completos, então só cito os resultados que os próprios resumos declaram.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Estudo&lt;/th&gt;
&lt;th&gt;Cenário&lt;/th&gt;
&lt;th&gt;O que os autores relatam&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2312.05934" rel="noopener noreferrer"&gt;Ovadia et al. (2023)&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Tarefas intensivas em conhecimento, com conhecimento já visto e novo&lt;/td&gt;
&lt;td&gt;Fine-tuning não supervisionado melhora um pouco, mas o RAG é melhor em ambos os casos. Os modelos têm dificuldade em aprender fatos novos por esse caminho&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2403.01432" rel="noopener noreferrer"&gt;Soudani et al. (2024)&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Perguntas e respostas sintéticas sobre conhecimento pouco popular, em doze modelos&lt;/td&gt;
&lt;td&gt;O fine-tuning melhora em todos os níveis de popularidade, mas o RAG o supera por ampla margem, principalmente nos fatos menos populares&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2401.08406" rel="noopener noreferrer"&gt;Balaguer et al. (2024), Microsoft&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Estudo de caso em agricultura, com Llama2-13B, GPT-3.5 e GPT-4&lt;/td&gt;
&lt;td&gt;O fine-tuning elevou a acurácia em mais de 6 pontos percentuais, e o RAG somou cerca de 5 pontos por cima. Os efeitos se acumularam&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2403.10131" rel="noopener noreferrer"&gt;Zhang et al. (2024), RAFT&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Adaptação a domínios específicos com RAG&lt;/td&gt;
&lt;td&gt;Treinar o modelo para ignorar documentos irrelevantes e citar os relevantes trouxe ganhos consistentes em três conjuntos de dados&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2404.00213" rel="noopener noreferrer"&gt;Mecklenburg et al. (2024), Microsoft&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Eventos esportivos recentes&lt;/td&gt;
&lt;td&gt;O fine-tuning supervisionado consegue injetar fatos, desde que os dados de treino sejam bem construídos. A cobertura foi desigual com escala por tokens e mais uniforme com escala por fatos&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2601.07054" rel="noopener noreferrer"&gt;Yang et al. (2026)&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Perguntas de múltiplos saltos sobre conhecimento novo, em três modelos de 7B&lt;/td&gt;
&lt;td&gt;O RAG deu ganhos consistentes, mas o fine-tuning supervisionado teve a maior acurácia geral&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2508.06178" rel="noopener noreferrer"&gt;Abonizio et al. (2025)&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Injeção de conhecimento com poucos dados&lt;/td&gt;
&lt;td&gt;A injeção por RAG tendeu a degradar o desempenho em conjuntos de controle mais do que os métodos paramétricos&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;O que dá para afirmar com cuidado.&lt;/strong&gt; Para fatos novos ou raros, o fine-tuning não supervisionado sobre texto bruto é um caminho fraco, e o RAG foi superior em todas as comparações diretas que li. Ao mesmo tempo, o fine-tuning supervisionado com &lt;a href="https://arxiv.org/abs/2404.00213" rel="noopener noreferrer"&gt;dados bem construídos&lt;/a&gt; consegue acrescentar conhecimento e, em um &lt;a href="https://arxiv.org/abs/2601.07054" rel="noopener noreferrer"&gt;estudo de 2026&lt;/a&gt;, teve a melhor acurácia geral. E os efeitos podem se somar, como no &lt;a href="https://arxiv.org/abs/2401.08406" rel="noopener noreferrer"&gt;caso da agricultura&lt;/a&gt;. A leitura mais honesta é de que as técnicas atuam em lugares diferentes, o que sustenta H1, e não de que uma vence a outra.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;O que a evidência não diz.&lt;/strong&gt; Há quatro limites que o leitor deve ter em mente:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Qualidade da evidência.&lt;/strong&gt; Quase todos os estudos acima são preprints, e a maioria usa benchmarks pequenos ou sintéticos, como agricultura, eventos esportivos e entidades inventadas. Não há razão para supor que os resultados se transfiram para um acervo corporativo.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nenhum deles testa o nosso problema.&lt;/strong&gt; Os estudos medem se o modelo aprende ou recupera fatos. Nenhum mede a distinção entre decisão vigente e obsoleta, que é o centro do cenário do artigo, e portanto nenhum testa H2.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RAG não é sempre melhor.&lt;/strong&gt; Além do &lt;a href="https://arxiv.org/abs/2508.06178" rel="noopener noreferrer"&gt;estudo de Abonizio&lt;/a&gt;, o exemplo do &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;guia da OpenAI&lt;/a&gt; mostra um caso em que o RAG reduziu a nota por ruído. O ganho depende de a recuperação trazer o conteúdo certo.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Efeitos de fine-tuning dependem do cenário.&lt;/strong&gt; O aumento de alucinação relatado por &lt;a href="https://arxiv.org/abs/2405.05904" rel="noopener noreferrer"&gt;Gekhman et al.&lt;/a&gt; foi medido em perguntas e respostas sem consulta externa, com conhecimento novo. Não se estende a fine-tuning para formato ou estilo.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Em resumo, a pesquisa justifica a tese de que as duas abordagens não competem no mesmo terreno, mas não substitui um teste no seu próprio acervo e com as suas falhas reais. É por isso que um dos próximos artigos da série trata do experimento.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Os trade-offs
&lt;/h2&gt;

&lt;p&gt;Escolher uma capacidade é aceitar um conjunto de custos que não aparecem no slide da proposta. A tabela abaixo compara as três capacidades em seis dimensões, aplicadas ao assistente de ADRs. É uma comparação &lt;strong&gt;qualitativa&lt;/strong&gt;: este artigo não mediu latência nem custo, e deixo essas medições para um artigo futuro da série, em vez de repetir números de origem duvidosa.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimensão&lt;/th&gt;
&lt;th&gt;Recuperação (RAG)&lt;/th&gt;
&lt;th&gt;Fine-tuning&lt;/th&gt;
&lt;th&gt;Ferramentas&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Atualização&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Reindexar o que mudou. Quando um ADR é substituído, basta atualizar o acervo&lt;/td&gt;
&lt;td&gt;Retreinar para refletir a mudança, e reavaliar&lt;/td&gt;
&lt;td&gt;Imediata: a ferramenta consulta o sistema atual&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Custo recorrente&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Infraestrutura de indexação e mais texto de contexto a cada consulta&lt;/td&gt;
&lt;td&gt;Dados de treino, treinamento, avaliação e nova rodada a cada mudança de modelo base&lt;/td&gt;
&lt;td&gt;Integração e manutenção de cada ferramenta e do seu contrato&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Latência&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Acrescenta pelo menos uma etapa de busca antes da geração&lt;/td&gt;
&lt;td&gt;Em geral não acrescenta etapas na inferência&lt;/td&gt;
&lt;td&gt;Acrescenta idas e voltas a cada chamada&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Auditabilidade&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Permite apontar a fonte, mas a citação precisa ser verificada&lt;/td&gt;
&lt;td&gt;Conhecimento difuso nos pesos, difícil de rastrear ou remover&lt;/td&gt;
&lt;td&gt;Chamadas registráveis e reproduzíveis&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Segurança&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Permissões e vazamento entre contextos precisam ser tratados na recuperação&lt;/td&gt;
&lt;td&gt;Dados de treino passam a fazer parte do modelo&lt;/td&gt;
&lt;td&gt;Ferramentas são execução de código e exigem confiança e limites&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Manutenção&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Curadoria contínua do acervo (H2)&lt;/td&gt;
&lt;td&gt;Atrelamento ao modelo base e ao ciclo de vida do provedor&lt;/td&gt;
&lt;td&gt;Versionamento e compatibilidade das integrações&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Alguns pontos da tabela merecem uma explicação.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Atualização e custo seguem a velocidade da mudança.&lt;/strong&gt; No nosso cenário, decisões de arquitetura mudam com frequência razoável. Um ADR substituído hoje exige, no RAG, atualizar o acervo, e no fine-tuning, retreinar e reavaliar o modelo para que ele esqueça a decisão anterior. Quanto mais volátil o conteúdo, mais o custo recorrente do fine-tuning pesa. Por outro lado, métodos eficientes em parâmetros reduzem o custo do treino em si: o &lt;a href="https://arxiv.org/abs/2106.09685" rel="noopener noreferrer"&gt;LoRA&lt;/a&gt; relata cerca de 10.000 vezes menos parâmetros treináveis e cerca de 3 vezes menos memória de GPU, em comparação com o ajuste completo do GPT-3 175B com Adam. Isso baixa a barreira de entrada, mas não elimina o custo de dados, avaliação e manutenção.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Auditabilidade não é o mesmo que verificabilidade.&lt;/strong&gt; O RAG permite mostrar de onde veio a resposta, e isso é uma vantagem real sobre conhecimento difuso nos pesos. Mas citar uma fonte não garante que a fonte sustente a afirmação. Em uma &lt;a href="https://arxiv.org/abs/2304.09848" rel="noopener noreferrer"&gt;auditoria humana de buscadores generativos&lt;/a&gt;, apenas 51,5% das frases estavam totalmente sustentadas pelas citações, e 74,5% das citações sustentavam a frase correspondente. São sistemas de 2023 e não assistentes corporativos, então o número não se transfere, mas o aviso permanece: uma citação precisa ser verificada, não apenas exibida.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Segurança entra pela recuperação.&lt;/strong&gt; O OWASP trata vetores e embeddings como &lt;a href="https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/" rel="noopener noreferrer"&gt;risco próprio no seu Top 10 de 2025 (LLM08)&lt;/a&gt;, inclusive o vazamento de contexto entre usuários em bancos vetoriais compartilhados, e recomenda armazenamento com controle de acesso por permissão e registros imutáveis das recuperações. Em um assistente de ADRs, nem todo ADR deve ser visível a todos os times. Esse tema merece um artigo próprio na série.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fine-tuning prende você ao ciclo de vida do modelo base.&lt;/strong&gt; Provedores aposentam modelos. A Anthropic informa um aviso de &lt;a href="https://platform.claude.com/docs/en/about-claude/model-deprecations" rel="noopener noreferrer"&gt;pelo menos 60 dias&lt;/a&gt; antes da aposentadoria de modelos públicos, e a OpenAI, de &lt;a href="https://developers.openai.com/api/docs/deprecations" rel="noopener noreferrer"&gt;pelo menos 6 meses&lt;/a&gt; para modelos de disponibilidade geral e cerca de 2 semanas para prévias, segundo as páginas consultadas em outubro de 2026. Um modelo ajustado sobre uma base que sai de linha precisa ser refeito. É uma inferência do autor, não uma regra documentada para todos os casos, mas é um custo que merece entrar na conta desde o início.&lt;/p&gt;

&lt;p&gt;O ponto de equilíbrio entre essas dimensões depende do caso. Em geral, e esta é uma regra de bolso do autor e não um resultado de pesquisa, um acervo estável, com latência apertada e formato rígido de saída, pende para o fine-tuning, enquanto um acervo que muda toda semana e exige rastreabilidade pende para recuperação e ferramentas. O que a tabela entrega é a lista de perguntas a fazer antes de decidir, e é com ela que a árvore de decisão da próxima seção é construída.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. O framework de decisão
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuofdg16nd2a0fygpoz5h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuofdg16nd2a0fygpoz5h.png" alt="Árvore de decisão: da falha à intervenção. Passo 0: instrução ambígua ou raciocínio fraco levam à especificação da tarefa, decomposição ou modelo mais capaz. Passo 1: informação privada ou mutável leva a RAG e curadoria. Passo 2: forma e estilo levam a fine-tuning. Passo 3: cálculo ou ação levam a ferramentas." width="800" height="553"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figura 4. Da falha à intervenção. Síntese do autor, para validar com dados próprios.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A Figura 4 transforma as seções anteriores em um procedimento. Antes de usá-la, uma ressalva: ela é uma &lt;strong&gt;síntese do autor&lt;/strong&gt;, construída a partir dos eixos da seção 3 e da &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;divisão que a OpenAI faz entre contexto e comportamento&lt;/a&gt;. Não é uma taxonomia publicada nem foi validada empiricamente. Trate-a como ponto de partida para o seu time, não como regra.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Como usar.&lt;/strong&gt; A árvore parte de uma falha concreta, já observada e registrada, nunca de uma preferência. A árvore tem um passo zero e três passos. Para cada pergunta, o teste diagnóstico da seção 3 diz como responder com evidência em vez de opinião:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;em&gt;A instrução é ambígua ou o raciocínio é insuficiente?&lt;/em&gt; Ajuste a especificação da tarefa (instrução, critérios de sucesso e formato de saída), decomponha a tarefa ou considere um modelo mais capaz. É o passo mais barato de verificar, e o &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;guia da OpenAI&lt;/a&gt; recomenda começar pela engenharia de prompt.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;A resposta depende de informação privada, recente ou mutável?&lt;/em&gt; Faça o teste do contexto manual: com a informação certa à mão, o modelo acerta? Se sim, o caminho é recuperação, e a primeira checagem é a curadoria do acervo (H2) antes de mexer no retrieval.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;O conteúdo está certo, mas forma, estilo ou formato falham?&lt;/em&gt; Se o erro persiste com contexto e instrução corretos, fine-tuning é candidato, mas só depois de provar que instrução e exemplos não bastam. Esse "provar" é o tema de um dos próximos artigos da série.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;A tarefa exige cálculo, consulta a um sistema ou ação?&lt;/em&gt; Prefira uma ferramenta que consulte a fonte autoritativa a uma resposta baseada em texto.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;O passo zero vem primeiro porque corrigir a especificação da tarefa é o mais barato e porque nenhuma capacidade compensa uma tarefa mal especificada. Depois de cada correção, reavalie a falha original antes de seguir. E se nenhuma pergunta receber "sim", a conclusão é voltar ao diagnóstico: a falha pode ter mais de uma causa, ou a causa pode ainda não ter sido identificada.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Aplicando ao nosso cenário.&lt;/strong&gt; A tabela mostra como quatro falhas hipotéticas do assistente de ADRs percorreriam a árvore. São exemplos ilustrativos, não resultados.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Falha observada&lt;/th&gt;
&lt;th&gt;Teste diagnóstico&lt;/th&gt;
&lt;th&gt;Caminho na árvore&lt;/th&gt;
&lt;th&gt;Intervenção&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;O ADR que substituiu o padrão antigo nunca foi indexado&lt;/td&gt;
&lt;td&gt;Com o ADR novo no contexto, o modelo acerta&lt;/td&gt;
&lt;td&gt;Passo 1: sim&lt;/td&gt;
&lt;td&gt;Indexar e corrigir o processo de ingestão&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;O ADR antigo, "superseded", é recuperado no lugar do novo&lt;/td&gt;
&lt;td&gt;O modelo acerta quando só o ADR vigente é fornecido&lt;/td&gt;
&lt;td&gt;Passo 1: sim&lt;/td&gt;
&lt;td&gt;Curadoria: status e relação de substituição, filtro na recuperação&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;O modelo recebe o ADR certo e ainda recomenda o padrão antigo&lt;/td&gt;
&lt;td&gt;Com contexto e instrução claros, o erro persiste, ou some com uma instrução melhor?&lt;/td&gt;
&lt;td&gt;Passo 0 primeiro; depois passo 2&lt;/td&gt;
&lt;td&gt;Reforçar a instrução de priorizar a decisão vigente; fine-tuning só se persistir&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A pergunta é se uma API ainda está ativa hoje&lt;/td&gt;
&lt;td&gt;A fonte de verdade é o catálogo, não um documento&lt;/td&gt;
&lt;td&gt;Passo 3: sim&lt;/td&gt;
&lt;td&gt;Ferramenta que consulta o catálogo de APIs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Observe o que as quatro linhas têm em comum: nenhuma pede, como primeiro passo, que se escolha entre RAG e fine-tuning. Três delas se resolvem sem tocar nos pesos do modelo. É exatamente o que a tese do artigo antecipa, mas vale repetir o limite: são quatro cenários escolhidos para ilustrar o procedimento, e não uma estimativa de com que frequência cada causa aparece na prática.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Checklist antes de agir.&lt;/strong&gt; Em qualquer equipe, antes de aprovar uma intervenção, vale conseguir responder:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Qual é a falha, e ela está registrada com entrada, saída esperada e saída obtida?&lt;/li&gt;
&lt;li&gt;Qual teste diagnóstico confirmou a categoria da causa?&lt;/li&gt;
&lt;li&gt;Qual é a alternativa mais barata que já foi tentada?&lt;/li&gt;
&lt;li&gt;Como saberemos que a intervenção funcionou, com qual conjunto de avaliação e qual critério definido antes de rodar?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  8. O princípio de engenharia
&lt;/h2&gt;

&lt;p&gt;Voltemos, uma última vez, ao time que viu o assistente recomendar um padrão arquitetural abandonado. Na versão da conversa que começa pela técnica, eles saem com uma decisão: fine-tuning ou RAG. Na versão que começa pela falha, saem com uma pergunta respondida: o que, exatamente, deu errado nessa resposta? A segunda versão é mais lenta no primeiro dia e mais barata em todos os outros.&lt;/p&gt;

&lt;p&gt;O princípio que este artigo propõe cabe em uma frase: &lt;strong&gt;diagnostique a falha antes de escolher a técnica.&lt;/strong&gt; Para aplicá-lo, três ideias:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Separe saber, comportar-se e fazer.&lt;/strong&gt; Conhecimento, comportamento e execução são problemas diferentes (H1), e cada um aponta para uma capacidade diferente: recuperação, adaptação do modelo ou ferramentas. Antes deles, confirme que a tarefa está bem especificada.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cure o acervo antes de culpar o modelo.&lt;/strong&gt; Uma parte das falhas de conhecimento é falha de curadoria (H2): o acervo contém, por desenho, o vigente e o obsoleto, e nenhuma técnica sobre o modelo distingue um do outro sem status, datas e relações explícitas.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Exija evidência antes de intervir.&lt;/strong&gt; Registre a falha, confirme a categoria com um teste diagnóstico, tente a alternativa mais barata e defina, antes de rodar, como saberá que funcionou.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Uma ressalva final, por coerência com o que o artigo pede. H1 e H2 são &lt;strong&gt;hipóteses de trabalho, não resultados&lt;/strong&gt;: a pesquisa que revisamos mostra que as técnicas atuam em lugares diferentes, mas nenhum dos estudos testa a distinção entre decisão vigente e obsoleta, e a árvore de decisão é uma síntese do autor. O valor dela está em oferecer um procedimento que o seu time pode testar e corrigir com os seus próprios dados, não em substituir esse teste.&lt;/p&gt;

&lt;p&gt;A pergunta "RAG ou fine-tuning?" não está errada por ser inútil. Está errada por chegar cedo demais. Quando o diagnóstico vem antes, ela volta como uma pergunta melhor: &lt;em&gt;qual destas capacidades a minha falha está pedindo?&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A série Beyond the Model
&lt;/h2&gt;

&lt;p&gt;Este é o primeiro artigo da série &lt;strong&gt;Beyond the Model&lt;/strong&gt;, sobre as decisões que transformam modelos de linguagem em sistemas de software confiáveis. Nos próximos, a série aprofunda os dois pontos que ficaram em aberto aqui:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stop Adding RAG to Everything&lt;/strong&gt; detalha a taxonomia de falhas: como uma resposta errada se classifica em conhecimento, recuperação, instrução, raciocínio, ferramentas ou comportamento, e por que cada causa aponta para uma intervenção diferente.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Before You Fine-Tune an LLM, Prove That You Need To&lt;/strong&gt; mostra como transformar a decisão em experimento: baseline, conjunto de avaliação, métricas e critério de decisão definidos antes de treinar.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Se você já passou por uma conversa como a do início deste artigo, conte nos comentários qual foi a falha que motivou a discussão e qual intervenção o time escolheu. Esses casos são exatamente o material de que os próximos artigos precisam.&lt;/p&gt;

&lt;h2&gt;
  
  
  Referências
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Barnett, S. et al. (2024). &lt;a href="https://arxiv.org/abs/2401.05856" rel="noopener noreferrer"&gt;&lt;em&gt;Seven Failure Points When Engineering a Retrieval Augmented Generation System&lt;/em&gt;&lt;/a&gt;. arXiv:2401.05856&lt;/li&gt;
&lt;li&gt;Lewis, P. et al. (2020). &lt;a href="https://arxiv.org/abs/2005.11401" rel="noopener noreferrer"&gt;&lt;em&gt;Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks&lt;/em&gt;&lt;/a&gt;. NeurIPS 2020&lt;/li&gt;
&lt;li&gt;Hu, E. J. et al. (2021). &lt;a href="https://arxiv.org/abs/2106.09685" rel="noopener noreferrer"&gt;&lt;em&gt;LoRA: Low-Rank Adaptation of Large Language Models&lt;/em&gt;&lt;/a&gt;. arXiv:2106.09685&lt;/li&gt;
&lt;li&gt;Ovadia, O. et al. (2023). &lt;a href="https://arxiv.org/abs/2312.05934" rel="noopener noreferrer"&gt;&lt;em&gt;Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs&lt;/em&gt;&lt;/a&gt;. arXiv:2312.05934&lt;/li&gt;
&lt;li&gt;OpenAI. &lt;a href="https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy" rel="noopener noreferrer"&gt;&lt;em&gt;Optimizing LLM Accuracy&lt;/em&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gekhman, Z. et al. (2024). &lt;a href="https://arxiv.org/abs/2405.05904" rel="noopener noreferrer"&gt;&lt;em&gt;Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?&lt;/em&gt;&lt;/a&gt;. EMNLP 2024&lt;/li&gt;
&lt;li&gt;Yang et al. (2026). &lt;a href="https://arxiv.org/abs/2601.07054" rel="noopener noreferrer"&gt;&lt;em&gt;Fine-Tuning vs. RAG for Multi-Hop Question Answering with Novel Knowledge&lt;/em&gt;&lt;/a&gt;. arXiv:2601.07054 (preprint)&lt;/li&gt;
&lt;li&gt;Nygard, M. (2011). &lt;a href="https://www.cognitect.com/blog/2011/11/15/documenting-architecture-decisions" rel="noopener noreferrer"&gt;&lt;em&gt;Documenting Architecture Decisions&lt;/em&gt;&lt;/a&gt;. Cognitect blog, 15 Nov 2011&lt;/li&gt;
&lt;li&gt;Gruber, T. R. (1993). &lt;a href="https://tomgruber.org/writing/ontolingua-kaj-1993/" rel="noopener noreferrer"&gt;&lt;em&gt;A Translation Approach to Portable Ontology Specifications&lt;/em&gt;&lt;/a&gt;. Knowledge Acquisition, 5(2), 199–220&lt;/li&gt;
&lt;li&gt;Pan, S. et al. (2024). &lt;a href="https://arxiv.org/abs/2306.08302" rel="noopener noreferrer"&gt;&lt;em&gt;Unifying Large Language Models and Knowledge Graphs: A Roadmap&lt;/em&gt;&lt;/a&gt;. IEEE TKDE (arXiv:2306.08302)&lt;/li&gt;
&lt;li&gt;Edge, D. et al. (2024). &lt;a href="https://arxiv.org/abs/2404.16130" rel="noopener noreferrer"&gt;&lt;em&gt;From Local to Global: A Graph RAG Approach to Query-Focused Summarization&lt;/em&gt;&lt;/a&gt;. arXiv:2404.16130&lt;/li&gt;
&lt;li&gt;Zhang, T. et al. (2024). &lt;a href="https://arxiv.org/abs/2403.10131" rel="noopener noreferrer"&gt;&lt;em&gt;RAFT: Adapting Language Model to Domain Specific RAG&lt;/em&gt;&lt;/a&gt;. arXiv:2403.10131 (preprint)&lt;/li&gt;
&lt;li&gt;Yao, S. et al. (2022). &lt;a href="https://arxiv.org/abs/2210.03629" rel="noopener noreferrer"&gt;&lt;em&gt;ReAct: Synergizing Reasoning and Acting in Language Models&lt;/em&gt;&lt;/a&gt;. arXiv:2210.03629 (ICLR 2023)&lt;/li&gt;
&lt;li&gt;Schick, T. et al. (2023). &lt;a href="https://arxiv.org/abs/2302.04761" rel="noopener noreferrer"&gt;&lt;em&gt;Toolformer: Language Models Can Teach Themselves to Use Tools&lt;/em&gt;&lt;/a&gt;. arXiv:2302.04761&lt;/li&gt;
&lt;li&gt;Model Context Protocol. &lt;a href="https://modelcontextprotocol.io/specification/latest" rel="noopener noreferrer"&gt;&lt;em&gt;Specification&lt;/em&gt;&lt;/a&gt;. (revisão 2026-07-28)&lt;/li&gt;
&lt;li&gt;Zaharia, M. et al. (2024). &lt;a href="https://bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/" rel="noopener noreferrer"&gt;&lt;em&gt;The Shift from Models to Compound AI Systems&lt;/em&gt;&lt;/a&gt;. BAIR Blog, 18 Feb 2024&lt;/li&gt;
&lt;li&gt;Anthropic (2024). &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;&lt;em&gt;Building Effective Agents&lt;/em&gt;&lt;/a&gt;. 19 Dec 2024&lt;/li&gt;
&lt;li&gt;Soudani, H., Kanoulas, E., Hasibi, F. (2024). &lt;a href="https://arxiv.org/abs/2403.01432" rel="noopener noreferrer"&gt;&lt;em&gt;Fine Tuning vs. Retrieval Augmented Generation for Less Popular Knowledge&lt;/em&gt;&lt;/a&gt;. arXiv:2403.01432&lt;/li&gt;
&lt;li&gt;Balaguer, A. et al. (2024). &lt;a href="https://arxiv.org/abs/2401.08406" rel="noopener noreferrer"&gt;&lt;em&gt;RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture&lt;/em&gt;&lt;/a&gt;. arXiv:2401.08406 (preprint)&lt;/li&gt;
&lt;li&gt;Abonizio, H., Almeida, T. S., Lotufo, R., Nogueira, R. (2025). &lt;a href="https://arxiv.org/abs/2508.06178" rel="noopener noreferrer"&gt;&lt;em&gt;Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime&lt;/em&gt;&lt;/a&gt;. arXiv:2508.06178 (preprint)&lt;/li&gt;
&lt;li&gt;Mecklenburg, N. et al. (2024). &lt;a href="https://arxiv.org/abs/2404.00213" rel="noopener noreferrer"&gt;&lt;em&gt;Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning&lt;/em&gt;&lt;/a&gt;. arXiv:2404.00213 (preprint)&lt;/li&gt;
&lt;li&gt;Liu, N. F., Zhang, T., Liang, P. (2023). &lt;a href="https://arxiv.org/abs/2304.09848" rel="noopener noreferrer"&gt;&lt;em&gt;Evaluating Verifiability in Generative Search Engines&lt;/em&gt;&lt;/a&gt;. Findings of EMNLP 2023 (arXiv:2304.09848)&lt;/li&gt;
&lt;li&gt;OWASP GenAI Security Project. &lt;a href="https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/" rel="noopener noreferrer"&gt;&lt;em&gt;LLM08:2025 Vector and Embedding Weaknesses&lt;/em&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Anthropic. &lt;a href="https://platform.claude.com/docs/en/about-claude/model-deprecations" rel="noopener noreferrer"&gt;&lt;em&gt;Model deprecations&lt;/em&gt;&lt;/a&gt;. (consultado em 2026-10-10)&lt;/li&gt;
&lt;li&gt;OpenAI. &lt;a href="https://developers.openai.com/api/docs/deprecations" rel="noopener noreferrer"&gt;&lt;em&gt;Deprecations&lt;/em&gt;&lt;/a&gt;. (consultado em 2026-10-10)&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>llm</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
