<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: TankDev-Tech</title>
    <description>The latest articles on DEV Community by TankDev-Tech (tankdev-tech).</description>
    <link>https://dev.to/tankdev-tech</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F14953%2F0ed9884e-1918-4276-96a0-7687020ea67c.png</url>
      <title>DEV Community: TankDev-Tech</title>
      <link>https://dev.to/tankdev-tech</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tankdev-tech"/>
    <language>en</language>
    <item>
      <title>RAG vs. Fine-Tuning: How to Choose the Right Technique for Enterprise LLM Systems</title>
      <dc:creator>Murat Onur Kaderoglu</dc:creator>
      <pubDate>Wed, 30 Sep 2026 11:00:00 +0000</pubDate>
      <link>https://dev.to/tankdev-tech/rag-vs-fine-tuning-how-to-choose-the-right-technique-for-enterprise-llm-systems-49ej</link>
      <guid>https://dev.to/tankdev-tech/rag-vs-fine-tuning-how-to-choose-the-right-technique-for-enterprise-llm-systems-49ej</guid>
      <description>&lt;h1&gt;
  
  
  RAG vs. Fine-Tuning: How to Choose the Right Technique for Enterprise LLM Systems
&lt;/h1&gt;

&lt;p&gt;Choose between RAG and fine-tuning through knowledge freshness, behaviour, data security, cost, and evaluation.&lt;/p&gt;

&lt;p&gt;RAG and fine-tuning are often presented as competing answers to one question: should we give company documents to the model, or train the model on company data? That framing is incomplete. RAG changes which &lt;strong&gt;knowledge&lt;/strong&gt; is available at answer time. Fine-tuning changes how a model &lt;strong&gt;behaves&lt;/strong&gt; on a task. A sound production system chooses where rules, RAG, fine-tuning, and human review are actually needed.&lt;/p&gt;

&lt;p&gt;This is TankDev's decision framework: retrieve the source first when the work requires a current contract, price, procedure, or customer record; evaluate fine-tuning experimentally when the problem is a repeatable response format, classification boundary, or behaviour; and make deterministic services—not an LLM—the decision maker for exact calculations, authorization, and irreversible actions.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Separate the problem first: knowledge, behaviour, or rule
&lt;/h2&gt;

&lt;p&gt;When a user asks, “What is the return window for this customer?” the problem is access to current knowledge. The policy may change tomorrow, and the evidence for a correct answer is the relevant policy revision. That is a RAG problem. “Classify an incoming email into one of six operations categories and return valid JSON only” is a consistency and output-contract problem; it may be a fine-tuning candidate. “A second approval is required above a 5,000 TRY refund” is a deterministic business rule and belongs in SQL, a workflow, or a rules service.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Ftankdev.tech%2Fuploads%2Fblog%2Frag-decision-en.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Ftankdev.tech%2Fuploads%2Fblog%2Frag-decision-en.svg" alt="Decision diagram showing that RAG, fine-tuning, and deterministic rules solve different problems" width="960" height="580"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Need&lt;/th&gt;
&lt;th&gt;First choice&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Current policy, product, or contract&lt;/td&gt;
&lt;td&gt;RAG&lt;/td&gt;
&lt;td&gt;The source revision changes&lt;/td&gt;
&lt;td&gt;Citation + document ID&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repeated classification or format&lt;/td&gt;
&lt;td&gt;Prompt → fine-tuning experiment&lt;/td&gt;
&lt;td&gt;Behaviour repeats&lt;/td&gt;
&lt;td&gt;Held-out test set&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exact calculation, permission, limit&lt;/td&gt;
&lt;td&gt;Deterministic code&lt;/td&gt;
&lt;td&gt;An error is costly&lt;/td&gt;
&lt;td&gt;Tests + audit log&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ambiguous high-risk case&lt;/td&gt;
&lt;td&gt;Human-in-the-loop&lt;/td&gt;
&lt;td&gt;Outcome is irreversible&lt;/td&gt;
&lt;td&gt;Review record&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  2. What RAG is—and is not
&lt;/h2&gt;

&lt;p&gt;Retrieval-augmented generation selects evidence from authorized sources for a user's question and supplies that evidence to the model context. A production path acquires sources, prepares text and metadata, retrieves candidates, reranks evidence, produces a cited answer, and records observability events. A vector database can be one component of this path; it is not a synonym for RAG.&lt;/p&gt;

&lt;p&gt;Good RAG does not end at splitting a PDF into smaller pieces. It preserves revision, effective date, product scope, language, access control, heading hierarchy, and source link. Retrieval may combine lexical and semantic search, then use a reranker to narrow evidence. The model should answer from evidence or explicitly report that evidence is insufficient.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Ftankdev.tech%2Fuploads%2Fblog%2Frag-pipeline-en.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Ftankdev.tech%2Fuploads%2Fblog%2Frag-pipeline-en.svg" alt="Production RAG pipeline showing source preparation, retrieval, response, observability, and evaluation" width="960" height="580"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here, &lt;code&gt;actor_id&lt;/code&gt; is not merely a log field: retrieval must filter out customer or contract fragments that the actor cannot access. Adding “do not reveal private information” to a prompt afterward is not access control. If retrieval returns no sufficient evidence, the system should enter an &lt;code&gt;insufficient_evidence&lt;/code&gt; state rather than generate a plausible answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. What does fine-tuning change?
&lt;/h2&gt;

&lt;p&gt;Fine-tuning moves a base model toward selected input-output examples for a specific task. Classification labels, structured output format, domain language, concise response style, or preferred tool selection can be suitable targets. A provider treats it as a separate operation with a training dataset and job; for example, the OpenAI API specifies JSONL training data for a fine-tuning job. &lt;a href="https://platform.openai.com/docs/api-reference/fine-tuning/resume?lang=python" rel="noopener noreferrer"&gt;OpenAI fine-tuning reference&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Fine-tuning does not turn training facts into a reliable document store. When a policy or price changes, it is hard to identify every affected answer. A memorized-looking example is neither citable, current, nor access-controlled evidence. Knowledge-base questions can therefore require RAG alongside a fine-tuned model.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Ftankdev.tech%2Fuploads%2Fblog%2Ffinetuning-pipeline-en.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Ftankdev.tech%2Fuploads%2Fblog%2Ffinetuning-pipeline-en.svg" alt="Fine-tuning pipeline from behaviour target and examples to evaluation and controlled release" width="960" height="580"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A training example represents the ideal work outcome and permitted output contract.&lt;/li&gt;
&lt;li&gt;A validation set supports prompt, hyperparameter, or decision-threshold choices.&lt;/li&gt;
&lt;li&gt;A test set stays untouched until final comparison; the same customer, template, or document family must not leak across splits.&lt;/li&gt;
&lt;li&gt;Model, dataset, prompt, and evaluation versions belong in the same release record.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. Concrete example: a dealer support assistant
&lt;/h2&gt;

&lt;p&gt;A B2B support assistant can perform three different jobs. When a user asks for the current return condition, the system retrieves the authorized contract revision with RAG and displays the clause link. When a user submits free-form text, a model classifies it as &lt;code&gt;returns&lt;/code&gt;, &lt;code&gt;pricing&lt;/code&gt;, &lt;code&gt;delivery&lt;/code&gt;, &lt;code&gt;account&lt;/code&gt;, &lt;code&gt;technical&lt;/code&gt;, or &lt;code&gt;other&lt;/code&gt;. Before an action creates a return, a deterministic service validates the customer, order, time window, and authorization.&lt;/p&gt;

&lt;p&gt;For classification, compare a baseline prompt against 600 labelled examples on a held-out test set. If failures are mostly caused by changing policy knowledge, fine-tuning is the wrong investment; repair retrieval, filters, or source quality. If evidence is present but category format or domain language stays inconsistent, a fine-tuning experiment has a meaningful hypothesis. This distinction prevents the costly reflex to train first.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. How do you measure RAG quality?
&lt;/h2&gt;

&lt;p&gt;A convincing answer is not proof that the system is good. RAG evaluation needs at least two layers: did retrieval bring back the correct evidence, and did generation provide a correct and sufficient answer using only that evidence? A wrong document can be summarized fluently. A correct document can be used to support the wrong clause. These failures need different owners and different fixes.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measure&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Example signal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Recall@k&lt;/td&gt;
&lt;td&gt;Is the correct source in the first k results?&lt;/td&gt;
&lt;td&gt;Critical clause in top 5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Citation precision&lt;/td&gt;
&lt;td&gt;Does the displayed source support the claim?&lt;/td&gt;
&lt;td&gt;Human review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grounded correctness&lt;/td&gt;
&lt;td&gt;Does the answer follow evidence and the business rule?&lt;/td&gt;
&lt;td&gt;Expected-answer rubric&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Abstention&lt;/td&gt;
&lt;td&gt;Does it stop when evidence is absent?&lt;/td&gt;
&lt;td&gt;Review rather than answer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency / cost&lt;/td&gt;
&lt;td&gt;Are time and spend acceptable?&lt;/td&gt;
&lt;td&gt;p95 + cost per query&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;An evaluation set should include easy FAQs alongside name collisions, stale revisions, multi-document answers, permission boundaries, OCR errors, empty retrieval, and conflicting sources. Google Cloud's RAG guidance also emphasizes that retrieval quality is central and irrelevant retrieval can lead to off-topic or incorrect generation. &lt;a href="https://cloud.google.com/use-cases/retrieval-augmented-generation" rel="noopener noreferrer"&gt;Google Cloud RAG guide&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Make fine-tuning a measured experiment
&lt;/h2&gt;

&lt;p&gt;Fine-tuning is not a deployment decision; it is an experiment. State a measurable objective first: “the JSON schema is valid” or “priority classification reaches at least X accuracy under expert adjudication.” Compare a base model with a strong system prompt, a base model with RAG, and a fine-tuned version using blind evaluation on the same test set. More complexity does not make an approach more correct.&lt;/p&gt;

&lt;p&gt;Classify failure examples: wrong knowledge, wrong retrieval, schema failure, wrong tool selection, missing context, safety violation, or evaluation ambiguity. Adding every failure to training can harm dataset quality. Examples should show ideal behaviour; personal data, access keys, stale conflicting rules, and unverified model output do not belong in training data.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Cost and latency are different accounting models
&lt;/h2&gt;

&lt;p&gt;RAG cost includes source preparation and indexing plus query-time embedding, search, reranking, context tokens, and generation. When a document changes, only affected fragments can be processed again. Fine-tuning has dataset preparation, training-job, evaluation, and custom-model inference costs. The material cost, however, includes correction effort, operational delay, and the impact of a wrong answer—not the model invoice alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Security, access, and data lifecycle
&lt;/h2&gt;

&lt;p&gt;Before a RAG system adds a document to context, it applies tenant, role, record-level, contract-scope, and effective-date filters. Source fragments, answer, and actor identity may be recorded for observability with data minimization. When a document is deleted or permission is revoked, remove it from the index and caches as well.&lt;/p&gt;

&lt;p&gt;Fine-tuning data adds a second governance question: examine the provider's data retention, usage, and deletion terms. Governance is not solved by calling a file “fine-tuning data.” Define which data class may leave the boundary, how long it is retained, and how a model revision is retired.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. A hybrid architecture is often the right result
&lt;/h2&gt;

&lt;p&gt;Mature systems can use both methods: a fine-tuned or carefully prompted model structures the query and stabilizes response form; RAG brings current, authorized evidence; a rules service validates permissions, calculations, and side effects; human review closes uncertain or high-risk cases. This assigns each component the responsibility it can actually carry instead of trying to teach everything to a model.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. A practical decision sequence
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Write the success contract: what are the correct source, required fields, side effect, and deadline?&lt;/li&gt;
&lt;li&gt;Establish a simple baseline: strong prompt, schema validation, rule checks, and a measured example set.&lt;/li&gt;
&lt;li&gt;When knowledge changes, start with versioned sources and RAG; evaluate retrieval separately.&lt;/li&gt;
&lt;li&gt;When a behaviour problem repeats, test the fine-tuning hypothesis on a fixed test set.&lt;/li&gt;
&lt;li&gt;Do not execute a high-risk decision without deterministic validation or human approval.&lt;/li&gt;
&lt;li&gt;Monitor versions, cost, false acceptance, source evidence, and rollback in production.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is RAG or fine-tuning more accurate?
&lt;/h3&gt;

&lt;p&gt;Use RAG first for current, citable knowledge; fine-tuning can be evaluated for repeated behaviour, format, or classification. Deterministic code remains the right layer for exact business rules. The same held-out work scenarios should decide.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is fine-tuning suitable for teaching company documents to a model?
&lt;/h3&gt;

&lt;p&gt;It is usually not the first choice for frequently changing policies, prices, contracts, or customer records. Those require versioning, citation, deletion, and access control, which RAG supports more directly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does RAG eliminate hallucinations?
&lt;/h3&gt;

&lt;p&gt;No. Retrieval can be wrong or incomplete, and a model can still go beyond evidence. Use source filters, reranking, citation checks, abstention when evidence is absent, and evaluation together.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is a vector database enough for RAG?
&lt;/h3&gt;

&lt;p&gt;No. Source quality, chunking, metadata, access filters, hybrid search, reranking, version management, and evaluation matter at least as much as the database choice.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can RAG and fine-tuning be combined?
&lt;/h3&gt;

&lt;p&gt;Yes. Fine-tuning or a strong prompt can improve behaviour and response form while RAG supplies current evidence. Permissions, calculations, and persistent side effects still need deterministic validation.&lt;/p&gt;

&lt;h3&gt;
  
  
  What should the first production release measure?
&lt;/h3&gt;

&lt;p&gt;Measure business-outcome correctness, correct-source retrieval, citation support, schema validity, abstention without evidence, p95 latency, cost per query, and the rate of human review.&lt;/p&gt;

&lt;p&gt;Author&lt;/p&gt;

&lt;h2&gt;
  
  
  Murat Onur Kaderoğlu
&lt;/h2&gt;

&lt;p&gt;Algorithm and Architecture Developer&lt;/p&gt;




&lt;p&gt;Published by &lt;a href="https://tankdev.tech/en/" rel="noopener noreferrer"&gt;TankDev&lt;/a&gt; · Original article: &lt;a href="https://tankdev.tech/en/blog/rag-vs-fine-tuning" rel="noopener noreferrer"&gt;RAG vs. Fine-Tuning&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>python</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Ending LLM bloat: How to integrate Jev and System One decision models into production architecture</title>
      <dc:creator>Murat Onur Kaderoglu</dc:creator>
      <pubDate>Sun, 27 Sep 2026 19:59:37 +0000</pubDate>
      <link>https://dev.to/tankdev-tech/ending-llm-bloat-how-to-integrate-jev-and-system-one-decision-models-into-production-architecture-2hmo</link>
      <guid>https://dev.to/tankdev-tech/ending-llm-bloat-how-to-integrate-jev-and-system-one-decision-models-into-production-architecture-2hmo</guid>
      <description>&lt;h1&gt;
  
  
  Ending LLM bloat: How to integrate Jev and System One decision models into production architecture
&lt;/h1&gt;

&lt;p&gt;Not every workflow needs generated text. A practical guide to turning Jev Choice, Score, and Noul decisions into a production layer with thresholds, auditability, and human review.&lt;/p&gt;

&lt;p&gt;A meaningful share of AI automation cost comes from asking a large language model to make decisions that do not require generated text. Routing a ticket to billing, sales, or support; selecting a tool; or screening a request for risk does not inherently require a paragraph. Software needs a value it can execute, not prose for a person to read.&lt;/p&gt;

&lt;p&gt;Jev is TypeSafe's System One decision model: an application sends state and typed questions, and receives a choice, score, or probability instead of generated text. &lt;a href="https://docs.typesafe.ai/introduction" rel="noopener noreferrer"&gt;The official documentation&lt;/a&gt; defines confidence for Choice and Score and a 0–1 probability for Noul. This does not make LLMs obsolete. It gives decisions and expression separate places in an architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem: structured output is not automatically a structured decision
&lt;/h2&gt;

&lt;p&gt;JSON Schema, function calling, and Pydantic can make an LLM response parseable. The model still interprets instructions, generates tokens autoregressively, and makes the application wait for that generation. A valid &lt;code&gt;route: billing&lt;/code&gt; field does not establish that the route is correct, that uncertainty is handled, or that the call is economical for the workload.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; When a model result causes code to run, the design question is not merely ‘does it return JSON?’ It is ‘which incorrect decisions can this operation tolerate, and what happens when the system is uncertain?’&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  System 1 and System 2: a useful architecture metaphor
&lt;/h2&gt;

&lt;p&gt;Kahneman's System 1 / System 2 distinction is useful here as a technical metaphor. A System 1 layer makes a fast, narrow judgment: which queue, whether risk exists, which tool to select. A System 2 layer reasons across context, explains alternatives, and writes text. A conventional LLM often serves the latter role; decision models such as Jev fit narrow, measurable questions in the former. This is a division of work, not a hierarchy of intelligence.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User request
     │
     ▼
[Validation + authorization] ── invalid ──► reject / explain
     │
     ▼
[Jev: route, risk, human_needed]  ← one state, parallel questions
     │
     ├─ confidence ≥ threshold ─► deterministic workflow / tool
     ├─ confidence &amp;lt; threshold ─► human-review queue
     └─ prose required ─────────► LLM (draft, summary, explanation)
                                      │
                                      ▼
                              audit log + metrics
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Three decision primitives: Choice, Score, and Noul
&lt;/h2&gt;

&lt;p&gt;Jev constrains the shape of a decision before the request is sent. Instead of leaving the answer open-ended, the application states which kind of value it will consume. Independent judgments can be collected as one decision package because multiple questions on the same state are evaluated in parallel.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Primitive&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Production example&lt;/th&gt;
&lt;th&gt;Value consumed by code&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Choice&lt;/td&gt;
&lt;td&gt;Which option?&lt;/td&gt;
&lt;td&gt;Which team owns this ticket?&lt;/td&gt;
&lt;td&gt;`billing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Score&lt;/td&gt;
&lt;td&gt;What level on a defined rubric?&lt;/td&gt;
&lt;td&gt;How urgent is the incident?&lt;/td&gt;
&lt;td&gt;probability-weighted score + confidence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Noul&lt;/td&gt;
&lt;td&gt;Is this statement true?&lt;/td&gt;
&lt;td&gt;Does this require human review?&lt;/td&gt;
&lt;td&gt;{% raw %}&lt;code&gt;P(true) ∈ [0, 1]&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;The three decision types documented by TypeSafe.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"state"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"subject"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"My card was charged twice"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"channel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"web"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"account_tier"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"enterprise"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"questions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"route"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"choice"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Route to the accountable team."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"criteria"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"billing"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"payment, invoice, or collection"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"support"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"product issue"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"sales"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"purchase or quote"&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"urgency"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"score"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Assess operational urgency."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"criteria"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"low"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"normal"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"high"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"critical"&lt;/span&gt;&lt;span class="p"&gt;]},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"human_needed"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"noul"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Is human review needed because of financial harm or account security?"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here, &lt;code&gt;route&lt;/code&gt; is not a label for a dashboard: it selects a queue, SLA, and access boundary. &lt;code&gt;urgency&lt;/code&gt; can drive prioritization. &lt;code&gt;human_needed&lt;/code&gt; can stop an automated reply. Keep questions atomic. ‘Route this safely, urgently, and correctly’ combines unrelated judgments and is difficult to evaluate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost and latency: compare the right unit
&lt;/h2&gt;

&lt;p&gt;A universal benchmark would be misleading because context length, caching, output length, region, and concurrency change the result. TypeSafe advertises Jev at &lt;strong&gt;$42 per billion input tokens&lt;/strong&gt;, or about $0.042 per million input tokens. An LLM decision cost includes more than input pricing: system instructions, generated structured output, possible repair calls, and waiting time. Measure p50/p95 latency, tokens per call, fallback rate, and the cost of incorrect decisions with your own traffic.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measure&lt;/th&gt;
&lt;th&gt;Jev / System One decision layer&lt;/th&gt;
&lt;th&gt;Structured-output LLM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;Typed choice, score, or probability&lt;/td&gt;
&lt;td&gt;Generated JSON matching a schema&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Execution&lt;/td&gt;
&lt;td&gt;Narrow parallel questions over one state&lt;/td&gt;
&lt;td&gt;Autoregressive generation with output tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost model&lt;/td&gt;
&lt;td&gt;Input tokens plus platform cost&lt;/td&gt;
&lt;td&gt;Input + output + possible retries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;Measure on the target short decision&lt;/td&gt;
&lt;td&gt;Varies with model, load, and output length&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best fit&lt;/td&gt;
&lt;td&gt;Routing, gates, classification, risk signals&lt;/td&gt;
&lt;td&gt;Explanation, summary, advice, and generation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Not a product benchmark; a framework for a production measurement.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The TankDev decision benchmark: a measurement contract
&lt;/h2&gt;

&lt;p&gt;A provider demo cannot prove that an architecture is fast or inexpensive for a production domain. Decision quality depends on the data domain, the definition of options, and the cost of a wrong result. TankDev's recommended benchmark runs one closed evaluation set through three paths: deterministic rules, Jev/System One, and a structured-output LLM. The aim is not the largest headline percentage; it is to show &lt;strong&gt;which decisions can safely be automated in which layer&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```benchmark flow&lt;br&gt;
Labeled evaluation set (for example, 500 anonymized requests)&lt;br&gt;
                  │&lt;br&gt;
                  ├──► Deterministic rule ─────────► result + duration + cost&lt;br&gt;
                  ├──► Jev / System One ───────────► result + distribution + confidence&lt;br&gt;
                  └──► Structured-output LLM ──────► result + output tokens&lt;/p&gt;

&lt;p&gt;For every result: correct label • p50/p95 • call cost • fallback • human override&lt;br&gt;
Segment separately: channel • language • account type • novel/rare category&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


| Measure | How to calculate it | Why it matters |
| --- | --- | --- |
| Selective accuracy | Correct rate among decisions actually automated | Shows automation quality above the confidence threshold |
| Coverage | Share of all records that received an automatic action | Prevents looking good only on easy examples |
| Risk-weighted error | Weight each wrong result by business impact | Separates a misroute from a money or access error |
| p50 / p95 latency | Time from client request to decision | Exposes queues and tail values that an average hides |
| Fallback and override | Rate of human review or later correction | Shows threshold and question-design maturity |

*A publishable benchmark reports coverage and error cost, not accuracy alone.*

This article makes no numerical benchmark claim; it provides a **results template**. Before publishing actual results, document the data source, sample size, labeling method, model/question version, date range, and exclusion criteria. Without that, a comparison is a marketing chart rather than engineering evidence.

| Method | Selective accuracy | Coverage | p95 | Call cost | Note |
| --- | --- | --- | --- | --- | --- |
| Rule engine | to measure | to measure | to measure | to measure | Reference for explicitly governed examples |
| Jev / System One | to measure | to measure | to measure | to measure | Report with confidence threshold and calibration |
| Structured-output LLM | to measure | to measure | to measure | to measure | Separate schema validity from decision correctness |

*TankDev benchmark result card — fields that must remain empty until measurement.*

## Hybrid production architecture: decision model, policy, and LLM

Do not attach a decision model directly to an external side effect. Put a policy layer between the decision and the action. The policy engine applies thresholds, user roles, transaction limits, operating hours, retry count, and risk classes. The model may say that a vendor payment is high risk; explicit application policy decides whether money can move.



```typescript
const decision = await jev.decide(ticket);
const autoRoute = decision.route.confidence &amp;gt;= 0.85;
const needsReview = decision.human_needed.noul &amp;gt;= 0.20;

if (!autoRoute || needsReview) {
  await reviewQueue.enqueue({ ticketId, decision, policyVersion: "2026-09-27" });
  return { status: "pending_review" };
}

await workflow.dispatch({ queue: decision.route.choice, ticketId });
// Call an LLM only when a human-readable draft or explanation is required.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Confidence is not authority: calibration and thresholds
&lt;/h2&gt;

&lt;p&gt;A confidence of 0.85 only approaches ‘85% correct’ when the model is calibrated on your domain. Begin in shadow mode: record recommendations, compare them with an existing rule or human outcome, and inspect error by segment. Language, customer tier, channel, product family, and newly introduced categories need separate measurement.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Log a safe state reference, model version, question/policy version, result, distribution, threshold, and final action for every decision.&lt;/li&gt;
&lt;li&gt;Choose thresholds from false-positive and false-negative cost, not headline accuracy.&lt;/li&gt;
&lt;li&gt;Do not turn low confidence into an automatic rejection; design review, an alternate path, or a request for richer context.&lt;/li&gt;
&lt;li&gt;Compare old and new versions under controlled traffic whenever the model or question text changes.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Jev does not do
&lt;/h2&gt;

&lt;p&gt;Jev does not write text. It is therefore not the right tool for explanations, long summaries, code, or multi-step research. A fixed output space prevents an invented schema value, but it does not prevent a high-confidence selection of the wrong defined option. ‘Zero hallucinations’ must not be interpreted as correct business decisions. If a deterministic rule already exists—such as &lt;code&gt;amount &amp;gt; 100000&lt;/code&gt;—use ordinary code rather than a model call.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Good automation does not send every classification to AI. Ask whether the rule is explicit. Use code for an explicit rule, a decision model for a narrow uncertain judgment, and an LLM when language or explanation must be produced.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Production release checklist
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Write the business goal, allowed actions, and reversibility of each decision.&lt;/li&gt;
&lt;li&gt;Describe every Choice option and Score level with observable criteria.&lt;/li&gt;
&lt;li&gt;Prepare an evaluation set from labeled or human-decided examples; do not reuse training examples.&lt;/li&gt;
&lt;li&gt;Specify a fallback for low confidence, timeout, and service failure.&lt;/li&gt;
&lt;li&gt;Monitor queue delay, model latency, confidence distribution, override rate, error rate, and call cost.&lt;/li&gt;
&lt;li&gt;Apply dual approval, limits, idempotency, and audit logs to high-impact actions.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Jev an LLM?
&lt;/h3&gt;

&lt;p&gt;No. TypeSafe defines Jev as a System One decision model that does not generate text. It accepts state and typed questions, then returns Choice, Score, or Noul results. It does not replace an LLM for conversation, long explanation, or content generation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why use Jev when Structured Outputs exist?
&lt;/h3&gt;

&lt;p&gt;Structured Outputs constrain shape; they do not by themselves solve an LLM's generation cost, latency, or decision uncertainty. Jev provides typed probability and confidence signals for narrow decisions that code consumes. The right approach depends on measured performance and the cost of an error.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does Jev avoid hallucinations?
&lt;/h3&gt;

&lt;p&gt;It cannot create a class outside the defined schema. It can still select the wrong existing option. Confidence, evaluation data, thresholds, fallback behavior, and human approval for high-impact actions remain necessary.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why does Jev not provide a written rationale?
&lt;/h3&gt;

&lt;p&gt;It is designed to produce a fast machine-consumable decision, not an explanation. When a user needs a rationale, build a separate explanation flow using the decision record and relevant business data with an LLM or human review.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where does Jev fit in an agent architecture?
&lt;/h3&gt;

&lt;p&gt;Place it at early decision points: tool selection, risk gating, routing, spam or suitability filters, and triggers for human review. A Jev result must not be authority on its own; combine it with policy, permissions, limits, and an audit log.&lt;/p&gt;

&lt;h3&gt;
  
  
  Should the confidence threshold always be 0.85?
&lt;/h3&gt;

&lt;p&gt;No. There is no universal threshold. Set it using calibration on your labeled examples and the cost of an incorrect outcome. Low-impact routing can tolerate a different threshold from a payment or permission change that requires human approval.&lt;/p&gt;




&lt;p&gt;Published by &lt;a href="https://tankdev.tech/en/" rel="noopener noreferrer"&gt;TankDev&lt;/a&gt; · Original article: &lt;a href="https://tankdev.tech/en/blog/jev-system-one-decision-models-production-architecture" rel="noopener noreferrer"&gt;Jev and System One decision models&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>architecture</category>
      <category>automation</category>
    </item>
  </channel>
</rss>
