<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: NEXT4I DEV</title>
    <description>The latest articles on DEV Community by NEXT4I DEV (@dev_next4i).</description>
    <link>https://dev.to/dev_next4i</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4057697%2Fb9191896-5c71-46c0-985e-6c58c3c85ffb.png</url>
      <title>DEV Community: NEXT4I DEV</title>
      <link>https://dev.to/dev_next4i</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dev_next4i"/>
    <language>en</language>
    <item>
      <title>Chunking in RAG: How to Split Documents Without Losing the Conditions</title>
      <dc:creator>NEXT4I DEV</dc:creator>
      <pubDate>Wed, 07 Oct 2026 18:27:50 +0000</pubDate>
      <link>https://dev.to/dev_next4i/chunking-in-rag-how-to-split-documents-without-losing-the-conditions-9kb</link>
      <guid>https://dev.to/dev_next4i/chunking-in-rag-how-to-split-documents-without-losing-the-conditions-9kb</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fql36spyguo2jxqcn2bam.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fql36spyguo2jxqcn2bam.webp" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
Your AI assistant has the employee benefits handbook. You ask: “Can someone who is still on probation claim the eyewear benefit?”&lt;/p&gt;

&lt;p&gt;It confidently answers: “Yes, up to 3,000 baht per year.” But the handbook says the benefit is only for permanent employees who have completed probation. The limit reached the model. The eligibility condition did not.&lt;/p&gt;

&lt;p&gt;That failure can start before generation, at the document boundary. &lt;strong&gt;Chunking in RAG is not about making text as small as possible. It is about creating searchable pieces without separating facts from the conditions that make them true.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The policy and amounts here are fictional examples, not NEXT4I or customer policies.&lt;/p&gt;
&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A chunk is a piece of source content used for retrieval. It is not inherently a vector.&lt;/li&gt;
&lt;li&gt;Large chunks can mix unrelated topics. Small chunks can lose subjects, conditions, and exceptions.&lt;/li&gt;
&lt;li&gt;Start with document structure, cap oversized sections, and test with real questions.&lt;/li&gt;
&lt;li&gt;Overlap and parent-child retrieval can help with missing context. They cannot repair incorrectly extracted text.&lt;/li&gt;
&lt;li&gt;Check retrieved evidence, context completeness, and answer correctness separately.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  What is chunking, and why not use the entire file?
&lt;/h2&gt;

&lt;p&gt;Chunking divides a document into pieces that can be indexed and retrieved. A simplified pipeline looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Document -&amp;gt; Extract and organize text -&amp;gt; Create chunks
         -&amp;gt; Store text and source references -&amp;gt; Build an index

Question -&amp;gt; Retrieve relevant chunks -&amp;gt; Assemble context -&amp;gt; LLM answer with sources
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Retrieval may use keyword search, vector search, or both. Not every chunk needs an embedding. When vector search is used, an embedding model creates a numerical representation of the content.&lt;/p&gt;

&lt;p&gt;The searchable unit also does not have to be the final reading unit. A system can retrieve a short passage, then supply its surrounding section to the model.&lt;/p&gt;

&lt;p&gt;A single chunk containing a 100-page handbook mixes many topics. That can make specific passages harder to retrieve and exceed model input limits. Sending unrelated material also costs tokens and can distract from the evidence that matters.&lt;/p&gt;

&lt;p&gt;Whole-document context can work for short documents. Tiny chunks are not automatically better: an amount without its eligibility rule is incomplete evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  One boundary can change the answer
&lt;/h2&gt;

&lt;p&gt;Consider this fictional passage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Section 8.1: Prescription eyewear benefit
The company reimburses up to 3,000 baht per year.
Only permanent employees who have completed probation are eligible.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A length-based split might produce:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Chunk A: The company reimburses up to 3,000 baht per year.
Chunk B: Only permanent employees who have completed probation are eligible.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without the heading, both pieces lose their subject. B never names the benefit, and vector search does not guarantee recovery of that missing relationship.&lt;/p&gt;

&lt;p&gt;If only A reaches the LLM, it has evidence about the amount, not enough evidence about eligibility. A prompt telling it not to invent information may help it acknowledge uncertainty. It cannot restore a sentence that was never supplied.&lt;/p&gt;

&lt;p&gt;Keeping the heading, amount, and condition together helps make evidence usable. It does not guarantee a correct answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main approaches and their trade-offs
&lt;/h2&gt;

&lt;p&gt;These approaches can be combined. Splitting by heading and then dividing oversized sections is often more useful than committing to one method everywhere.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Fixed-size splitting
&lt;/h3&gt;

&lt;p&gt;Choose a length, such as 500 tokens, and divide the text into consecutive pieces. It is straightforward, but the boundary may cut through a sentence, table, or exception.&lt;/p&gt;

&lt;p&gt;Use the actual model tokenizer when token limits matter. Character counts are not interchangeable with token counts, and a Thai word is not necessarily one token.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Fixed-size splitting with overlap
&lt;/h3&gt;

&lt;p&gt;Overlap repeats some text from one chunk in the next, giving material near the boundary a chance to stay together.&lt;/p&gt;

&lt;p&gt;It cannot fix distant conditions or scrambled reading order. Thai OCR with unreliable spacing still needs boundary checks.&lt;/p&gt;

&lt;p&gt;More overlap means more indexed content and duplication. Near-identical results can occupy retrieval slots that should contain other evidence. Deduplicate or merge overlapping context before sending it to the LLM where appropriate.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Recursive splitting
&lt;/h3&gt;

&lt;p&gt;A recursive splitter tries separators in a chosen order, such as paragraphs, sentences, then finer boundaries, until pieces fit the size limit.&lt;/p&gt;

&lt;p&gt;It avoids unnecessary cuts through sentences, but does not understand meaning. A paragraph can cover several topics; damaged OCR can remove boundaries.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Structure-aware splitting
&lt;/h3&gt;

&lt;p&gt;Use headings, lists, tables, and code blocks to guide boundaries. A subsection can carry its parent heading so a short passage still has a clear subject.&lt;/p&gt;

&lt;p&gt;Tables need special care. An isolated “3,000” is not enough. Preserve the row name, column name, and unit. In our fictional example, “Prescription eyewear: annual reimbursement limit of 3,000 baht” is clearer. Any eligibility note must accompany it too.&lt;/p&gt;

&lt;p&gt;Code and ordered requirements have similar dependencies. If an oversized block must be split, preserve its relationships and necessary conditions.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Semantic chunking
&lt;/h3&gt;

&lt;p&gt;Semantic methods estimate where nearby sentences or paragraphs change topic, often using embeddings or a model.&lt;/p&gt;

&lt;p&gt;They can help when structure is unclear, but add processing cost and tuning. Results depend on the model and source quality. If headings already separate topics well, I would test that simpler structure first.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Parent-child retrieval and contextual chunking
&lt;/h3&gt;

&lt;p&gt;These solve different problems alongside the splitting methods above.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Parent-child retrieval&lt;/strong&gt; indexes small child passages for search, then expands a match to a larger parent section for reading. Maintain the links, control context size, and check access permissions for the expanded parent too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contextual chunking&lt;/strong&gt; adds useful context before indexing, such as a document title, section name, or short source-grounded explanation. If an LLM generates that explanation, check that it has not invented conditions. Keep original text available for citations instead of treating generated context as source evidence.&lt;/p&gt;

&lt;p&gt;Propositional or agentic approaches may use an LLM to extract standalone statements or choose content groupings. Implementations vary. The shared trade-offs are extra cost and the risk that rewriting drops a qualification. Compare them with a non-LLM baseline and retain the original source.&lt;/p&gt;

&lt;h2&gt;
  
  
  Match the approach to the document
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Document&lt;/th&gt;
&lt;th&gt;A reasonable first approach&lt;/th&gt;
&lt;th&gt;Check carefully&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Markdown or structured manuals&lt;/td&gt;
&lt;td&gt;Split by heading, subdivide oversized sections&lt;/td&gt;
&lt;td&gt;Parent headings, tables, footnotes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plain text with inconsistent structure&lt;/td&gt;
&lt;td&gt;Recursive splitting, limited overlap&lt;/td&gt;
&lt;td&gt;Conditions at boundaries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Contracts or rules requiring several paragraphs&lt;/td&gt;
&lt;td&gt;Retrieve small passages, expand or combine evidence&lt;/td&gt;
&lt;td&gt;Clause numbers, versions, permissions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scanned or multi-column PDFs&lt;/td&gt;
&lt;td&gt;Repair extraction and reading order first&lt;/td&gt;
&lt;td&gt;OCR errors, scrambled columns, broken tables&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are starting points, not substitutes for inspecting actual documents.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Thai tourism PDF that changed my priorities
&lt;/h3&gt;

&lt;p&gt;While experimenting with a RAG pipeline, I used a beautifully designed tourism PDF about Thailand. The extracted text was much less attractive: Thai vowels and tone marks were misplaced, while colored backgrounds, watermarks, and illustrations interfered with reading.&lt;/p&gt;

&lt;p&gt;For some pages, I tried black-and-white conversion to make the letters clearer. That could also lose visual context. I compared outputs with help from multiple models and had a person review spelling again.&lt;/p&gt;

&lt;p&gt;I described this in my &lt;a href="https://www.next4i.com/dev-notes/en/markdown-the-ultimate-ai-native-dev-en" rel="noopener noreferrer"&gt;earlier PDF and Markdown write-up, in Thai&lt;/a&gt;. It is one experience, not evidence that every PDF needs the same workflow.&lt;/p&gt;

&lt;p&gt;The lesson was simple: changing chunk size cannot bring back words that extraction has already lost.&lt;/p&gt;

&lt;p&gt;Markdown makes structure easier to identify, not evaluation unnecessary. Sections can still be oversized, and table rows need headers.&lt;/p&gt;

&lt;h2&gt;
  
  
  How large should chunks and overlap be?
&lt;/h2&gt;

&lt;p&gt;There is no universal setting. &lt;strong&gt;300-500 tokens with 10-20% overlap&lt;/strong&gt; can be an initial experiment for paragraph-based text. Those numbers are not a benchmark result from my system or a standard to apply everywhere.&lt;/p&gt;

&lt;p&gt;Before tuning, check four things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Meaning:&lt;/strong&gt; Does each piece identify the subject and preserve necessary conditions?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Limits:&lt;/strong&gt; Use the actual tokenizer and budget for instructions, the question, all retrieved context, and the answer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Duplication:&lt;/strong&gt; Does overlap add useful evidence or just repeat it?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Question type:&lt;/strong&gt; A lookup for an amount and a question about eligibility may require different evidence.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Start with structure and a size cap. Add overlap or parent expansion when tests reveal missing context, not simply because more context feels safer.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to tell whether chunking works
&lt;/h2&gt;

&lt;p&gt;Build questions with known supporting documents, sections, or pages. Include boundary-sensitive questions, amount lookups, and questions the documents cannot answer.&lt;/p&gt;

&lt;p&gt;Compare two or three approaches on the same sources and questions with comparable context budgets. Then inspect three layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval:&lt;/strong&gt; Are the labeled supporting passages found? Recall@k measures retrieval of predefined relevant evidence in the top k results. Separately check whether the evidence available is sufficient for the question.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context:&lt;/strong&gt; Are conditions and exceptions present? Do duplicates crowd out other evidence? Are table headers and units intact?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answer:&lt;/strong&gt; Does it follow the evidence? Does each citation support the associated claim? Can the system acknowledge insufficient information?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Store metadata such as &lt;code&gt;document_id&lt;/code&gt;, &lt;code&gt;section&lt;/code&gt;, &lt;code&gt;page/offset&lt;/code&gt;, and &lt;code&gt;version&lt;/code&gt; so passages can be traced and stale index entries managed.&lt;/p&gt;

&lt;p&gt;For organizational data, enforce permissions before content reaches the model, including any expanded parent context. Hiding the final answer is not a substitute for controlling what the model receives.&lt;/p&gt;

&lt;p&gt;If source text is wrong, fix ingestion. If evidence exists but is not found, inspect chunking and retrieval. If complete evidence reaches the model but the answer is wrong, inspect context assembly and generation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The goal is usable evidence
&lt;/h2&gt;

&lt;p&gt;A RAG system generally supplies selected evidence, not the entire library, for each question. Chunk boundaries help determine what the model gets to see, though they are not the only factor in answer quality.&lt;/p&gt;

&lt;p&gt;My starting point is accurate source text, structure-aware boundaries, sensible size limits, source references, and tests with real questions. Add complexity where those tests show a need.&lt;/p&gt;

&lt;p&gt;The best chunk is not necessarily the smallest one. It is a piece the system can find and use without losing what makes the answer valid.&lt;/p&gt;

&lt;p&gt;If you are building RAG, share which document types cause missing conditions and where in the pipeline you catch them.&lt;/p&gt;




&lt;p&gt;Explore the NEXT4I journey and read the original article at: &lt;a href="https://go.next4i.com/next4i-devto-en" rel="noopener noreferrer"&gt;https://go.next4i.com/next4i-devto-en&lt;/a&gt;&lt;/p&gt;

</description>
      <category>rag</category>
      <category>aie</category>
      <category>ai</category>
      <category>next4i</category>
    </item>
    <item>
      <title>From Ctrl + F to AI-Powered Search</title>
      <dc:creator>NEXT4I DEV</dc:creator>
      <pubDate>Wed, 30 Sep 2026 18:11:58 +0000</pubDate>
      <link>https://dev.to/dev_next4i/from-ctrl-f-to-ai-powered-search-192l</link>
      <guid>https://dev.to/dev_next4i/from-ctrl-f-to-ai-powered-search-192l</guid>
      <description>&lt;h1&gt;
  
  
  How Do Computers Find Content and Answer Questions from Files?
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;(A Deep Dive into Search &amp;amp; Retrieval Fundamentals: From Ctrl+F to AI Semantic Search)&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;In the daily work of modern tech teams and knowledge workers, there is a classic problem that never really goes away:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"We know this information exists somewhere, but we have no idea which file, which page, or which line it is on."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ul&gt;
&lt;li&gt;In the past, we relied on &lt;code&gt;Ctrl + F&lt;/code&gt; to search through documents one file at a time.&lt;/li&gt;
&lt;li&gt;Later, search servers and file search tools introduced indexing, allowing us to find file names and keywords much faster across storage systems.&lt;/li&gt;
&lt;li&gt;Today, in the era of Generative AI, we can simply ask a question in a chat interface:
&lt;em&gt;"Summarize the vendor contract for me: what is the late delivery penalty?"&lt;/em&gt;
And within seconds, the system digs through a 200-page document and extracts the exact numbers and contractual conditions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Behind the scenes, though, how does a computer actually "search" and "understand" the content inside our files?&lt;br&gt;&lt;br&gt;
Why does traditional keyword search frequently miss critical answers? And why do modern search engines seem to understand what we mean, even when our queries share zero words with the source document?&lt;/p&gt;

&lt;p&gt;This article walks through the &lt;strong&gt;engineering foundations of Search &amp;amp; Retrieval&lt;/strong&gt;, from brute-force byte scans to vector embeddings and AI-powered synthesis. It provides the essential mental model you need before diving deeper into Retrieval-Augmented Generation (RAG).&lt;/p&gt;


&lt;h2&gt;
  
  
  1. The Fundamental Gap: Humans Read Meaning, Computers Read Bytes
&lt;/h2&gt;

&lt;p&gt;Before discussing search algorithms, we need to understand the fundamental limitation of computers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Humans:&lt;/strong&gt; When we read the phrase &lt;em&gt;"an employee is unwell"&lt;/em&gt;, we instantly connect it with &lt;em&gt;"sick leave"&lt;/em&gt;, &lt;em&gt;"medical certificate"&lt;/em&gt;, &lt;em&gt;"hospital"&lt;/em&gt;, or &lt;em&gt;"health insurance benefits"&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Computers:&lt;/strong&gt; At a baseline level, the phrase &lt;code&gt;"sick leave"&lt;/code&gt; is nothing more than an array of UTF-8 bytes. The computer does not inherently know that this sequence of bytes relates to physical health, time off work, or payroll deductions.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This gap created two distinct eras of information retrieval:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Lexical Search:&lt;/strong&gt; Cares strictly about character matching: "Are the words spelled the same way?"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semantic Search:&lt;/strong&gt; Cares about conceptual intent: "Are the meaning and context aligned?"&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  2. Era 1: Exact Character Matching (Lexical / Exact Keyword Search)
&lt;/h2&gt;

&lt;p&gt;This is where search technology started. In production environments, it primarily takes two forms:&lt;/p&gt;
&lt;h3&gt;
  
  
  2.1 Linear Scan (&lt;code&gt;Ctrl + F&lt;/code&gt; or &lt;code&gt;grep&lt;/code&gt;)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;How it works:&lt;/strong&gt;
When a user searches for &lt;code&gt;loan interest&lt;/code&gt;, the system opens the file from the very first byte and iterates through every line sequentially until it reaches the end.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Complexity:&lt;/strong&gt; $O(N)$, where $N$ is the total number of characters or bytes in the dataset.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Advantages:&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Real-time accuracy: if you edit a file, the next search reflects the update immediately without an indexing delay.&lt;/li&gt;
&lt;li&gt;Zero storage overhead: no dedicated search database or pre-computed structures are needed.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disadvantages:&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Does not scale:&lt;/strong&gt; If you have 10,000 files with hundreds of millions of words, opening and scanning every file sequentially will quickly exhaust CPU cycles and disk I/O.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;


&lt;h3&gt;
  
  
  2.2 Inverted Index (The Backbone of Search Engines)
&lt;/h3&gt;

&lt;p&gt;To eliminate the bottleneck of linear scans, computer scientists developed the &lt;strong&gt;Inverted Index&lt;/strong&gt;, which powers engines like Elasticsearch, Apache Lucene, and OpenSearch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;How it works:&lt;/strong&gt;
Instead of waiting for a query before opening files, the system analyzes and indexes documents ahead of time (Index Time).
The indexing pipeline reads the documents, splits text into tokens (Tokenization), removes irrelevant filler words (Stopwords), and records a lookup table: &lt;strong&gt;"Which files and positions contain this specific word?"&lt;/strong&gt;, much like the index at the back of a textbook.
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[ Source Documents ]
File_A: "Employee is entitled to sick leave for 30 days"
File_B: "Submitting sick leave requires a medical certificate"

[ Inverted Index Table ]
Term                 │ Posting List (Occurrences)
─────────────────────┼─────────────────────────────────────────────
certificate          │ File_B (pos 7)
days                 │ File_A (pos 8)
employee             │ File_A (pos 1)
leave                │ File_A (pos 5), File_B (pos 3)
medical              │ File_B (pos 6)
sick                 │ File_A (pos 4), File_B (pos 2)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;At query time (Query Time):&lt;/strong&gt;
When a user searches for &lt;code&gt;"sick leave"&lt;/code&gt;, the system does not rescan File A or File B from scratch. It directly looks up the table and returns matching documents in sub-millisecond time ($O(1)$ or $O(\log M)$, where $M$ is the vocabulary size of the index).&lt;/li&gt;
&lt;/ul&gt;


&lt;h3&gt;
  
  
  Limitations of Lexical Search: Why Exact Matching Falls Short
&lt;/h3&gt;

&lt;p&gt;While inverted indexes are fast and power production systems through algorithms like BM25, exact character matching suffers from three inherent language barriers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The Synonym Problem:&lt;/strong&gt;
If a policy document says &lt;em&gt;"Rules for absence while undergoing medical treatment"&lt;/em&gt; but the user searches for &lt;em&gt;"how to take sick leave"&lt;/em&gt;, the system returns zero results, even though the content is a 100% match in meaning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Polysemy Problem (Word Ambiguity):&lt;/strong&gt;
The same word can carry completely different meanings depending on context. For example, searching for &lt;em&gt;"crane"&lt;/em&gt; might return documents about construction machinery, a bird species, or an industrial pulley system without distinction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Typos &amp;amp; Natural Questions:&lt;/strong&gt;
A single typo can break a match. Furthermore, when users ask conversational questions like &lt;em&gt;"If someone gets burned in the cafeteria kitchen, who handles the medical claim?"&lt;/em&gt;, tokenizers break the query into fragmented words, which distorts keyword relevance scores.&lt;/li&gt;
&lt;/ol&gt;


&lt;h2&gt;
  
  
  3. Era 2: Searching by Meaning (Semantic / Vector Search)
&lt;/h2&gt;

&lt;p&gt;To overcome the limitations of spelling, modern AI leverages deep learning to create &lt;strong&gt;Text Embeddings&lt;/strong&gt; and Vector Search.&lt;/p&gt;

&lt;p&gt;In early hands-on experiments, when seeding test data with different wording, paraphrased sentences, or even cross-language queries, running a top-k similarity search using embedding models like &lt;code&gt;BAAI/bge-m3&lt;/code&gt; consistently retrieved documents with matching intent, despite having zero exact word overlap.&lt;/p&gt;
&lt;h3&gt;
  
  
  3.1 What is Text Embedding? (Turning Language into Coordinates)
&lt;/h3&gt;

&lt;p&gt;For those interested in exploring how text embeddings power Retrieval-Augmented Generation, check out &lt;a href="https://www.next4i.com/journey/en/what-is-rag-en" rel="noopener noreferrer"&gt;NEXT4I's Guide on RAG&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;An embedding model (such as OpenAI's &lt;code&gt;text-embedding-3&lt;/code&gt;, Cohere Embed, or open-source models like BAAI/bge) takes text passages as input and maps them into a high-dimensional dense vector, such as 768 or 1,536 dimensions.&lt;/p&gt;

&lt;p&gt;The core principle:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Passages with similar meaning or shared context end up close together as clusters in this multi-dimensional vector space."&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;


&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                     [ Semantic Cluster: Illness &amp;amp; Healthcare ]
                     • "Employee is feeling sick"
                     • "Medical leave guidelines"
                     • "Healthcare expense claims"

  [ Semantic Cluster: Finance ]              [ Semantic Cluster: Time Off ]
  • "Q3 balance sheet review"                • "Submitting annual leave"
  • "Withholding tax deduction"              • "Official public holidays"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3.2 Measuring Proximity Mathematically (Similarity Metrics)
&lt;/h3&gt;

&lt;p&gt;Once texts are converted into vectors, a computer determines conceptual similarity through geometry, most commonly via &lt;strong&gt;Cosine Similarity&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;$$\text{Cosine Similarity}(\vec{A}, \vec{B}) = \frac{\vec{A} \cdot \vec{B}}{|\vec{A}| |\vec{B}|}$$&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Value close to 1.0:&lt;/strong&gt; The vectors point in nearly identical directions, indicating closely aligned conceptual meaning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Value close to 0:&lt;/strong&gt; The vectors are orthogonal, indicating little to no contextual correlation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Query: &lt;em&gt;"Caught the flu, who should I notify?"&lt;/em&gt; $\rightarrow$ Vector $\vec{Q}$&lt;/li&gt;
&lt;li&gt;Document A: &lt;em&gt;"Procedure for submitting doctor's note to team lead"&lt;/em&gt; $\rightarrow$ Vector $\vec{D_1}$ (Cosine Similarity = &lt;strong&gt;0.88&lt;/strong&gt;)&lt;/li&gt;
&lt;li&gt;Document B: &lt;em&gt;"Setting up email passwords for sales department"&lt;/em&gt; $\rightarrow$ Vector $\vec{D_2}$ (Cosine Similarity = &lt;strong&gt;0.12&lt;/strong&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The search system immediately retrieves Document A because its cosine similarity is exceptionally high, &lt;strong&gt;even though Document A contains neither the word "flu" nor "notify".&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Why Not Convert an Entire 100-Page File into a Single Vector?
&lt;/h2&gt;

&lt;p&gt;A frequent question is: &lt;em&gt;"If vector search is so powerful, why not embed an entire 100-page PDF into one vector and be done with it?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Doing so fails in production for three key engineering reasons:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Information Dilution:&lt;/strong&gt;
If you take an employee handbook covering procurement, vacation policies, legal compliance, IT security, and annual bonuses, and compress all 100 pages into a single vector, that coordinate becomes a vague mathematical average of all topics. When a user asks a specific question, the combined vector will sit far away from the query. It is like blending dozens of different fruits into one container: you end up with a diluted mixture where individual flavors are unrecognizable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context Window Limits of Embedding Models:&lt;/strong&gt;
Embedding models have token input limits, such as 512, 2,048, or 8,192 tokens. They cannot ingest megabytes of raw text in a single pass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Precision for Answer Generation:&lt;/strong&gt;
When retrieving evidence, we need to pass only the specific, relevant passages into the language model. Feeding 100 pages of irrelevant noise increases latency, cost, and hallucination risks.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  The Solution: The Art of Chunking (Chunking Strategy)
&lt;/h3&gt;

&lt;p&gt;Before indexing documents, we split files into smaller passages (chunks):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Typically 300 to 500 words per chunk.&lt;/li&gt;
&lt;li&gt;A slight overlap (e.g., 50 words) to prevent context from breaking across chunk boundaries.&lt;/li&gt;
&lt;li&gt;Each chunk maintains a focused conceptual topic and receives its own unique vector representation.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  5. From Finding Content to Answering Questions: The 3-Step Pipeline
&lt;/h2&gt;

&lt;p&gt;Search and retrieval systems only identify &lt;strong&gt;where the evidence lives&lt;/strong&gt;. In real-world workflows, users rarely want ten raw links or disconnected paragraphs: they want a concise, reliable, and actionable answer.&lt;/p&gt;

&lt;p&gt;This is where &lt;strong&gt;Retrieval&lt;/strong&gt; connects with &lt;strong&gt;Generative AI (LLMs)&lt;/strong&gt; through a three-step pipeline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[ 1. User Asks a Question ]
  "Can an employee who is still on probation claim eyeglasses reimbursement?"
       │
       ▼
[ 2. Step 1: RETRIEVAL (Locate Relevant Chunks) ]
  - The system converts the user's question into a query vector.
  - Performs an ANN search across a Vector Database or Inverted Index.
  - Retrieves the top-k chunks with the highest relevance scores.
       │
       ▼
[ 3. Step 2: CONTEXT AUGMENTATION (Assemble Prompt for AI) ]
  - Combines the user query with the retrieved evidence passages:

    ┌────────────────────────────────────────────────────────┐
    │ SYSTEM: Answer the question using ONLY the provided    │
    │ context. Do not extrapolate. If the context does not   │
    │ contain the answer, reply with "Information not found".│
    │                                                        │
    │ CONTEXT:                                               │
    │ [Chunk #42]: Section 8.1 Eyeglasses Benefit            │
    │ The company provides an annual allowance of $100 for   │
    │ prescription glasses. This benefit is reserved exclusively│
    │ for permanent full-time employees who have successfully│
    │ passed their probationary period...                    │
    │                                                        │
    │ QUESTION: Can an employee on probation claim glasses?  │
    └────────────────────────────────────────────────────────┘
       │
       ▼
[ 4. Step 3: GENERATION &amp;amp; REASONING (Synthesize the Answer) ]
  - The LLM reads only the verified context provided.
  - Formulates a concise answer with source citation:

    👉 "No, employees on probation are not eligible. The optical
        allowance ($100/year) is reserved exclusively for permanent
        employees who have completed probation (Ref: Section 8.1)."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  6. Comparison: 4 Paradigms of Document Retrieval
&lt;/h2&gt;

&lt;p&gt;Here is a side-by-side comparison of how search methods have evolved:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Linear Scan (&lt;code&gt;Ctrl+F&lt;/code&gt; / &lt;code&gt;grep&lt;/code&gt;)&lt;/th&gt;
&lt;th&gt;Inverted Index (BM25 / Search Engines)&lt;/th&gt;
&lt;th&gt;Semantic Vector Search (Embeddings)&lt;/th&gt;
&lt;th&gt;Retrieval + LLM (RAG Paradigm)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Core Mechanism&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Scans and compares characters line by line&lt;/td&gt;
&lt;td&gt;Looks up pre-compiled word dictionary and postings&lt;/td&gt;
&lt;td&gt;Measures geometric distance between dense vectors&lt;/td&gt;
&lt;td&gt;Retrieves matching evidence chunks, then asks AI to synthesize&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Search Speed&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Slows down linearly with data size ($O(N)$)&lt;/td&gt;
&lt;td&gt;Sub-millisecond lookup&lt;/td&gt;
&lt;td&gt;Fast using Approximate Nearest Neighbor (ANN/HNSW)&lt;/td&gt;
&lt;td&gt;Moderate (depends on LLM inference time)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Intent Understanding&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;None (tracks term frequency and statistics)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Understands semantic concepts and context&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Deep contextual reasoning with synthesis&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Handling Synonyms &amp;amp; Typos&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Fails completely&lt;/td&gt;
&lt;td&gt;Limited (Fuzzy matching, lemmatization)&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Best-in-class&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Output Format&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Line positions or character highlights&lt;/td&gt;
&lt;td&gt;Document IDs and line offsets&lt;/td&gt;
&lt;td&gt;Relevant text chunks&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Direct synthesized answer with references&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  7. Conclusion: What Should Real-World Systems Use?
&lt;/h2&gt;

&lt;p&gt;When building production document intelligence systems, we do not have to pick one method and discard the rest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If you rely solely on &lt;strong&gt;Keyword Search&lt;/strong&gt;, you will miss documents whenever user terminology diverges from source wording.&lt;/li&gt;
&lt;li&gt;If you rely solely on &lt;strong&gt;Vector Search&lt;/strong&gt;, you may struggle with exact identifiers such as SKU numbers (&lt;code&gt;NK-9920-X&lt;/code&gt;), tax IDs, version strings, or domain-specific jargon absent from pre-trained embeddings.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is why modern enterprise search architectures adopt &lt;strong&gt;Hybrid Search&lt;/strong&gt;, combining lexical keyword indexes with vector embeddings, running the candidates through a Re-ranking stage, and passing the most relevant evidence to an LLM.&lt;/p&gt;

&lt;p&gt;At &lt;strong&gt;NEXT4I&lt;/strong&gt;, these search principles guide how we explore and architect document intelligence solutions, ensuring retrieval systems deliver speed, accuracy, and provable evidence. Stay tuned for future deep dives as we continue building and sharing our engineering journey.&lt;/p&gt;




&lt;p&gt;Explore the NEXT4I journey and read the original article at: &lt;a href="https://go.next4i.com/next4i-devto-en" rel="noopener noreferrer"&gt;https://go.next4i.com/next4i-devto-en&lt;/a&gt;&lt;/p&gt;

</description>
      <category>rag</category>
      <category>aie</category>
      <category>knowledgemanagement</category>
      <category>next4i</category>
    </item>
    <item>
      <title>RAG Architecture Beyond the Demo: Retrieval, Thai Chunking, and Production Boundaries</title>
      <dc:creator>NEXT4I DEV</dc:creator>
      <pubDate>Wed, 23 Sep 2026 15:37:19 +0000</pubDate>
      <link>https://dev.to/dev_next4i/what-is-rag-why-ai-needs-an-open-book-exam-igf</link>
      <guid>https://dev.to/dev_next4i/what-is-rag-why-ai-needs-an-open-book-exam-igf</guid>
      <description>&lt;p&gt;Connecting an LLM to an application is the easy part of a RAG demo.&lt;/p&gt;

&lt;p&gt;The difficult part starts when the source is a real document instead of a clean string in an array.&lt;/p&gt;

&lt;p&gt;I learned this while feeding a Thai tourism PDF into a document pipeline. (You can read more about that story here: "Why Markdown Is the Ultimate AI-Native File Format A War Story from Building NEXT4I" &lt;a href="https://www.next4i.com/dev-notes/en/markdown-the-ultimate-ai-native-dev-en" rel="noopener noreferrer"&gt;https://www.next4i.com/dev-notes/en/markdown-the-ultimate-ai-native-dev-en&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Direct text extraction produced fragmented characters and misplaced vowels or tone marks. Rendering pages as images preserved the layout but introduced colorful backgrounds and watermarks. Converting them to black and white made some text clearer while removing useful context from charts and photographs. Different models produced different interpretations, so the outputs still required synthesis and human review.&lt;/p&gt;

&lt;p&gt;Think of it like OCR processing for an ID card selfie: even with modern vision models, apps still often need human input fields to verify or correct the data. Document ingestion is just step one of RAG, and this is where the real world begins.&lt;/p&gt;

&lt;p&gt;That experience changed how I look at Retrieval-Augmented Generation. RAG is not only a prompt pattern. It is a data and retrieval system with an LLM at the end.&lt;/p&gt;

&lt;p&gt;This article unpacks that system through:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;separate ingestion and query pipelines&lt;/li&gt;
&lt;li&gt;document extraction and chunking&lt;/li&gt;
&lt;li&gt;embeddings and retrieval&lt;/li&gt;
&lt;li&gt;dense, sparse, and hybrid search&lt;/li&gt;
&lt;li&gt;access control in the retrieval path&lt;/li&gt;
&lt;li&gt;evaluation boundaries&lt;/li&gt;
&lt;li&gt;a conceptual implementation in Python&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The code below is generic and sanitized. It demonstrates mechanics, not a production-ready application. Provider names, model IDs, SDK calls, and executable integration still need verification against the versions you use.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three limits RAG is trying to address
&lt;/h2&gt;

&lt;p&gt;When an application uses an LLM out of the box, three limits appear quickly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Knowledge boundaries
&lt;/h3&gt;

&lt;p&gt;The model does not automatically know a new event, a private company policy, or a document created yesterday.&lt;/p&gt;

&lt;h3&gt;
  
  
  Context and cost boundaries
&lt;/h3&gt;

&lt;p&gt;Passing an entire document collection in every request is usually impractical. A large context also does not guarantee that the model will use every passage equally well.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verifiability
&lt;/h3&gt;

&lt;p&gt;A generated answer does not prove itself. If a user cannot trace a statement back to a permitted source, confidence is not evidence.&lt;/p&gt;

&lt;p&gt;RAG addresses these limits by finding a small set of relevant passages and placing them in the context before generation. It can improve grounding and traceability, but it does not guarantee either accuracy or security.&lt;/p&gt;

&lt;h2&gt;
  
  
  RAG is two pipelines, not one request
&lt;/h2&gt;

&lt;p&gt;A useful mental model separates offline ingestion from online querying.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[Ingestion pipeline]

Raw documents
    |
    v
Extract and clean
    |
    v
Split into chunks
    |
    v
Create embeddings and metadata
    |
    v
Store in a searchable index


[Query pipeline]

User question
    |
    v
Apply identity and permission scope
    |
    v
Retrieve candidate chunks
    |
    v
Filter and optionally rerank
    |
    v
Build prompt with citations
    |
    v
Generate answer
    |
    v
Return answer and inspectable sources
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pipelines are decoupled for a reason. Documents change on their own schedule. User queries arrive on another. Extraction failures should not be rediscovered during every chat request, and permission changes should not require retraining a model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ingestion quality becomes retrieval quality
&lt;/h2&gt;

&lt;p&gt;The phrase “garbage in, garbage out” is almost too familiar, but it describes RAG accurately.&lt;/p&gt;

&lt;p&gt;A document can look correct to a human and still become poor retrieval material after extraction. Typical problems include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;repeated headers and footers&lt;/li&gt;
&lt;li&gt;tables flattened into an unreadable order&lt;/li&gt;
&lt;li&gt;text stored as positioned glyphs rather than paragraphs&lt;/li&gt;
&lt;li&gt;watermarks mixed with body text&lt;/li&gt;
&lt;li&gt;diagrams whose meaning exists only in the visual layout&lt;/li&gt;
&lt;li&gt;scanned pages without a reliable text layer&lt;/li&gt;
&lt;li&gt;old and current versions indexed together&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The ingestion pipeline needs to preserve more than plain text. Useful metadata can include a stable document ID, version, page, section heading, timestamps, source location, tenant, and permission attributes.&lt;/p&gt;

&lt;p&gt;That metadata supports citations, updates, deletion, filtering, and review. Without it, a retrieved paragraph becomes an orphan.&lt;/p&gt;

&lt;h2&gt;
  
  
  Thai documents make naive chunking easier to break
&lt;/h2&gt;

&lt;p&gt;Many chunking examples assume that spaces and punctuation provide reliable boundaries. Thai text does not always behave that way.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ambiguous boundaries
&lt;/h3&gt;

&lt;p&gt;Consider the Thai text:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ตากลมนั่งมองตากลม
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Depending on segmentation and context, &lt;code&gt;ตากลม&lt;/code&gt; can relate to round eyes, while &lt;code&gt;ตากลม&lt;/code&gt; can also be read through a boundary involving exposure to wind. A bad split changes the meaning represented by the chunk and therefore affects retrieval.&lt;/p&gt;

&lt;h3&gt;
  
  
  Long dependencies
&lt;/h3&gt;

&lt;p&gt;A policy sentence may introduce the subject and action early, then place the condition or consequence much later. A fixed-size cut can leave one chunk with the violation and another with the disciplinary action. Neither passage is complete enough to answer the question reliably.&lt;/p&gt;

&lt;h3&gt;
  
  
  Periods that are not sentence endings
&lt;/h3&gt;

&lt;p&gt;Thai abbreviations, titles, legal terms, locations, and time expressions can contain periods:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ศ.ดร.สมชาย ... พ.ร.บ. ... อ.เมือง จ.เชียงใหม่ ... 09.00 น.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A generic splitter that treats every period as an end of sentence can create tiny, meaningless chunks.&lt;/p&gt;

&lt;p&gt;Possible approaches from the original engineering notes include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;semantic chunking around topic changes&lt;/li&gt;
&lt;li&gt;Thai-aware tokenization, for example with PyThaiNLP, plus a domain dictionary&lt;/li&gt;
&lt;li&gt;overlap around chunk boundaries&lt;/li&gt;
&lt;li&gt;hierarchy-aware parent and child chunks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The correct strategy and parameters are corpus-specific. Any fixed chunk size or overlap percentage should be evaluated against real questions rather than copied as a universal setting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Chunking strategies and their trade-offs
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Fixed size with overlap
&lt;/h3&gt;

&lt;p&gt;This is simple and predictable. It also cuts across structure when the chosen length does not match the document.&lt;/p&gt;

&lt;h3&gt;
  
  
  Recursive splitting
&lt;/h3&gt;

&lt;p&gt;This method tries larger structural separators first, such as sections and paragraphs, then falls back to smaller boundaries. It works better when the extracted structure is trustworthy.&lt;/p&gt;

&lt;h3&gt;
  
  
  Semantic chunking
&lt;/h3&gt;

&lt;p&gt;This approach compares nearby passages and creates a new chunk when the topic shifts. It can preserve meaning better, but it adds model dependency, threshold tuning, and processing cost.&lt;/p&gt;

&lt;h3&gt;
  
  
  Parent-child chunking
&lt;/h3&gt;

&lt;p&gt;Small child chunks can support precise retrieval while a larger parent section supplies enough context to answer. The trade-off is more complex indexing, deduplication, and prompt assembly.&lt;/p&gt;

&lt;p&gt;Chunking should be treated as a retrieval decision, not a formatting task. The unit you retrieve determines the evidence the model can see.&lt;/p&gt;

&lt;h2&gt;
  
  
  Embeddings are coordinates, not facts
&lt;/h2&gt;

&lt;p&gt;An embedding converts content into a numeric vector. Content with related meaning may appear near each other in that vector space, depending on the model and data.&lt;/p&gt;

&lt;p&gt;That makes semantic retrieval possible, but an embedding does not verify truth. It also does not know that an older policy is invalid unless version and filtering logic provide that boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conceptual implementation in Python (In-Memory Retrieval &amp;amp; Query Pipeline)
&lt;/h2&gt;

&lt;p&gt;To illustrate how a query pipeline works in one place, the following Python example simulates cosine similarity calculation, user permission and tenant filtering prior to retrieval, top-K ranking, and prompt augmentation for the LLM:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;cosine_similarity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vec_a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;vec_b&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Compute cosine similarity between two numeric vectors.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vec_a&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vec_b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vec_a&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Vectors must have the same non-zero length&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;dot_product&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vec_a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;vec_b&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;norm_a&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sqrt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;vec_a&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;norm_b&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sqrt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;vec_b&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;norm_a&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;norm_b&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Cosine similarity is undefined for a zero vector&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;dot_product&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;norm_a&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;norm_b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;answer_with_rag&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;knowledge_base&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;top_k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;create_embedding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;generate_answer&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Demonstrate an end-to-end RAG query flow with permission scoping.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="c1"&gt;# 1. Convert user question into an embedding vector
&lt;/span&gt;    &lt;span class="n"&gt;query_vector&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;create_embedding&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# 2. Apply identity and permission scope before retrieval
&lt;/span&gt;    &lt;span class="c1"&gt;# Critical: The model must never receive context the user is not allowed to see
&lt;/span&gt;    &lt;span class="n"&gt;authorized_chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;knowledge_base&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tenant_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tenant_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;allowed_roles&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="c1"&gt;# 3. Retrieval: calculate similarity and rank Top-K
&lt;/span&gt;    &lt;span class="n"&gt;ranked_chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;score&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;cosine_similarity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query_vector&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;embedding&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])}&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;authorized_chunks&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;ranked_chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sort&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;score&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;reverse&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;selected_chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ranked_chunks&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="n"&gt;top_k&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="c1"&gt;# 4. Augmentation: build context with verifiable source citations
&lt;/span&gt;    &lt;span class="n"&gt;context_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[ID: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;] &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;selected_chunks&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

    &lt;span class="c1"&gt;# 5. Generation: ask the model to ground its response in the provided context
&lt;/span&gt;    &lt;span class="n"&gt;instruction&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Answer only from the supplied context. If the context is insufficient, state that clearly.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;generate_answer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;instruction&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;instruction&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;context_text&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sources&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;score&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;score&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;selected_chunks&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;The core takeaway is that &lt;strong&gt;permission filtering must happen before passages become model context&lt;/strong&gt;. In production enterprise systems, this is handled through row-level security (RLS), policy engines, document ACLs, or metadata filtering on a vector store rather than an in-memory loop.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Dense search is not enough for every query
&lt;/h2&gt;

&lt;p&gt;Dense vector search is useful for semantic similarity. It can connect related wording such as synonyms and paraphrases.&lt;/p&gt;

&lt;p&gt;Exact identifiers are a different problem. Product SKUs, serial numbers, legal section numbers, error codes, and names can be better served by sparse keyword retrieval.&lt;/p&gt;

&lt;p&gt;That creates three common choices:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Dense retrieval:&lt;/strong&gt; strong for semantic similarity&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sparse retrieval such as BM25:&lt;/strong&gt; strong for exact terms&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hybrid retrieval:&lt;/strong&gt; combines candidates from both systems, then fuses or reranks them&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Reciprocal Rank Fusion and cross-encoder reranking are possible tools in that pipeline. They are not automatic improvements for every corpus. Evaluate them using your own question set and latency boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five production boundaries worth testing
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Language and embedding fit
&lt;/h3&gt;

&lt;p&gt;Do not assume an embedding model that looks good on English examples will behave the same way on Thai policies, abbreviations, and domain terms. Compare models against labeled queries from the actual corpus.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Extraction and version quality
&lt;/h3&gt;

&lt;p&gt;Track failed pages, missing sections, table quality, repeated boilerplate, and document versions. Retrieval cannot recover content that ingestion lost.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Context pollution
&lt;/h3&gt;

&lt;p&gt;More chunks are not always better. Irrelevant context can distract the model and increase cost. Evaluate the complete answer path rather than maximizing &lt;code&gt;topK&lt;/code&gt; by intuition.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Security and data isolation
&lt;/h3&gt;

&lt;p&gt;Permission-aware retrieval needs to be enforced outside the model. Test cross-tenant access, role changes, deleted permissions, cached results, and citation links.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Evaluation
&lt;/h3&gt;

&lt;p&gt;The original notes use three useful questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context relevance:&lt;/strong&gt; Did retrieval find evidence related to the question?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Groundedness or faithfulness:&lt;/strong&gt; Does the answer stay within the supplied evidence?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answer relevance:&lt;/strong&gt; Does it answer what the user asked?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Add operational checks that match the system, such as extraction failures, stale indexes, permission denials, and citation integrity. Do not invent target scores. Establish them from your risk and user requirements.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main engineering lesson
&lt;/h2&gt;

&lt;p&gt;RAG connects language-model capability with external knowledge, but the connection is only as useful as the pipeline around it.&lt;/p&gt;

&lt;p&gt;A local demo can hide document quality, authorization, stale versions, and evaluation because the sample chunks are already clean. Production exposes all of them.&lt;/p&gt;

&lt;p&gt;For me, the work starts before prompt engineering:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;prepare and version the knowledge&lt;/li&gt;
&lt;li&gt;choose chunk boundaries that preserve meaning&lt;/li&gt;
&lt;li&gt;retrieve with semantic and exact-match needs in mind&lt;/li&gt;
&lt;li&gt;apply permissions before context construction&lt;/li&gt;
&lt;li&gt;preserve citations&lt;/li&gt;
&lt;li&gt;evaluate retrieval and generation separately&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The LLM writes the final response. Most of the evidence it can use has already been decided by then.&lt;/p&gt;




&lt;p&gt;Explore the NEXT4I journey and read the original article at: &lt;a href="https://go.next4i.com/next4i-devto-en" rel="noopener noreferrer"&gt;https://go.next4i.com/next4i-devto-en&lt;/a&gt;&lt;/p&gt;

</description>
      <category>rag</category>
      <category>knowledgemanagement</category>
      <category>llm</category>
      <category>next4i</category>
    </item>
    <item>
      <title>How to Write an AI Agent Skill File That Reduces Guesswork</title>
      <dc:creator>NEXT4I DEV</dc:creator>
      <pubDate>Mon, 14 Sep 2026 16:58:08 +0000</pubDate>
      <link>https://dev.to/dev_next4i/how-to-write-an-ai-agent-skill-file-that-reduces-guesswork-2cn9</link>
      <guid>https://dev.to/dev_next4i/how-to-write-an-ai-agent-skill-file-that-reduces-guesswork-2cn9</guid>
      <description>&lt;p&gt;A prompt can tell an agent what to do once. A skill file can document how a team handles that class of work repeatedly.&lt;/p&gt;

&lt;p&gt;For developers, the distinction matters. An agent may have repository access and the right tools, yet still has to make decisions about scope, references, validation, and failure handling. If those decisions are undefined, a technically valid result can still be the wrong result.&lt;/p&gt;

&lt;p&gt;This guide builds a small, generic skill for reviewing API changes. The point is not to prescribe one universal Agent Skills format. It is to show how to turn vague intent into an executable workflow.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Skill discovery, instruction loading, and tool use vary by model, agent runtime, system instructions, and available tools. Treat the structure below as a portable design pattern, then adapt it to your platform.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The problem with a reasonable instruction
&lt;/h2&gt;

&lt;p&gt;Consider this prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Review this API and make sure it follows best practices.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It sounds clear, but the agent still has to decide:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which API contract is authoritative?&lt;/li&gt;
&lt;li&gt;Does the review include authentication and permissions?&lt;/li&gt;
&lt;li&gt;Should it modify code or only produce a report?&lt;/li&gt;
&lt;li&gt;Which tests should it run?&lt;/li&gt;
&lt;li&gt;What should it do when information is missing?&lt;/li&gt;
&lt;li&gt;What does a completed review contain?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Different models may fill those gaps differently. The result can be reasonable without matching the workflow you intended.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prompt versus skill
&lt;/h2&gt;

&lt;p&gt;I use this distinction:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A prompt says, “Do this task now.”&lt;/li&gt;
&lt;li&gt;A skill says, “This is how we handle this type of task.”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A useful skill does not make a model magically smarter. It makes the operating boundaries visible.&lt;/p&gt;

&lt;p&gt;At minimum, it should define:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Trigger: when the skill applies&lt;/li&gt;
&lt;li&gt;Goal: what outcome it should produce&lt;/li&gt;
&lt;li&gt;Scope: what is included and excluded&lt;/li&gt;
&lt;li&gt;Workflow: steps the agent can follow&lt;/li&gt;
&lt;li&gt;References: information to read under specific conditions&lt;/li&gt;
&lt;li&gt;Constraints: actions and claims that are prohibited&lt;/li&gt;
&lt;li&gt;Uncertainty handling: what to do when evidence is missing&lt;/li&gt;
&lt;li&gt;Completion criteria: how to validate the result&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  A practical directory structure
&lt;/h2&gt;

&lt;p&gt;Start with one file while the workflow is small. Split it when references, examples, or reusable scripts become large enough to distract from the main instructions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;api-review-skill/
├── SKILL.md
├── references/
│   ├── api-contract.md
│   ├── security-rules.md
│   ├── output-examples.md
│   └── troubleshooting.md
├── scripts/
│   └── validate.sh
└── assets/
    └── review-template.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact file names and supported directories depend on the platform and runtime. More importantly, models may not select or load these files in the same way. The main file should therefore explain why and when each reference matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  A generic &lt;code&gt;SKILL.md&lt;/code&gt; example
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;api-change-review&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Review&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;API&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;changes&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;contract&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;compatibility,"&lt;/span&gt;
  &lt;span class="s"&gt;security boundaries, and required validation. Use when an&lt;/span&gt;
  &lt;span class="s"&gt;endpoint, request schema, response schema, authentication,&lt;/span&gt;
  &lt;span class="s"&gt;or permission behavior changes.&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="gh"&gt;# API Change Review&lt;/span&gt;

&lt;span class="gu"&gt;## Goal&lt;/span&gt;

Produce a review report that identifies contract-breaking changes,
security-sensitive changes, missing validation, and unresolved assumptions.

&lt;span class="gu"&gt;## Scope&lt;/span&gt;

Review the proposed change and related tests.
Do not modify implementation files unless the user requests edits.

&lt;span class="gu"&gt;## Workflow&lt;/span&gt;
&lt;span class="p"&gt;
1.&lt;/span&gt; Read the changed files and the API contract.
&lt;span class="p"&gt;2.&lt;/span&gt; Identify changes to endpoints, fields, status codes, and behavior.
&lt;span class="p"&gt;3.&lt;/span&gt; If authentication, secrets, or permissions are involved,
   read the security rules.
&lt;span class="p"&gt;4.&lt;/span&gt; Compare the tests with the changed behavior.
&lt;span class="p"&gt;5.&lt;/span&gt; Produce the required report.
&lt;span class="p"&gt;6.&lt;/span&gt; Run the validation command when the environment supports it.

&lt;span class="gu"&gt;## Reference routing&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Read &lt;span class="sb"&gt;`references/api-contract.md`&lt;/span&gt; for endpoint or schema changes.
&lt;span class="p"&gt;-&lt;/span&gt; Read &lt;span class="sb"&gt;`references/security-rules.md`&lt;/span&gt; for authentication,
  secrets, or permission changes.
&lt;span class="p"&gt;-&lt;/span&gt; Read &lt;span class="sb"&gt;`references/output-examples.md`&lt;/span&gt; only when the report format
  is unclear.
&lt;span class="p"&gt;-&lt;/span&gt; Read &lt;span class="sb"&gt;`references/troubleshooting.md`&lt;/span&gt; when validation fails.

&lt;span class="gu"&gt;## Constraints&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Do not invent endpoints, fields, metrics, or test results.
&lt;span class="p"&gt;-&lt;/span&gt; Do not expose credentials, customer data, or internal URLs.
&lt;span class="p"&gt;-&lt;/span&gt; Do not present an unimplemented requirement as released behavior.
&lt;span class="p"&gt;-&lt;/span&gt; Flag assumptions that require human review.

&lt;span class="gu"&gt;## Required output&lt;/span&gt;
&lt;span class="p"&gt;
1.&lt;/span&gt; Summary
&lt;span class="p"&gt;2.&lt;/span&gt; Contract-breaking changes
&lt;span class="p"&gt;3.&lt;/span&gt; Security-sensitive changes
&lt;span class="p"&gt;4.&lt;/span&gt; Missing or affected tests
&lt;span class="p"&gt;5.&lt;/span&gt; Assumptions requiring review
&lt;span class="p"&gt;6.&lt;/span&gt; Validation performed

&lt;span class="gu"&gt;## Completion criteria&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Every changed behavior is mapped to the contract.
&lt;span class="p"&gt;-&lt;/span&gt; Security-sensitive changes are explicitly identified.
&lt;span class="p"&gt;-&lt;/span&gt; Validation results state what was and was not run.
&lt;span class="p"&gt;-&lt;/span&gt; Unresolved assumptions are visible.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This example is intentionally generic. It demonstrates an instruction contract, not a claim about one platform's required syntax.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before and after
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Before
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Review this API and make sure it follows best practices.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  After
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Review the API change using the api-change-review skill.

Use `references/api-contract.md` as the contract source.
If the change touches authentication, secrets, or permissions,
also apply `references/security-rules.md`.

Produce the required review report. Do not modify implementation files.
State which validation commands were run and flag unresolved assumptions.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second instruction is not better because it is longer. It is better because fewer operational decisions are left undefined.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turn the index into a routing table
&lt;/h2&gt;

&lt;p&gt;A file list tells an agent what exists. A routing table explains when a file becomes relevant.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Reference&lt;/th&gt;
&lt;th&gt;Read when&lt;/th&gt;
&lt;th&gt;Skip when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;api-contract.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;An endpoint, schema, status code, or behavior changes&lt;/td&gt;
&lt;td&gt;The task does not involve an API contract&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;security-rules.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Authentication, secrets, roles, or permissions are involved&lt;/td&gt;
&lt;td&gt;The task has no security-sensitive behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;output-examples.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The required report shape is unclear&lt;/td&gt;
&lt;td&gt;The output contract is already explicit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;troubleshooting.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Validation cannot run or fails&lt;/td&gt;
&lt;td&gt;Validation succeeds normally&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This table is guidance, not an enforcement mechanism. Actual reference selection still depends on the model, runtime, context strategy, and tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Single file or multiple files?
&lt;/h2&gt;

&lt;p&gt;Keep the skill in one file when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;it has one short workflow&lt;/li&gt;
&lt;li&gt;all constraints fit without hiding the main steps&lt;/li&gt;
&lt;li&gt;examples are small&lt;/li&gt;
&lt;li&gt;no reusable scripts are needed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Split the skill when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;references are long or domain-specific&lt;/li&gt;
&lt;li&gt;several workflows share the same rules&lt;/li&gt;
&lt;li&gt;examples make the main file difficult to scan&lt;/li&gt;
&lt;li&gt;validation scripts are reusable&lt;/li&gt;
&lt;li&gt;sensitive rules require separate ownership or review&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The main file should remain enough to answer two questions: What should the agent do, and where should it look next?&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat skills like software
&lt;/h2&gt;

&lt;p&gt;A skill can contain instructions, references, and executable code. Review an external skill before using it, especially when it can access files, call a network service, or run scripts.&lt;/p&gt;

&lt;p&gt;Check for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;file access outside the expected scope&lt;/li&gt;
&lt;li&gt;network and API destinations&lt;/li&gt;
&lt;li&gt;hardcoded credentials&lt;/li&gt;
&lt;li&gt;destructive file operations&lt;/li&gt;
&lt;li&gt;hidden instructions that attempt to bypass system rules&lt;/li&gt;
&lt;li&gt;scripts whose behavior has not been reviewed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For organizational use, sandboxing and coexistence tests may also be appropriate. Different models can respond to the same instruction or tool surface differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Validation checklist
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="p"&gt;-&lt;/span&gt; [ ] The description states when the skill should trigger.
&lt;span class="p"&gt;-&lt;/span&gt; [ ] The goal and scope are explicit.
&lt;span class="p"&gt;-&lt;/span&gt; [ ] The workflow contains observable steps.
&lt;span class="p"&gt;-&lt;/span&gt; [ ] Each reference has a condition for when to read it.
&lt;span class="p"&gt;-&lt;/span&gt; [ ] Prohibited actions and claims are listed.
&lt;span class="p"&gt;-&lt;/span&gt; [ ] Missing information has a defined handling rule.
&lt;span class="p"&gt;-&lt;/span&gt; [ ] The output contract is reusable and testable.
&lt;span class="p"&gt;-&lt;/span&gt; [ ] Completion criteria can be checked.
&lt;span class="p"&gt;-&lt;/span&gt; [ ] No secrets or internal URLs are embedded.
&lt;span class="p"&gt;-&lt;/span&gt; [ ] Executable scripts have been reviewed.
&lt;span class="p"&gt;-&lt;/span&gt; [ ] The skill was tested on tasks that should trigger it.
&lt;span class="p"&gt;-&lt;/span&gt; [ ] The skill was tested on tasks that should not trigger it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Final takeaway
&lt;/h2&gt;

&lt;p&gt;Before handing a skill to an agent, read it as if you were a developer joining the project today:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can I complete the work from this information? What would I still have to guess? If I must make a decision, do I know which direction to take?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A good skill does not remove reasoning. It removes avoidable guessing.&lt;/p&gt;

&lt;p&gt;Read the canonical article on NEXT4I: &lt;a href="https://go.next4i.com/next4i-devto-en" rel="noopener noreferrer"&gt;https://go.next4i.com/next4i-devto-en&lt;/a&gt;&lt;/p&gt;

</description>
      <category>aiskill</category>
      <category>ai</category>
      <category>llm</category>
      <category>next4i</category>
    </item>
    <item>
      <title>What Is Token &amp; LLM Cost Optimization</title>
      <dc:creator>NEXT4I DEV</dc:creator>
      <pubDate>Mon, 07 Sep 2026 14:35:04 +0000</pubDate>
      <link>https://dev.to/dev_next4i/what-is-token-llm-cost-optimization-398l</link>
      <guid>https://dev.to/dev_next4i/what-is-token-llm-cost-optimization-398l</guid>
      <description>&lt;h1&gt;
  
  
  What Is a Token in LLMs? A Developer's Guide to Cost Optimization and Architecture
&lt;/h1&gt;

&lt;p&gt;Understanding how LLM tokens work is the difference between an AI feature that costs $50/month and one that runs up a $5,000 AWS/OpenAI bill.&lt;/p&gt;

&lt;p&gt;Here is the engineering breakdown of tokenization algorithms, input vs. output pricing asymmetries, multilingual overhead, model pricing dynamics, and six production strategies implemented at NEXT4I to cut inference costs by up to 50-80%.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;code&gt;#Buildinpublic&lt;/code&gt;, &lt;code&gt;#ModelAI&lt;/code&gt;, &lt;code&gt;#LLM&lt;/code&gt;, &lt;code&gt;#AIArchitecture&lt;/code&gt;, &lt;code&gt;#TokenOptimization&lt;/code&gt;, &lt;code&gt;#AIDeveloper&lt;/code&gt;&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  1. What Is a Token Under the Hood?
&lt;/h2&gt;

&lt;p&gt;LLMs do not process raw text or strings. They consume &lt;strong&gt;Tokens&lt;/strong&gt;—subword representations mapped to high-dimensional embedding vectors via algorithms like &lt;strong&gt;Byte-Pair Encoding (BPE)&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Hello world" -&amp;gt; ["Hello", " world"] (2 tokens)
"Unstoppable" -&amp;gt; ["Un", "stoppable"] (2 tokens)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The Multilingual Token Penalty
&lt;/h3&gt;

&lt;p&gt;Because vocabulary dictionaries are predominantly trained on English corpora, languages without whitespace delimitation (such as Thai) suffer severe subword fragmentation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;English: &lt;code&gt;1 Token ≈ 0.75 words&lt;/code&gt; (~4 characters).&lt;/li&gt;
&lt;li&gt;Thai: &lt;code&gt;1 Word ≈ 3 to 8 Tokens&lt;/code&gt; (frequently split into byte-level representations).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A prompt written in Thai can cost up to 6x more in raw token usage and introduce noticeable latency compared to its English equivalent.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Formatting Mechanics: Linebreaks &amp;amp; The "Emoji Tax"
&lt;/h2&gt;

&lt;p&gt;Every character in your prompt payload carries a token cost:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Newlines (&lt;code&gt;\n&lt;/code&gt;) and Numbered Markdown:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Consumes 1 token per newline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verdict:&lt;/strong&gt; Highly recommended. Clear formatting provides structural anchors for transformer attention heads, significantly reducing hallucination.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;The Emoji Tax:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Emojis are complex multibyte Unicode sequences, often consuming &lt;strong&gt;2 to 6+ tokens each&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Avoid embedding decorative emojis in static system prompts that execute millions of times.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  3. Input vs. Output Tokens: Why Output Costs 3x–5x More
&lt;/h2&gt;

&lt;p&gt;API rate cards price Output Tokens significantly higher than Input Tokens. Why?&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Input Tokens (Parallelized Compute):&lt;/strong&gt; Processed simultaneously across GPU tensor cores in a single matrix multiplication pass (Prefill / Encoding). It is fast and hardware-efficient.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output Tokens (Sequential Autoregression):&lt;/strong&gt; Generated one token at a time (Autoregressive Decoding). The model generates token N, appends it back to the context history, and re-computes attention for token N+1. This locks GPU resources over the entire generation cycle.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  4. Why Model Pricing Varies by 100x (Dense vs. MoE Architecture)
&lt;/h2&gt;

&lt;p&gt;You may have noticed that API pricing spans from $0.30 to $50.00+ per 1 million tokens across models:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek-V4:&lt;/strong&gt; $0.50 – $1.70 / 1M Tokens (Massive architecture, exceptionally cost-efficient)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 3.7 Flash:&lt;/strong&gt; $0.30 – $1.80 / 1M Tokens (High-efficiency edge tier)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Sonnet 5:&lt;/strong&gt; $2.00 – $10.00 / 1M Tokens (Mid-tier balanced powerhouse)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI GPT-5.6 Terra:&lt;/strong&gt; $2.00 – $12.00 / 1M Tokens (Enterprise mid-tier)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Opus 5:&lt;/strong&gt; $5.00 – $25.00 / 1M Tokens (Premium reasoning class)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Fable 5:&lt;/strong&gt; $10.00 – $50.00 / 1M Tokens (Mythos-class model from Anthropic)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI GPT-5.6 Sol:&lt;/strong&gt; $4.00 – $30.00 / 1M Tokens (OpenAI's flagship)&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Model pricing per token fluctuates frequently across providers and should be used strictly for relative comparative analysis.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This pricing variance stems from three fundamental drivers:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Architectural Design: Dense Models vs. Mixture-of-Experts (MoE)
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Dense Models
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mechanics:&lt;/strong&gt; Every incoming vector/token is processed through &lt;strong&gt;100% of the model's parameters&lt;/strong&gt;, from the initial layer to the final output layer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Analogy:&lt;/strong&gt; Like a company where every single employee must sit in every meeting and vote on every decision—regardless of whether the task is simple arithmetic or legal compliance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trade-offs:&lt;/strong&gt; Highly compute-intensive (massive FLOPs per token), resulting in higher per-token inference costs. However, memory management and GPU scheduling remain straightforward.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Mixture-of-Experts (MoE)
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mechanics:&lt;/strong&gt; An intelligent gating mechanism called a &lt;strong&gt;Router (or Gating Network)&lt;/strong&gt; acts as a dispatcher, evaluating incoming tokens and dynamically routing them to specialized sub-networks (&lt;strong&gt;Experts&lt;/strong&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Execution Flow:&lt;/strong&gt;

&lt;ol&gt;
&lt;li&gt;The Router analyzes each token and activates only a sparse subset of experts (e.g., selecting 2 out of 8 or 64 total experts).&lt;/li&gt;
&lt;li&gt;Compute flows exclusively through the parameters of the chosen experts.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Analogy:&lt;/strong&gt; Like a well-structured organization with an executive dispatcher. A calculus question is routed strictly to the math experts without distracting the linguistics team.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Advantages:&lt;/strong&gt; Enables total model capacity (Total Parameters) to scale massively while keeping active compute per token (Active Parameters / FLOPs) extremely low. This allows lightning-fast generation and drastically lower API prices (e.g., DeepSeek, Mixtral).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trade-offs:&lt;/strong&gt; Requires massive VRAM/RAM pools to keep all expert weights loaded in memory simultaneously.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  The MoE Achilles' Heel: When Routers Fail
&lt;/h3&gt;

&lt;p&gt;While MoE unlocks unmatched cost efficiency, its performance is tightly bound to routing stability:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Loss of Nuance &amp;amp; Context Disruption:&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Natural language is rich with subtle subtext. If a router misinterprets a token and dispatches it to the wrong expert, nuanced meaning collapses. The output may stay grammatically intact but lose analytical depth or fail to address the core prompt intent.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Router Collapse &amp;amp; Expert Imbalance:&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Routers naturally develop bias toward a handful of frequently trained experts, causing severe load imbalance. The favored experts hit computational bottlenecks while neglected experts become &lt;strong&gt;Dead Parameters&lt;/strong&gt;, defeating the entire purpose of modular specialization.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cascading Errors:&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;LLMs process representations layer by layer. If an early-layer router misroutes a token, downstream layers receive corrupted intermediate activations, amplifying routing errors across subsequent layers.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;h4&gt;
  
  
  🛠️ Engineering Safeguards Used by Frontier Labs:
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Auxiliary Load Balancing Loss:&lt;/strong&gt; Incorporating penalty penalties into the loss function during pre-training to enforce uniform token distribution across all expert sub-networks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capacity Factor Enforcement:&lt;/strong&gt; Setting strict token buffer caps per expert. Once an expert's capacity threshold is reached, excess tokens are spilled over to secondary experts to prevent execution bottlenecks.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  2. Reasoning Overhead (Thinking / Chain-of-Thought Tokens)
&lt;/h3&gt;

&lt;p&gt;Reasoning-focused models (like Claude Opus, OpenAI GPT Sol) generate thousands of internal, hidden &lt;strong&gt;Chain-of-Thought tokens&lt;/strong&gt; before emitting their first visible output token. Providers meter and bill for every single background reasoning step.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Hardware Sovereignty &amp;amp; Custom Silicon
&lt;/h3&gt;

&lt;p&gt;Hyperscalers operating proprietary custom silicon (such as Google’s TPU clusters for Gemini) achieve significantly lower baseline operating costs than providers renting general-purpose Nvidia H100/H200 GPU clusters.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. 6 Production Strategies to Cut Token Costs by 80%
&lt;/h2&gt;

&lt;p&gt;Here are some of the production-level strategies we use in building NEXT4I:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Prompt Engineering for Token Efficiency
&lt;/h3&gt;

&lt;p&gt;Eliminate fluff and instructions that don't add semantic value. Use concise formats like YAML or Markdown instead of verbose JSON schemas.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Leverage Prompt Caching
&lt;/h3&gt;

&lt;p&gt;Major LLM providers (Anthropic, OpenAI, DeepSeek, Google) offer prompt caching. Placing static context (system instructions, tool definitions, schemas) at the prompt root allows providers to cache the KV-cache, reducing input costs by &lt;strong&gt;75%–90%&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Intelligent Model Cascading (Model Routing)
&lt;/h3&gt;

&lt;p&gt;Never route every query to flagship models. Route simple classification, data extraction, and formatting tasks to lightweight models or budget-friendly models (Gemini Flash, DeepSeek), escalating only complex reasoning tasks to larger models (Claude Sonnet, GPT Terra, GPT Sol, Claude Opus / Fable).&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Sliding Context Windows &amp;amp; Conversation Pruning
&lt;/h3&gt;

&lt;p&gt;Chat histories grow quadratically &lt;code&gt;($O(n^2)$)&lt;/code&gt; if sent in their entirety on every turn. Maintain a rolling sliding window of the last 5–10 messages, or summarize older context into a single concise paragraph.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Pre-Retrieval RAG Filtering
&lt;/h3&gt;

&lt;p&gt;In RAG pipelines, do not inject full documents into the context window. Use embedding similarity and rerankers to select top-k (3 to 5) chunks, applying semantic deduplication before prompt construction.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Multilingual Translation Layer
&lt;/h3&gt;

&lt;p&gt;For bulk data extraction or batch processing on non-Latin languages, translating text to English with a lightweight model prior to deep inference on flagship models can reduce total token usage and improve execution latency.&lt;/p&gt;




&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+-------------------+---------------------------------------------------+
| Metric            | Engineering Reality                               |
+-------------------+---------------------------------------------------+
| Token Ratio (EN)  | ~1 Token ≈ 0.75 words (~4 characters)             |
| Token Ratio (TH)  | ~1 Word ≈ 3–8 Tokens (byte-level inflation)       |
| Cost Ratio        | Output is 3x–5x more expensive than Input         |
| Pricing Deltas    | Dense vs MoE, TPU/ASIC custom silicon, CoT tokens |
| Core Optimizers   | Caching + Routing + Windowing + RAG Filtering     |
+-------------------+---------------------------------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Treating tokens as finite compute bandwidth ensures your AI infrastructure remains fast, scalable, and economically sustainable.&lt;/p&gt;




&lt;p&gt;Explore the NEXT4I journey and read the original article at: &lt;a href="https://go.next4i.com/next4i-devto-en" rel="noopener noreferrer"&gt;https://go.next4i.com/next4i-devto-en&lt;/a&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>buildinpublic</category>
      <category>next4i</category>
    </item>
    <item>
      <title>The RAG War Story: How I Learned to Stop Worrying and Love Markdown</title>
      <dc:creator>NEXT4I DEV</dc:creator>
      <pubDate>Tue, 01 Sep 2026 11:23:36 +0000</pubDate>
      <link>https://dev.to/dev_next4i/the-rag-war-story-how-i-learned-to-stop-worrying-and-love-markdown-4f5p</link>
      <guid>https://dev.to/dev_next4i/the-rag-war-story-how-i-learned-to-stop-worrying-and-love-markdown-4f5p</guid>
      <description>&lt;p&gt;&lt;strong&gt;TLDR;&lt;/strong&gt; When building an AI knowledge retrieval pipeline that extracts text from documents, I discovered that PDF is the worst format for AI and Markdown is the best. Here's the multi-step pipeline I had to build just to handle PDFs, why it was necessary, and the generic pattern you can steal to handle document ingestion in your own RAG systems.&lt;/p&gt;




&lt;h3&gt;
  
  
  The Problem: PDFs Are Pixel-Perfect Hell for AI Parsers
&lt;/h3&gt;

&lt;p&gt;I was building a RAG (Retrieval-Augmented Generation) pipeline for NEXT4I the kind of system that reads your documents first, then answers questions from them. Standard stuff: document ingestion → chunking → embedding → vector search → LLM answer generation.&lt;/p&gt;

&lt;p&gt;I chose a beautiful Thai tourism PDF as my test document. Professional design, complex Thai typography, images, tables, charts the works. Real-world document, real-world pain.&lt;/p&gt;

&lt;p&gt;Here's what the naive approach looked like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PDF File → PDF Parser → Extracted Text → Chunk → Embed → Search
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And here's what actually worked:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PDF File
  ├─→ PDF Parser → Raw Text (broken Thai, missing punctuation)
  ├─→ Page Renderer → Full-Color Images
  │     └─→ B&amp;amp;W Converter → High-Contrast Images
  ├─→ AI Vision Model (color images) → Image Descriptions
  ├─→ AI Vision Model (B&amp;amp;W images) → Text Extraction
  └─→ Cross-Validation Layer
        ├─→ Multi-Model Synthesis
        ├─→ Spell-Check Model (critical for Thai)
        └─→ Human Review
              └─→ Final Structured Text → Chunk → Embed → Search
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why the complexity?&lt;/strong&gt; Because PDF is fundamentally a &lt;em&gt;presentation&lt;/em&gt; format, not a &lt;em&gt;data&lt;/em&gt; format. When you extract text from a PDF, you're not reading structured data you're reverse-engineering a rendered page layout. For languages with complex typography like Thai (where vowels can appear above, below, left, or right of consonants, and tone marks float above), this is especially brutal.&lt;/p&gt;




&lt;h3&gt;
  
  
  The Generic Pattern: Multi-Path Document Ingestion with Cross-Validation
&lt;/h3&gt;

&lt;p&gt;If you're building any system that ingests arbitrary documents, you'll inevitably hit the PDF wall. Here's the reusable pattern I settled on:&lt;/p&gt;

&lt;h4&gt;
  
  
  Architecture
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    ┌──────────────┐
                    │   Document   │
                    │   Ingest     │
                    └──────┬───────┘
                           │
              ┌────────────┼────────────┐
              ▼            ▼            ▼
        ┌──────────┐ ┌──────────┐ ┌──────────┐
        │  Direct  │ │  Image   │ │  Image   │
        │  Text    │ │  (Color) │ │  (B&amp;amp;W)   │
        │ Extract  │ │  Render  │ │  Render  │
        └────┬─────┘ └────┬─────┘ └────┬─────┘
             │            │            │
             ▼            ▼            ▼
        ┌──────────┐ ┌──────────┐ ┌──────────┐
        │  Text    │ │  Vision  │ │  Vision  │
        │  Output  │ │  Model   │ │  Model   │
        │          │ │  (Desc)  │ │  (OCR)   │
        └────┬─────┘ └────┬─────┘ └────┬─────┘
             │            │            │
             └────────────┼────────────┘
                          │
                          ▼
                 ┌─────────────────┐
                 │  Cross-Validate │
                 │  &amp;amp; Synthesize   │
                 │  (Multi-Model)  │
                 └────────┬────────┘
                          │
                          ▼
                 ┌─────────────────┐
                 │  Spell-Check    │
                 │  &amp;amp; Normalize    │
                 └────────┬────────┘
                          │
                          ▼
                 ┌─────────────────┐
                 │  Human Review   │
                 │  (Optional)     │
                 └────────┬────────┘
                          │
                          ▼
                 ┌─────────────────┐
                 │  Structured     │
                 │  Output → Embed │
                 └─────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key insight: &lt;strong&gt;no single extraction path is reliable enough on its own.&lt;/strong&gt; You need multiple independent paths producing results, then a synthesis layer that cross-validates. Think of it like sensor fusion each path is a noisy sensor, and the truth emerges from the overlap.&lt;/p&gt;

&lt;h4&gt;
  
  
  Why Spell-Check is Non-Negotiable for Non-English Languages
&lt;/h4&gt;

&lt;p&gt;For English, you might get away without a dedicated spell-check pass. For Thai where a single misplaced tone mark changes the entire word you absolutely cannot. OCR and vision models hallucinate characters constantly on decorated fonts or text-over-image backgrounds. A dedicated language model fine-tuned for spell correction is the difference between "usable" and "garbage."&lt;/p&gt;




&lt;h3&gt;
  
  
  The Real Takeaway: Markdown is AI-Native. Everything Else Is Legacy.
&lt;/h3&gt;

&lt;p&gt;After building this entire pipeline, I had a moment of clarity. If that same document had been authored in Markdown:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Top Destinations&lt;/span&gt;

| Province | Highlight | Best Season |
|----------|-----------|-------------|
| Krabi    | Islands   | Nov–Apr     |
| Chiang Mai | Mountains | Nov–Feb   |

See the &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;full itinerary&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;#itinerary&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; for details.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
mermaid&lt;br&gt;
graph TD&lt;br&gt;
    A[Arrive Bangkok] --&amp;gt; B[Fly to Krabi]&lt;br&gt;
    B --&amp;gt; C[Island Hopping]&lt;br&gt;
    C --&amp;gt; D[Return]&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
markdown&lt;/p&gt;

&lt;p&gt;...the entire pipeline collapses to: &lt;strong&gt;read the file → chunk → embed → search.&lt;/strong&gt; That's it.&lt;/p&gt;

&lt;p&gt;No OCR. No vision models. No B&amp;amp;W conversion. No multi-path cross-validation. No spell-check model. No human review for format-induced errors.&lt;/p&gt;

&lt;p&gt;Markdown is structured, plain-text, and both human-readable and machine-parseable by default. It's the only format where:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Headings&lt;/strong&gt; are unambiguously &lt;code&gt;#&lt;/code&gt; / &lt;code&gt;##&lt;/code&gt; not inferred from font size&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tables&lt;/strong&gt; are &lt;code&gt;| column | row |&lt;/code&gt; syntax not pixel grids&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Diagrams&lt;/strong&gt; are Mermaid text not flattened raster images&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code&lt;/strong&gt; is fenced not monospaced-font heuristics&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;"And here is how the human user experiences it:"&lt;/p&gt;
&lt;h2&gt;
  
  
  Top Destinations
&lt;/h2&gt;
&lt;/blockquote&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Province&lt;/th&gt;
&lt;th&gt;Highlight&lt;/th&gt;
&lt;th&gt;Best Season&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Krabi&lt;/td&gt;
&lt;td&gt;Islands&lt;/td&gt;
&lt;td&gt;Nov–Apr&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chiang Mai&lt;/td&gt;
&lt;td&gt;Mountains&lt;/td&gt;
&lt;td&gt;Nov–Feb&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;See the full itinerary for details.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;graph TD
    A[Arrive Bangkok] --&amp;gt; B[Fly to Krabi]
    B --&amp;gt; C[Island Hopping]
    C --&amp;gt; D[Return]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;&lt;em&gt;"In reality, we can't always control the documents we ingest, and we can't just ignore them because they might contain critical data. But if we were to start from scratch, Markdown is definitely the go-to choice."&lt;/em&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  How This Shapes Our Architecture at NEXT4I
&lt;/h3&gt;

&lt;p&gt;At NEXT4I, we treat Markdown as a first-class format throughout our stack. When building AI knowledge retrieval systems for everyday users and organizations, we encourage Markdown as the source of truth and handle PDFs as a necessary-but-painful compatibility layer.&lt;/p&gt;

&lt;p&gt;The design principle is simple: &lt;strong&gt;AI Integration by Design.&lt;/strong&gt; Make AI a first-class citizen of your content architecture, not something you bolt on later and hope it works. The format you choose today determines the ceiling of your AI capabilities tomorrow.&lt;/p&gt;




&lt;p&gt;Thanks for reading all the way to the end, I'll keep working on more articles like this.&lt;/p&gt;




&lt;p&gt;Explore the NEXT4I journey and read the original article at: &lt;a href="https://go.next4i.com/next4i-devto-en" rel="noopener noreferrer"&gt;https://go.next4i.com/next4i-devto-en&lt;/a&gt;&lt;/p&gt;

</description>
      <category>rag</category>
      <category>markdown</category>
      <category>pdf</category>
      <category>next4i</category>
    </item>
    <item>
      <title>What Is LLM Actually Doing? A Fellow Engineer's Take on Vectors, Next-Token Prediction, and Fail-back Routing</title>
      <dc:creator>NEXT4I DEV</dc:creator>
      <pubDate>Tue, 01 Sep 2026 05:30:51 +0000</pubDate>
      <link>https://dev.to/dev_next4i/what-is-llm-actually-doing-a-fellow-engineers-take-on-vectors-next-token-prediction-and-bb5</link>
      <guid>https://dev.to/dev_next4i/what-is-llm-actually-doing-a-fellow-engineers-take-on-vectors-next-token-prediction-and-bb5</guid>
      <description>&lt;p&gt;&lt;strong&gt;Why treating an LLM as a probability engine, not a brain, changes how you architect around it. The reasoning behind NEXT4I's AI layer.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;code&gt;#LLM&lt;/code&gt; &lt;code&gt;#BuildinPublic&lt;/code&gt; &lt;code&gt;#SystemArchitecture&lt;/code&gt; &lt;code&gt;#AI&lt;/code&gt; &lt;code&gt;#Model AI&lt;/code&gt; &lt;code&gt;#AI Router&lt;/code&gt; &lt;code&gt;#AI Stable&lt;/code&gt;&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;I used to wonder why the model confidently gives you a wrong number. It's not a bug in the traditional sense, it's the model doing exactly what it's built to do: predicting the next token from probability, with zero actual arithmetic happening underneath.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TLDR;&lt;/strong&gt; An LLM (large language model) converts text into vectors (numeric coordinates in a high-dimensional meaning-space) and generates output via next-token prediction, sampled with parameters like top-k and temperature. Because it's fundamentally a probability engine and not a calculator or a database, I designed NEXT4I with automatic model fail-back routing and task-based model selection instead of trusting any single model as a source of truth.&lt;/p&gt;




&lt;h3&gt;
  
  
  The Mental Model: Vectors, Not Meaning
&lt;/h3&gt;

&lt;p&gt;Every token gets embedded into a vector, often with hundreds or thousands of dimensions. Semantically similar tokens end up close together in that space. That's why semantic search (vector-based retrieval) can match "large flying animal consumes insects" to "big bird eats worms" even with zero shared keywords, unlike old-school lexical search (TF-IDF/BM25) which needs literal term overlap.&lt;/p&gt;

&lt;p&gt;GPUs handle this well because they're already wired for massive parallel floating-point math (the same math used to shade millions of pixels per frame), so throwing billions of similarly-directed vectors at a GPU is a natural fit, not a coincidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Generation Is Sampling, Not Retrieval
&lt;/h3&gt;

&lt;p&gt;Given a prompt, the model doesn't look up an answer, it samples one token at a time from a probability distribution. Two knobs matter in practice:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;top_k: 2        # only sample from the top-2 most likely next tokens
temperature: 0.2  # low = deterministic/precise, high = creative/varied
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Low temperature + low top_k gives you consistent, "boring" output, good for structured extraction. High temperature gives you variety, good for brainstorming, bad for anything requiring precision.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why LLMs Hallucinate on Arithmetic
&lt;/h3&gt;

&lt;p&gt;There's no calculator inside the model. It doesn't evaluate &lt;code&gt;x * y&lt;/code&gt;, it predicts digits that are statistically plausible given the prompt, one token at a time. It gets &lt;code&gt;2 * 2&lt;/code&gt; right because that pattern is everywhere in training data. It confidently botches large multiplication because it's still just sampling digits, not computing. This is exactly why production systems now delegate real math to a tool call (a Python sandbox, a calculator function) instead of trusting raw model output.&lt;/p&gt;




&lt;h3&gt;
  
  
  Core Value: A Generic Model-Tier Fail-back Router
&lt;/h3&gt;

&lt;p&gt;Here's the simple pattern, stripped of any specific business logic, a reusable fail-back wrapper for any set of same-tier model clients:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;package&lt;/span&gt; &lt;span class="n"&gt;modelrouter&lt;/span&gt;

&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s"&gt;"context"&lt;/span&gt;
    &lt;span class="s"&gt;"errors"&lt;/span&gt;
    &lt;span class="s"&gt;"fmt"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c"&gt;// ModelClient is any backend that can answer a prompt.&lt;/span&gt;
&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;ModelClient&lt;/span&gt; &lt;span class="k"&gt;interface&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Name&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;
    &lt;span class="n"&gt;Complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c"&gt;// TieredRouter tries each client in a tier in order until one succeeds.&lt;/span&gt;
&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;TieredRouter&lt;/span&gt; &lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;tier&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="n"&gt;ModelClient&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;NewTieredRouter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;clients&lt;/span&gt; &lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="n"&gt;ModelClient&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;TieredRouter&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;TieredRouter&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;tier&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;clients&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c"&gt;// Complete attempts each model in the tier, fail-back on error.&lt;/span&gt;
&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;TieredRouter&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;Complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="n"&gt;errs&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="kt"&gt;error&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="k"&gt;range&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tier&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="n"&gt;errs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;errs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fmt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Errorf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"%s: %w"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Name&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;errs&lt;/span&gt;&lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is intentionally boring: try the next model in the same tier on failure, return the first success. No retries with backoff yet, no circuit breaker, just the core fail-back idea. In NEXT4I's actual implementation, tiers are populated dynamically and health state feeds back into ordering, but that logic is abstracted here on purpose, the generic version above is what's actually useful to share.&lt;/p&gt;

&lt;h3&gt;
  
  
  Core Value: Task-Based Routing, the Simple Version
&lt;/h3&gt;

&lt;p&gt;A minimal router that inspects task complexity before picking a tier:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;package&lt;/span&gt; &lt;span class="n"&gt;modelrouter&lt;/span&gt;

&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;Complexity&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt;

&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;Simple&lt;/span&gt; &lt;span class="n"&gt;Complexity&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;iota&lt;/span&gt;
    &lt;span class="n"&gt;Complex&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;ClassifyAndRoute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;simpleTier&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;complexTier&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;TieredRouter&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;TieredRouter&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;estimateComplexity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;Simple&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;simpleTier&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;complexTier&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;estimateComplexity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;Complexity&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nb"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;Simple&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;Complex&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;estimateComplexity&lt;/code&gt; can be as crude as a length/keyword heuristic or as sophisticated as a small classifier model, the point is the routing &lt;em&gt;decision&lt;/em&gt; happens before the expensive call, not after.&lt;/p&gt;




&lt;h3&gt;
  
  
  Trade-offs I Made
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fail-back within a tier, not across tiers.&lt;/strong&gt; Swapping a cheap model in for an expensive one silently would change output quality without anyone noticing. Tiers exist specifically to avoid that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No cross-request state in the router.&lt;/strong&gt; Keeps it stateless and trivially horizontally scalable, at the cost of not learning from past failures within a single request lifecycle.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routing is a toggle, not a mandate.&lt;/strong&gt; Users can pin a specific model when they need deterministic behavior from one exact provider, the router only kicks in by default.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;If there's one thing worth taking away: treat the model as a probability engine you can't fully trust, and let the system around it, fail-back, routing, tool calls for math, carry the reliability burden instead.&lt;/p&gt;




&lt;p&gt;Thanks for reading all the way to the end, I'll keep working on more articles like this.&lt;/p&gt;




&lt;p&gt;Explore the NEXT4I journey and read the original article at: &lt;a href="https://go.next4i.com/next4i-devto-en" rel="noopener noreferrer"&gt;https://go.next4i.com/next4i-devto-en&lt;/a&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>buildinpublic</category>
      <category>next4i</category>
      <category>architecture</category>
    </item>
    <item>
      <title>How to Build a "Second Brain" with Obsidian That Your AI Agent Can Read, Without Building a Custom RAG Pipeline</title>
      <dc:creator>NEXT4I DEV</dc:creator>
      <pubDate>Mon, 17 Aug 2026 11:40:46 +0000</pubDate>
      <link>https://dev.to/dev_next4i/how-to-build-a-second-brain-with-obsidian-that-your-ai-agent-can-read-without-building-a-custom-29h3</link>
      <guid>https://dev.to/dev_next4i/how-to-build-a-second-brain-with-obsidian-that-your-ai-agent-can-read-without-building-a-custom-29h3</guid>
      <description>&lt;p&gt;Ever run into this? You wrote a detailed technical spec three months ago, and today someone asks "how did we design this module again?" You end up spending 20 minutes hitting &lt;code&gt;Cmd+F&lt;/code&gt; across Google Docs, Trello, and &lt;code&gt;README.md&lt;/code&gt; files scattered across different repos, or sometimes there's no documentation at all, so you have to dig through the code and reverse-engineer it by eye.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TLDR;&lt;/strong&gt; I built a knowledge-base system for NEXT4I using Obsidian (plain Markdown) together with Git and a VS Code AI agent, without writing a single line of custom integration. The core of the system is choosing tools that already "speak the same language" from day one (plain text, open formats), so AI can read our knowledge base practically for free, no RAG pipeline required.&lt;/p&gt;




&lt;h3&gt;
  
  
  Pain Point: Scattered Knowledge and Requirements Make Search Hard and Nothing Stays in Sync
&lt;/h3&gt;

&lt;p&gt;Before this system, all of NEXT4I's knowledge was scattered across many places: notebooks, Apple Notes, Google Docs, Google Sheets, Trello, &lt;code&gt;README.md&lt;/code&gt; files, or, even worse, sometimes there was no documentation at all, just buried in the code, spread across multiple repos each written in a different language, frontend and backend alike.&lt;/p&gt;

&lt;p&gt;This isn't just an inconvenience. It makes search genuinely hard, sometimes it's &lt;code&gt;Cmd+F&lt;/code&gt; and pray. Worse: my AI coding agent could read the entire codebase, but &lt;strong&gt;it had zero visibility into the reasoning behind that code&lt;/strong&gt;, because that reasoning was scattered somewhere the AI couldn't reach.&lt;/p&gt;




&lt;h3&gt;
  
  
  Design Constraints: 3 Non-Negotiables
&lt;/h3&gt;

&lt;p&gt;Before picking a tool, I set 3 rules:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It has to be affordable, or free if possible.&lt;/li&gt;
&lt;li&gt;It has to be accessible online anytime, from my phone, and still work offline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plain text, zero lock-in.&lt;/strong&gt; If the tool disappears tomorrow, the files must remain immediately readable and usable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local-first, Git-friendly.&lt;/strong&gt; It has to be a normal folder I can &lt;code&gt;git init&lt;/code&gt; and track right away.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI-readable without building extra infrastructure.&lt;/strong&gt; My AI agent already lives in VS Code, so the knowledge base has to sit inside that workspace without me building an extra pipeline.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Obsidian passed everything, because at its core, an Obsidian vault is just a folder of &lt;code&gt;.md&lt;/code&gt; files.&lt;/p&gt;




&lt;h3&gt;
  
  
  Architecture: Vault Structure
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;vault/
├── Ideas/          # Raw concepts, brainstorming
├── Manifesto/      # Vision, mission, core policies
├── Principles/     # Design rules, engineering guidelines
├── Infrastructure/ # Deployment topology, IaC specs
├── Platform/       # Domain model, API contracts
├── Script/         # Utility scripts, automation, runbooks
├── Skill/          # Patterns, checklists, reusable knowledge
└── Appendix/       # Domain glossary, citations
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every file is a plain &lt;code&gt;.md&lt;/code&gt;. Links use &lt;code&gt;[[wiki-link]]&lt;/code&gt; syntax. Metadata lives in YAML frontmatter, and Graph View renders the relationships as a visible dependency graph.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Version control&lt;/strong&gt; is just &lt;code&gt;git init&lt;/code&gt; inside the vault folder, then committing every change with a rationale. &lt;code&gt;git log -- "Infrastructure/sharding-strategy.md"&lt;/code&gt; shows the full decision history for that topic. Push it to a private GitHub repo and you get backup, an audit trail, and branching for major revisions, all for free.&lt;/p&gt;




&lt;h3&gt;
  
  
  Key Insight: AI Access Without Building Anything
&lt;/h3&gt;

&lt;p&gt;This is the part that changed everything.&lt;/p&gt;

&lt;p&gt;Obsidian vault = a folder of &lt;code&gt;.md&lt;/code&gt; files.&lt;br&gt;
VS Code = opens any folder as a workspace.&lt;br&gt;
AI coding agent (running as a VS Code extension) = reads every file in that workspace.&lt;/p&gt;

&lt;p&gt;So the "integration" here is: &lt;strong&gt;open the vault folder in VS Code.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's it. No API, no embedding pipeline, no vector database, no chunking strategy. Just plain Markdown files that the AI agent reads natively.&lt;/p&gt;


&lt;h3&gt;
  
  
  Prompt Patterns I Actually Use
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Contextual search + reasoning:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Search all documents that mention our sharding strategy.
Summarize every trade-off we've considered
and tell me which approach we ultimately chose and why.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Gap analysis:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Look at everything in the Infrastructure/ folder
and tell me which architectural decisions are still undocumented,
compared against the template in Skill/.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Drafting from conventions, not from a blank page:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using the patterns in Skill/go-backend/ and the domain model in Platform/core/,
draft a design doc for a new message consumer
following our established conventions.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Impact analysis via link traversal:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;If I change the authentication rule in Principles/auth.md,
trace every file in Platform/ and Skill/ that links to it via [[links]]
and tell me what needs updating.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI reads across multiple files, follows &lt;code&gt;[[wiki-links]]&lt;/code&gt;, understands the relationships, and synthesizes an answer, without me writing a single line of integration code.&lt;/p&gt;




&lt;h3&gt;
  
  
  Philosophy: Seamless Integration by Design
&lt;/h3&gt;

&lt;p&gt;The pattern here isn't "integrate 3 tools." It's &lt;strong&gt;choosing components that already speak the same language.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Obsidian chose Markdown, the most universally readable format in computing, over a proprietary database. Git works with any plain text. AI agents already know how to read files in a VS Code workspace.&lt;/p&gt;




&lt;h3&gt;
  
  
  Bonus: Obsidian Canvas as a Visual Layer
&lt;/h3&gt;

&lt;p&gt;Obsidian's Canvas is an infinite whiteboard, place document cards, text, media, then draw connections between them.&lt;/p&gt;

&lt;p&gt;I use Canvas for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;System architecture sketches:&lt;/strong&gt; each service as a card, data flow arrows, real specs embedded right on the board&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decision trees:&lt;/strong&gt; "if we pick X, then Y and Z are affected," with linked evidence&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan strategy &amp;amp; flow:&lt;/strong&gt; for planning work, sequencing, and various NEXT4I workflows&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Canvas files are Markdown too under the hood (JSON-like structure), so they're Git-versioned and AI-readable as well. Drop a &lt;code&gt;.canvas&lt;/code&gt; file into your VS Code workspace and ask the AI to analyze it for circular dependencies or single points of failure.&lt;/p&gt;




&lt;h3&gt;
  
  
  What I Learned
&lt;/h3&gt;

&lt;p&gt;What makes this system work isn't the technology, it's what I &lt;em&gt;didn't&lt;/em&gt; build. No middleware, no proprietary pipeline, no vendor lock-in.&lt;/p&gt;

&lt;p&gt;The discipline of plain text + Git + open formats is a feature, not a limitation.&lt;/p&gt;

&lt;p&gt;If you're a dev or a small team drowning in scattered documents, try this before jumping to a heavyweight knowledge-management platform. Plain Markdown with a good folder structure will take you further than you'd expect.&lt;/p&gt;




&lt;p&gt;Thanks for reading all the way to the end, I'll keep working on more articles like this.&lt;/p&gt;

&lt;p&gt;Explore the NEXT4I journey and read the original article at: &lt;a href="https://go.next4i.com/next4i-devto-en" rel="noopener noreferrer"&gt;https://go.next4i.com/next4i-devto-en&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Alternatively, you can register to join &lt;strong&gt;NEXT4I&lt;/strong&gt; the AI-Native Ecosystem I am currently building at: &lt;a href="https://go.next4i.com/cta" rel="noopener noreferrer"&gt;https://go.next4i.com/cta&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;tags: &lt;br&gt;
&lt;code&gt;#buildinpublic&lt;/code&gt;&lt;br&gt;
&lt;code&gt;#secondbrain&lt;/code&gt;&lt;br&gt;
&lt;code&gt;#next4i&lt;/code&gt;&lt;br&gt;
&lt;code&gt;#obsidian&lt;/code&gt;&lt;/p&gt;




</description>
      <category>ai</category>
      <category>markdown</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How I Built an AI-Readable Second Brain with Obsidian, Git, and a VS Code AI Agent</title>
      <dc:creator>NEXT4I DEV</dc:creator>
      <pubDate>Fri, 07 Aug 2026 07:55:45 +0000</pubDate>
      <link>https://dev.to/dev_next4i/how-i-built-an-ai-readable-second-brain-with-obsidian-git-and-a-vs-code-ai-agent-5hep</link>
      <guid>https://dev.to/dev_next4i/how-i-built-an-ai-readable-second-brain-with-obsidian-git-and-a-vs-code-ai-agent-5hep</guid>
      <description>&lt;p&gt;When you're a solo founder and lead architect, your knowledge base is your most valuable asset. Lose the thread on why a decision was made, and you spend hours — sometimes days — reconstructing context that you &lt;em&gt;already figured out once&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;I want to share the exact setup I use at NEXT4I to turn a folder of Markdown files into a fully searchable, version-controlled, AI-readable knowledge system. No proprietary SaaS, no vendor lock-in, no custom integration work.&lt;/p&gt;

&lt;p&gt;This is the Key Highlight of this post: a genuinely useful, generic pattern you can apply to your own projects today. The NEXT4I-specific business logic stays abstracted (per our security rules), but the pattern itself is 100% reusable.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. The Architecture: Three Layers, Zero Magic
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌─────────────────────────────────────────┐
│           AI Agent (VS Code)            │
│   Reads, Searches, Summarizes, Drafts   │
└──────────────────┬──────────────────────┘
                   │ reads plain .md files
┌──────────────────▼──────────────────────┐
│       Git-tracked Obsidian Vault        │
│  ├── Idea/           (brainstorms)      │
│  ├── Infrastructure/ (architecture docs)│
│  ├── Platform/       (product specs)    │
│  ├── Script/         (automation)       │
│  └── Skill/          (reusable limits)  │
└──────────────────┬──────────────────────┘
                   │ committed &amp;amp; pushed
┌──────────────────▼──────────────────────┐
│         GitHub (remote backup)          │
└─────────────────────────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Layer 1 — The Vault (Obsidian):&lt;/strong&gt; A folder of interconnected &lt;code&gt;.md&lt;/code&gt; files. The key insight is that Obsidian uses plain Markdown with &lt;code&gt;[[wiki-links]]&lt;/code&gt; for connections — no database, no proprietary format.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 2 — Version Control (Git):&lt;/strong&gt; Every vault is a git repo. Every change to any document has a commit message, a timestamp, and a diff. You can &lt;code&gt;git log --oneline -- Idea/&lt;/code&gt; to see the evolution of a concept.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 3 — AI Agent (VS Code Extension):&lt;/strong&gt; Because the vault is just a file tree of &lt;code&gt;.md&lt;/code&gt; files, any AI coding agent that can read a codebase can also read your knowledge base. Point the agent at the vault folder, and it has full context.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. The Setup: Step-by-Step
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Create the Vault
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;
&lt;span class="nb"&gt;mkdir &lt;/span&gt;next4i-knowledge

&lt;span class="nb"&gt;cd &lt;/span&gt;next4i-knowledge

&lt;span class="nb"&gt;mkdir &lt;/span&gt;Idea Infrastructure Platform Script Skill

git init
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open this folder in Obsidian: &lt;strong&gt;Open folder as vault&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Link Everything
&lt;/h3&gt;

&lt;p&gt;Inside a note, link to another note with &lt;code&gt;[[Note Name]]&lt;/code&gt;. Obsidian auto-suggests as you type. Over time, this builds a graph you can visualize with &lt;code&gt;Cmd/Ctrl + G&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pro tip:&lt;/strong&gt; Create a &lt;code&gt;_INDEX.md&lt;/code&gt; in each folder that links to the most important notes. This becomes a human-readable table of contents AND a search anchor for the AI.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Add Git Discipline
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;
git add &lt;span class="nt"&gt;-A&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"infra: initial sharding strategy decision"&lt;/span&gt;

git remote add origin git@github.com:your-org/knowledge-vault.git

git push &lt;span class="nt"&gt;-u&lt;/span&gt; origin main

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Treat commit messages like code. Use prefixes: &lt;code&gt;idea:&lt;/code&gt;, &lt;code&gt;infra:&lt;/code&gt;, &lt;code&gt;platform:&lt;/code&gt;, &lt;code&gt;script:&lt;/code&gt;, &lt;code&gt;skill:&lt;/code&gt;. This makes &lt;code&gt;git log --oneline --grep="infra:"&lt;/code&gt; instantly useful.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Open in VS Code and Activate the AI
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;
code /path/to/vault

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With an AI agent extension active (Copilot, Cline, Cody, etc.), try prompts like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;"Summarize the key architectural decisions in the Infrastructure folder."&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;"Find any contradiction between documents in /Platform/ and /Infrastructure/."&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;"Draft a new document in /Idea/ based on the sharding notes in /Infrastructure/."&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent reads the files as context, just like it would for code.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. The Design Pattern: Folder-Convention-as-API
&lt;/h2&gt;

&lt;p&gt;Here's the key pattern: &lt;strong&gt;your folder structure IS your API&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;By keeping a consistent vault structure, both humans and AI know where to look:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Folder&lt;/th&gt;
&lt;th&gt;Contains&lt;/th&gt;
&lt;th&gt;AI Use Case&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Idea/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Raw, unstructured thinking&lt;/td&gt;
&lt;td&gt;Generate summaries, find related concepts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Infrastructure/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;System topology, deployment, config&lt;/td&gt;
&lt;td&gt;Validate consistency, trace dependencies&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Platform/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Feature specs, user flows&lt;/td&gt;
&lt;td&gt;Draft task tickets, check requirement coverage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Script/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Automation, one-liners&lt;/td&gt;
&lt;td&gt;Explain what a script does, suggest improvements&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Skill/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Reusable patterns, checklists&lt;/td&gt;
&lt;td&gt;Retrieve relevant patterns for new tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is essentially a &lt;strong&gt;convention-based RAG (Retrieval-Augmented Generation)&lt;/strong&gt; setup without any vector database, embedding pipeline, or chunking strategy. The "chunking" is the natural boundary of each &lt;code&gt;.md&lt;/code&gt; file. The "retrieval" is the AI agent's file-reading capability.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Why This Beats a Wiki
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Wiki / Confluence&lt;/th&gt;
&lt;th&gt;This Setup (Obsidian + Git)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Vendor lock-in&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Plain &lt;code&gt;.md&lt;/code&gt; files, portable anywhere&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Search is siloed&lt;/strong&gt; within the tool&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;VS Code AI&lt;/strong&gt; searches across the whole vault&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;No version control&lt;/strong&gt; (or poor built-in)&lt;/td&gt;
&lt;td&gt;Full version control (&lt;code&gt;git blame&lt;/code&gt;, &lt;code&gt;git diff&lt;/code&gt;, &lt;code&gt;git log&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Hard to automate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Scriptable&lt;/strong&gt; — &lt;code&gt;grep&lt;/code&gt;, &lt;code&gt;sed&lt;/code&gt;, and AI prompts all work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;AI needs API integration&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;AI reads files natively&lt;/strong&gt;, zero setup required&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  5. What I Learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The biggest surprise:&lt;/strong&gt; The AI agent became better at finding connections in my own notes than I was. It doesn't have recency bias. It doesn't forget what I wrote 8 months ago. It reads everything with equal attention.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The biggest lesson:&lt;/strong&gt; AI-native doesn't mean "add an AI button." It means design your systems — including your thinking systems — so that AI can participate as a first-class citizen without special plumbing.&lt;/p&gt;




&lt;p&gt;I'm building NEXT4I as an AI-native ecosystem from the ground up. If you're interested in following a solo founder's engineering journey — or want early access — join here: &lt;br&gt;
&lt;a href="https://go.next4i.com/cta" rel="noopener noreferrer"&gt;Subscribe NEXT4I or want early access — join here&lt;/a&gt;&lt;/p&gt;

</description>
      <category>buildinpublic</category>
      <category>obsidian</category>
      <category>next4i</category>
      <category>secondbrain</category>
    </item>
  </channel>
</rss>
