The moment I stopped trusting my internal agent was when it gave me the wrong rollback steps with total confidence.
Not obviously wrong.
Worse: plausibly wrong.
I asked for the rollback procedure for one service. It retrieved a semantically similar chunk from a different runbook. Same kind of deployment. Same general wording. Wrong ENV vars.
That was the point where I realized my RAG stack was solving the fun problem, not the real one.
For a lot of internal bots, the hard part is not embeddings. It’s keeping docs clean, structured, and current.
And if your docs are already decent, markdown-first retrieval is often a better default than jumping straight to a vector DB.
The docs most internal agents actually search
A lot of internal agents are not doing open-ended research.
They’re answering questions from:
- incident runbooks
- SOPs
- onboarding docs
- deployment checklists
- feature flag docs
- integration notes
-
ENVvar references
That matters because these docs are full of exact strings that carry most of the meaning.
Things like:
STRIPE_WEBHOOK_SECRET--force-with-lease- error code
E1027 - table name
customer_billing_events - heading
Rollback procedure - heading
Required ENV vars
This is where vector-only retrieval gets weird.
If I’m asking about PAYMENTS_DB_URL, I do not want something semantically close. I want the section that literally contains PAYMENTS_DB_URL.
For ops bots, “close enough” is how you get bad answers that look good.
Your markdown already has retrieval boundaries
This is the part I think people underestimate.
When engineers write markdown, they already chunk the document for you.
Example:
# Payments Service
## Rollback procedure
1. Disable the `payments-v2` feature flag
2. Re-deploy previous image tag
3. Verify `PAYMENTS_DB_URL`
## Required ENV vars
- PAYMENTS_DB_URL
- STRIPE_WEBHOOK_SECRET
That structure is useful.
- The file path tells you the system
- The heading tells you the task
- The list tells you the exact steps
- The code spans tell you the exact identifiers
Then a lot of RAG pipelines throw that away and split every 500–1000 characters.
That’s a strange move when the human author already did the chunking.
LangChain and LlamaIndex both hint at this
This is not some anti-RAG hot take.
The frameworks themselves already expose structure-aware parsing.
LangChain example
from langchain_text_splitters import MarkdownHeaderTextSplitter
headers_to_split_on = [
("#", "Header 1"),
("##", "Header 2"),
("###", "Header 3"),
]
markdown_splitter = MarkdownHeaderTextSplitter(headers_to_split_on)
docs = markdown_splitter.split_text(markdown_text)
for doc in docs:
print(doc.metadata)
print(doc.page_content)
That exists for a reason.
A practical retrieval shape
Instead of flattening docs into generic chunks, preserve:
- file path
- heading hierarchy
- section body
- exact tokens like flags, table names, and env vars
Then retrieve by:
- exact term match
- heading match
- file path or service name
- neighboring section expansion if needed
That gets you surprisingly far.
The hidden cost of vector search for internal docs
Vector search is useful. I’m not anti-vector.
But for small internal agents, it adds a second system you now have to maintain.
Usually that means wiring together:
- document loading
- chunking
- embeddings
- vector storage
- retrieval logic
- re-indexing when docs change
That’s a lot of machinery just to read your own markdown.
And every layer can drift.
Somebody updates docs/deployments/payments.md in Git.
Now you need to know:
- did the chunk boundaries change?
- did the embedding job rerun?
- did the vector index update?
- are stale chunks still retrievable?
- did retrieval ranking change?
When your source of truth is Git, adding a second truth system is where things get brittle.
A boring markdown-first pipeline that works
For a bot over SOPs and runbooks, this is often enough:
- keep docs as
.mdin Git - parse headings into metadata
- build a lexical index over sections
- retrieve by exact term + heading + file path
- send the best section to the model
- expand to adjacent sections only if the answer is incomplete
That’s it.
No embedding jobs.
No Pinecone, Weaviate, Qdrant, or pgvector on day one.
No “why is the bot citing last week’s docs?” debugging session.
Example: indexing markdown sections in Python
Here’s a tiny example of what I mean.
import re
from pathlib import Path
from dataclasses import dataclass
@dataclass
class Section:
path: str
h1: str | None
h2: str | None
content: str
def parse_markdown_sections(path: Path) -> list[Section]:
text = path.read_text()
lines = text.splitlines()
sections = []
current_h1 = None
current_h2 = None
buffer = []
def flush():
if buffer and (current_h1 or current_h2):
sections.append(
Section(
path=str(path),
h1=current_h1,
h2=current_h2,
content="\n".join(buffer).strip(),
)
)
for line in lines:
if line.startswith("# "):
flush()
current_h1 = line[2:].strip()
current_h2 = None
buffer = []
elif line.startswith("## "):
flush()
current_h2 = line[3:].strip()
buffer = []
else:
buffer.append(line)
flush()
return sections
def score(query: str, section: Section) -> int:
haystack = " ".join([
section.path,
section.h1 or "",
section.h2 or "",
section.content,
]).lower()
terms = re.findall(r"[a-zA-Z0-9_\-\.]+", query.lower())
return sum(1 for term in terms if term in haystack)
def search(query: str, docs_dir: str) -> list[Section]:
sections = []
for path in Path(docs_dir).rglob("*.md"):
sections.extend(parse_markdown_sections(path))
ranked = sorted(sections, key=lambda s: score(query, s), reverse=True)
return [s for s in ranked if score(query, s) > 0][:5]
Query it like this:
results = search("payments rollback PAYMENTS_DB_URL", "./docs")
for r in results:
print("---")
print(r.path)
print(r.h1, "->", r.h2)
print(r.content)
Is this fancy? No.
Does it work well for internal docs with disciplined naming? Very often, yes.
Why this often improves answer quality
The obvious win is simpler retrieval.
The less obvious win is prompt quality.
A lot of bad agent behavior starts earlier than people think. The model gets a pile of vaguely relevant chunks and has to guess which one matters.
Markdown-first retrieval tends to send:
- fewer sections
- smaller context
- cleaner boundaries
- more exact identifiers
That means less prompt stuffing and less context-window waste.
And if you’re running automations all day in n8n, Make, Zapier, OpenClaw, or your own Python worker, retrieval mistakes get expensive fast.
More bad retrieval means:
- more retries
- bigger prompts
- more fallback calls
- more guardrail logic
That’s exactly the kind of thing that creates token anxiety when you’re billed per token.
This is one reason flat-rate inference is so useful for agent workflows. When you’re iterating on retrieval, prompts, and multi-step automations, per-token pricing punishes experimentation. Standard Compute is basically built for this kind of workload: OpenAI-compatible API, flat monthly pricing, and enough headroom to run internal agents without watching every request like a taxi meter.
Markdown-first vs vector RAG in practice
Here’s the tradeoff as I’ve seen it.
| Approach | What actually happens in practice |
|---|---|
| Markdown knowledge base + lexical/section retrieval | Low complexity, strong on exact terms like ENV vars, flags, error codes, and product names, easy to maintain in Git |
| Pure vector RAG stack | More setup and more moving parts, better for fuzzy search across messy corpora, easier to get semantically similar but operationally wrong matches |
| Hybrid retrieval | Usually the best long-term option once simple retrieval stops being enough, but still more operational overhead than markdown-first |
That hybrid row matters.
A lot of teams end up there anyway.
Which is basically the industry admitting that pure semantic similarity is not enough for documentation-heavy workloads.
When markdown-first wins
Markdown-first usually wins when:
- your docs are clean
- headings are meaningful
- service names are consistent
- exact strings matter
- the answer usually lives in one section
This is common for:
- support engineering bots
- internal ops assistants
- deployment helpers
- onboarding assistants
- docs Q&A bots for one team
If the right answer is under ## Rollback procedure, retrieving that exact section is usually better than retrieving some semantically related paragraph from another doc.
When vector search absolutely helps
There are plenty of cases where embeddings earn their keep.
Use vector or hybrid retrieval when:
- docs are inconsistent across teams
- people use different names for the same thing
- answers are spread across multiple documents
- users ask fuzzy natural-language questions
- your corpus is large and messy
This is not a religion.
It’s a sequencing argument.
The rollout path I’d recommend
If I were building an internal docs bot today, I’d do it in this order:
- Start with markdown files in Git
- Enforce strong headings
- Index by file path + heading + exact terms
- Retrieve the smallest useful section
- Add neighboring-section expansion
- Add BM25 or another lexical ranker
- Add hybrid retrieval if fuzzy queries become a real problem
- Only then consider a vector-heavy architecture
Most teams start at step 8 because that’s what demos show.
But demos are optimized to look smart.
Internal agents are optimized to not break on Tuesday.
Quick test before you spin up a vector DB
Before you reach for Pinecone, Weaviate, Qdrant, or pgvector, try this:
# 1. Put your docs in one place
mkdir -p docs/runbooks docs/sops docs/integrations
# 2. Standardize headings
# Example sections:
# - Overview
# - Rollback procedure
# - Required ENV vars
# - Common errors
# 3. Grep for exact terms first
rg "PAYMENTS_DB_URL|Rollback procedure|E1027" docs/
If a fast exact-match pass already finds the right section, that’s a strong sign you should start simple.
My actual takeaway
The weirdly effective thing I wish more people tried first is this:
- write clean markdown
- keep one procedure per section
- preserve heading metadata
- search exact terms first
- pass the smallest correct section to the model
A lot of prompt engineering is really retrieval regret.
We tell GPT-5, Claude, Qwen, or Llama to “be careful” because we gave them a pile of mush and hoped they’d sort it out.
A clean markdown knowledge base does the opposite.
It gives the model a sharply bounded slice of truth.
And for small internal agents, that is often the whole game.
Top comments (1)
The PAYMENTS_DB_URL example is exactly where vector-only retrieval bites on support and ops bots.
When the answer hangs on an env var name, a feature flag, or a heading like Rollback procedure, semantically close is the failure mode. You want the section that literally contains the string. Heading-aware markdown chunks already give you those boundaries. Throwing them away for fixed-size embeddings is how you get a confident wrong rollback.
I would keep lexical match on path, heading, and exact tokens as the first pass for internal runbooks, and only add vectors when the corpus is large enough that synonyms matter more than identifiers. For small SOP sets, the second truth system (embeddings, reindex lag, stale chunks) often costs more than it returns.
The other win you name is prompt quality: fewer, cleaner sections beat a pile of near-miss chunks. The model guesses less when retrieval already did the hard work.