At 3:15 AM, an on-call engineer paged for an upstream latency cliff does not need an LLM to rephrase their incident ticket. They do not need a four-paragraph preamble detailing the history of distributed consensus, nor do they need the model to burn 600 output tokens restating the stack trace they just pasted. They need the remediation command on line one.
Most production retrieval-augmented generation (RAG) pipelines fail here because teams optimize in reverse: they fine-tune embedding spaces, push chunk retrieval to top-k=20, and wonder why their token bills surge while mean time to resolution (MTTR) stays flat. The model is not lacking context—it is burying the operational payload under conversational debt.
That is the practical reason I integrated ayghri/i-have-adhd into our engineering assistant pipeline. The repository defines an explicit, zero-fluff output contract designed to stop coding and operations agents from hiding the answer: lead with the next action, number multi-step tasks, cap lists at five items, and terminate on a single actionable step. It is not an embedder, vector store, or caching layer; treating it strictly as a deterministic response-policy layer keeps our system architecture clean.
The Pipeline: Retrieval First, Policy Second
A production RAG pipeline separates stateful retrieval from generation governance across four stages:
- Query Normalization: Strip conversational noise and extract exact technical anchors—file paths, error codes, and flags. Do not pipe raw message threads into embedding endpoints.
- Retrieval and Ingestion Filtering: Pull candidate passages, deduplicate exact token spans, and enforce tenant-level authorization before documents hit the context window. Data isolation belongs here, not inside prompt instructions.
- Deterministic Prompt Assembly: Position immutable system instructions and policy contracts ahead of variable context chunks. Strict ordering is mandatory for upstream gateways that support prefix caching.
- Generation & Boundary Validation: Enforce structural output rules: require an immediate command on line one and reject unbounded verbosity. A policy layer enforces formatting discipline, but it cannot fix an upstream retrieval defect or an unauthenticated document leak.
The contract in ayghri/i-have-adhd provides a sharp boundary for stage four: instead of conversational pleasantries, the output starts with executable instructions, uses numbered sequences, and closes with a single verified next step [1].
Implementation: Deterministic Assembly and Prefix Alignment
Here is a minimal, production-tested Python harness demonstrating chunk deduplication, prefix cache preservation, and policy-first assembly:
from hashlib import sha256
from dataclasses import dataclass
@dataclass(frozen=True)
class Chunk:
source: str
text: str
def chunk_document(
source: str,
text: str,
size: int = 1200,
overlap: int = 120,
) -> list[Chunk]:
if size <= overlap or not text.strip():
raise ValueError("size must be greater than overlap")
step = size - overlap
return [
Chunk(source, text[i:i + size].strip())
for i in range(0, len(text), step)
if text[i:i + size].strip()
]
def dedupe(chunks: list[Chunk]) -> list[Chunk]:
seen: set[str] = set()
result = []
for chunk in chunks:
key = sha256(chunk.text.encode()).hexdigest()
if key not in seen:
seen.add(key)
result.append(chunk)
return result
def assemble(
policy: str,
question: str,
hits: list[Chunk],
limit: int = 6,
) -> str:
context = "\n\n".join(
f"[{i}] {hit.source}\n{hit.text}"
for i, hit in enumerate(dedupe(hits)[:limit], 1)
)
return f"""{policy}
<retrieved_context>
{context}
</retrieved_context>
User request: {question}
Return the answer directly. Do not repeat the request."""
POLICY = """You are an engineering assistant.
Lead with the next action. Number multi-step tasks.
Keep lists to five items or fewer. End with one concrete next step.
State uncertainty plainly. Never invent repository facts."""
Keep the POLICY block byte-for-byte identical across calls. Placing dynamic timestamps, session IDs, or ephemeral run tokens at the prompt header destroys KV-cache reuse upstream, transforming what should have been a 70% cache hit rate into a 100% full-billing token penalty.
Failure Modes Under Real Load
Naive RAG implementations simply concatenate search outputs and call completion endpoints. In production, this breaks across three failure modes:
-
Redundant Token Inflation: Overlapping chunks frequently repeat identical sentences. In a
top-k=8retrieval, duplicate tokens can account for 25% of your context window. Deterministic exact deduplication is practically free in CPU cycles and prevents paying twice for identical strings. -
Context Injection Bleed: Retrieved documentation may contain embedded directives or sample code formatted like prompt instructions. Encapsulating raw hits within explicit
<retrieved_context>XML tags prevents the model from conflating data with policy, though it is never a substitute for network-layer sanitization. - Template Drift & Rejection Loops: If your output validator rejects responses that deviate by a single character, generation retries will spike p99 latency. Validate only structural boundaries—verifying an actionable first line and list length caps—rather than policing exact wording.
The operational economics are straightforward. A 900-token system and policy prefix queried 10,000 times represents 9,000,000 input tokens. When served through a gateway supporting native prompt caching (such as B-Lost), caching that prefix dramatically cuts both time-to-first-token (TTFT) and inference overhead.
The Operational Boundary
Integrate i-have-adhd to control output shape: eliminate conversational filler, surface runbooks immediately, enforce strict item limits, and preserve operational focus. Do not ask it to solve vector drift, authorization boundaries, or ground-truth hallucination.
The central dilemma every systems engineer faces when tightening RAG output contracts is balancing brevity against actionable edge cases. If you clamp output tokens too aggressively, the model truncates critical rollback flags; if you loosen constraints, on-call operators drown in conversational fluff during an outage.
How does your team enforce output contracts across your RAG and agent gateways? Are you validating structured output through JSON Schema/Wasm layers, or relying on prompt-level response policies and KV-prefix alignment? Drop your architecture and battle scars in the comments below.
Disclosure: Compute infrastructure and multi-model benchmark relays for this writeup are sponsored by b-lost.com — an enterprise AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All benchmark metrics reflect independent reproducible testing.
Top comments (0)