<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Anand Kannan</title>
    <description>The latest articles on DEV Community by Anand Kannan (@anandindia93).</description>
    <link>https://dev.to/anandindia93</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F667776%2F3a968b03-ac1b-4aed-9a9a-e5979385c591.png</url>
      <title>DEV Community: Anand Kannan</title>
      <link>https://dev.to/anandindia93</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/anandindia93"/>
    <language>en</language>
    <item>
      <title>RAG vs OKF: What’s the Difference and Where Should SREs Use Them?</title>
      <dc:creator>Anand Kannan</dc:creator>
      <pubDate>Wed, 07 Oct 2026 17:52:17 +0000</pubDate>
      <link>https://dev.to/anandindia93/rag-vs-okf-whats-the-difference-and-where-should-sres-use-them-4b6e</link>
      <guid>https://dev.to/anandindia93/rag-vs-okf-whats-the-difference-and-where-should-sres-use-them-4b6e</guid>
      <description>&lt;h1&gt;
  
  
  RAG vs OKF: What’s the Difference and Where Should SREs Use Them?
&lt;/h1&gt;

&lt;p&gt;If you're working in SRE, you've probably heard a lot about &lt;strong&gt;RAG (Retrieval-Augmented Generation)&lt;/strong&gt; and &lt;strong&gt;OKF (Open Knowledge Framework)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;At first, they can sound like the same thing:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Give an AI access to our documentation and let it answer questions."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But they're actually solving different problems.&lt;/p&gt;

&lt;p&gt;Let's understand this without the AI buzzwords.&lt;/p&gt;




&lt;h2&gt;
  
  
  The 10-second explanation
&lt;/h2&gt;

&lt;p&gt;Think of an AI system as an SRE engineer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RAG says:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Before answering, search our runbooks, documentation and incident reports."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;OKF says:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Don't just give me documents. Organize knowledge, relationships and context so an AI can navigate it consistently."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RAG = Retrieve relevant information&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OKF = Organize knowledge and relationships for AI consumption&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;They aren't competitors. In a good SRE system, you can use both.&lt;/p&gt;




&lt;h1&gt;
  
  
  1. What is RAG?
&lt;/h1&gt;

&lt;p&gt;RAG stands for:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retrieval-Augmented Generation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The idea is surprisingly simple.&lt;/p&gt;

&lt;p&gt;Instead of asking an LLM a question and expecting it to know everything, we first retrieve relevant information from our own systems.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User:
Why are checkout 500 errors increasing?

        ↓

Search knowledge base

        ↓

Find relevant documents

- checkout-runbook.md
- incident-2026-08.md
- nginx-errors.md

        ↓

LLM

        ↓

Answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI might respond:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The most likely cause is database connection exhaustion. The checkout runbook recommends checking the PostgreSQL connection pool and active connections."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The important part is that the model didn't necessarily know this beforehand.&lt;/p&gt;

&lt;p&gt;We &lt;strong&gt;retrieved the information and gave it to the model&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  2. Why is RAG useful for SRE?
&lt;/h1&gt;

&lt;p&gt;Most SRE organizations already have huge amounts of operational knowledge:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Runbooks&lt;/li&gt;
&lt;li&gt;Incident reports&lt;/li&gt;
&lt;li&gt;Postmortems&lt;/li&gt;
&lt;li&gt;Architecture documents&lt;/li&gt;
&lt;li&gt;Terraform repositories&lt;/li&gt;
&lt;li&gt;Kubernetes manifests&lt;/li&gt;
&lt;li&gt;Jenkins pipelines&lt;/li&gt;
&lt;li&gt;Troubleshooting guides&lt;/li&gt;
&lt;li&gt;Known issues&lt;/li&gt;
&lt;li&gt;Service documentation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The problem isn't necessarily lack of knowledge.&lt;/p&gt;

&lt;p&gt;The problem is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Finding the right knowledge at the right time.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Imagine you're paged at 2 AM.&lt;/p&gt;

&lt;p&gt;You see:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;checkout-api
HTTP 500
Error rate: 35%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Instead of manually searching through 200 runbooks, you ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What should I check for checkout-api 500 errors?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;RAG retrieves the relevant documentation and gives you a concise answer.&lt;/p&gt;

&lt;p&gt;That's already extremely useful.&lt;/p&gt;




&lt;h1&gt;
  
  
  3. But RAG has a limitation
&lt;/h1&gt;

&lt;p&gt;Here's where things get interesting.&lt;/p&gt;

&lt;p&gt;Imagine your documentation contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;checkout-api.md

Checkout API depends on Payment API.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And another document says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;payment-api.md

Payment API depends on PostgreSQL.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And another says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;database-runbook.md

PostgreSQL connection exhaustion can cause payment failures.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A RAG system can retrieve these documents.&lt;/p&gt;

&lt;p&gt;But the relationships are mostly &lt;strong&gt;implicit inside the text&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The AI has to figure out:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Checkout
   ↓
Payment API
   ↓
PostgreSQL
   ↓
Connection pool
   ↓
500 errors
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OKF becomes interesting here because knowledge can be deliberately structured, linked and made machine-readable instead of being treated as unrelated chunks of text. Practical OKF implementations commonly use Markdown with metadata/frontmatter and links between concepts.&lt;/p&gt;




&lt;h1&gt;
  
  
  4. Think about OKF
&lt;/h1&gt;

&lt;p&gt;Instead of only storing documents, we explicitly represent our operational knowledge.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Checkout Service
      │
      ├── depends_on → Payment API
      ├── deployed_by → Jenkins
      ├── monitored_by → Prometheus
      └── owned_by → Payments Team

Payment API
      │
      └── depends_on → PostgreSQL

PostgreSQL
      │
      ├── monitored_by → Prometheus
      └── has_runbook → DB-POOL-001
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the system doesn't just know &lt;strong&gt;what the documents say&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It knows &lt;strong&gt;how things are connected&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  5. Let's actually build a tiny OKF for SRE
&lt;/h1&gt;

&lt;p&gt;You don't need Kubernetes, a graph database or a huge AI platform to understand the concept.&lt;/p&gt;

&lt;p&gt;Let's build a tiny knowledge base for one imaginary service:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;checkout-api
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Our goal is to answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Why is checkout-api returning 500 errors?"&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Step 1: Create the knowledge structure
&lt;/h2&gt;

&lt;p&gt;Start with something simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sre-okf/
├── index.md
├── services/
│   ├── checkout-api.md
│   └── payment-api.md
├── infrastructure/
│   └── postgres.md
├── alerts/
│   └── checkout-500.md
└── runbooks/
    └── db-connection-pool.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important thing isn't the exact folder names.&lt;/p&gt;

&lt;p&gt;The important thing is that the knowledge is &lt;strong&gt;structured, predictable and linked&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  Step 2: Create the service definition
&lt;/h1&gt;

&lt;p&gt;Create:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;services/checkout-api.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;service&lt;/span&gt;
&lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-api&lt;/span&gt;
&lt;span class="na"&gt;owner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;payments-team&lt;/span&gt;
&lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then add the actual knowledge:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Checkout API&lt;/span&gt;

The Checkout API handles customer checkout requests.

&lt;span class="gu"&gt;## Dependencies&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Payment API&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;payment-api.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="gu"&gt;## Monitoring&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Prometheus metric: &lt;span class="sb"&gt;`checkout_http_requests_total`&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Grafana dashboard: Checkout Overview

&lt;span class="gu"&gt;## Alerts&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Checkout 500 Alert&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;../alerts/checkout-500.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice something important.&lt;/p&gt;

&lt;p&gt;We're not just writing prose.&lt;/p&gt;

&lt;p&gt;We're creating &lt;strong&gt;machine-readable metadata + human-readable content + links&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This combination is one of the useful patterns seen in practical OKF implementations.&lt;/p&gt;




&lt;h1&gt;
  
  
  Step 3: Define the dependency
&lt;/h1&gt;

&lt;p&gt;Create:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;services/payment-api.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;service&lt;/span&gt;
&lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;payment-api&lt;/span&gt;
&lt;span class="na"&gt;owner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;payments-team&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Payment API&lt;/span&gt;

The Payment API processes payment transactions.

&lt;span class="gu"&gt;## Dependencies&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;PostgreSQL&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;../infrastructure/postgres.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="gu"&gt;## Known Failure Modes&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Database connection pool exhaustion
&lt;span class="p"&gt;-&lt;/span&gt; Payment provider timeout
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now we have:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Checkout API
     ↓
Payment API
     ↓
PostgreSQL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Step 4: Define the infrastructure
&lt;/h1&gt;

&lt;p&gt;Create:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;infrastructure/postgres.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;infrastructure&lt;/span&gt;
&lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgres-payments&lt;/span&gt;
&lt;span class="na"&gt;technology&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;PostgreSQL&lt;/span&gt;
&lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# PostgreSQL&lt;/span&gt;

Primary database for Payment API.

&lt;span class="gu"&gt;## Used By&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Payment API&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;../services/payment-api.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="gu"&gt;## Important Metrics&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Active connections
&lt;span class="p"&gt;-&lt;/span&gt; Connection pool utilization
&lt;span class="p"&gt;-&lt;/span&gt; Query latency

&lt;span class="gu"&gt;## Runbook&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Database Connection Pool Runbook&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;../runbooks/db-connection-pool.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the relationship exists in both directions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Payment API
     ↓
PostgreSQL
     ↓
DB Connection Pool Runbook
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Step 5: Define the alert
&lt;/h1&gt;

&lt;p&gt;Create:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;alerts/checkout-500.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;alert&lt;/span&gt;
&lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-http-500&lt;/span&gt;
&lt;span class="na"&gt;severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;critical&lt;/span&gt;
&lt;span class="na"&gt;service&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;checkout-api&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Checkout HTTP 500&lt;/span&gt;

Triggered when:

&lt;span class="sb"&gt;`checkout_http_5xx_rate &amp;gt; 10%`&lt;/span&gt;

&lt;span class="gu"&gt;## Affected Service&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Checkout API&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;../services/checkout-api.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="gu"&gt;## Investigation Path&lt;/span&gt;

Check:
&lt;span class="p"&gt;
1.&lt;/span&gt; Checkout API
&lt;span class="p"&gt;2.&lt;/span&gt; Payment API
&lt;span class="p"&gt;3.&lt;/span&gt; PostgreSQL connections
&lt;span class="p"&gt;4.&lt;/span&gt; Database connection pool

&lt;span class="gu"&gt;## Related Runbook&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;DB Connection Pool Runbook&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;../runbooks/db-connection-pool.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now an alert isn't just a string.&lt;/p&gt;

&lt;p&gt;It has:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Alert
 ↓
Service
 ↓
Dependency
 ↓
Infrastructure
 ↓
Runbook
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Step 6: Add the runbook
&lt;/h1&gt;

&lt;p&gt;Create:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;runbooks/db-connection-pool.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;runbook&lt;/span&gt;
&lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;db-connection-pool&lt;/span&gt;
&lt;span class="na"&gt;owner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;payments-team&lt;/span&gt;
&lt;span class="na"&gt;severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;critical&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Database Connection Pool Runbook&lt;/span&gt;

Use this runbook when database connection exhaustion is suspected.

&lt;span class="gu"&gt;## Check&lt;/span&gt;

Prometheus:

&lt;span class="sb"&gt;`pg_stat_activity_count`&lt;/span&gt;

Check connection pool utilization.

&lt;span class="gu"&gt;## Loki&lt;/span&gt;

Search:

&lt;span class="sb"&gt;`"connection pool exhausted"`&lt;/span&gt;

&lt;span class="gu"&gt;## Remediation&lt;/span&gt;
&lt;span class="p"&gt;
1.&lt;/span&gt; Confirm active connections.
&lt;span class="p"&gt;2.&lt;/span&gt; Check for connection leaks.
&lt;span class="p"&gt;3.&lt;/span&gt; Check recent deployments.
&lt;span class="p"&gt;4.&lt;/span&gt; Restart affected workers if approved.
&lt;span class="p"&gt;5.&lt;/span&gt; Escalate to the database team if the issue persists.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now we have a very small operational knowledge system.&lt;/p&gt;




&lt;h1&gt;
  
  
  Step 7: Create an index
&lt;/h1&gt;

&lt;p&gt;Finally, create:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;index.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# SRE Knowledge Base&lt;/span&gt;

&lt;span class="gu"&gt;## Services&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Checkout API&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;services/checkout-api.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Payment API&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;services/payment-api.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="gu"&gt;## Infrastructure&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;PostgreSQL&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;infrastructure/postgres.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="gu"&gt;## Alerts&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Checkout 500&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;alerts/checkout-500.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="gu"&gt;## Runbooks&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;DB Connection Pool&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;runbooks/db-connection-pool.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The index gives an agent a predictable starting point instead of forcing it to blindly search everything.&lt;/p&gt;

&lt;p&gt;This "index + linked knowledge" approach is used by real OKF-style implementations to make knowledge navigable by both humans and agents.&lt;/p&gt;




&lt;h1&gt;
  
  
  6. Now connect an AI agent
&lt;/h1&gt;

&lt;p&gt;At this point, you don't necessarily need a vector database.&lt;/p&gt;

&lt;p&gt;A very simple OKF retrieval flow can be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User question
      ↓
Find starting concept
      ↓
Read metadata
      ↓
Follow related links
      ↓
Collect relevant context
      ↓
Send context to LLM
      ↓
Generate answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Question:

"Why is checkout returning 500s?"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent might traverse:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;checkout-500.md
      ↓
checkout-api.md
      ↓
payment-api.md
      ↓
postgres.md
      ↓
db-connection-pool.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That traversal path is important.&lt;/p&gt;

&lt;p&gt;The AI can explain &lt;strong&gt;why those pieces of knowledge were used&lt;/strong&gt;, rather than simply saying:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"These five chunks looked similar to your question."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Graph/linked traversal is one of the approaches demonstrated by OKF implementations, where the retrieved context can be represented as a traversal path through related knowledge files.&lt;/p&gt;




&lt;h1&gt;
  
  
  7. Now add your real SRE systems
&lt;/h1&gt;

&lt;p&gt;This is where it becomes useful.&lt;/p&gt;

&lt;p&gt;The OKF itself doesn't have to contain your live metrics.&lt;/p&gt;

&lt;p&gt;Instead, it can tell the AI &lt;strong&gt;where and how to get them&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OKF
 │
 ├── Checkout API
 │     ├── Prometheus metrics
 │     └── Grafana dashboard
 │
 ├── Payment API
 │     └── Loki logs
 │
 └── PostgreSQL
       └── Prometheus metrics
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then your agent can call the actual systems:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OKF
 ↓
Understand architecture
 ↓
Prometheus
 ↓
Loki
 ↓
Jenkins
 ↓
LLM
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now you're moving from a static knowledge base toward an &lt;strong&gt;AI-powered SRE investigation system&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  8. A real incident
&lt;/h1&gt;

&lt;p&gt;Imagine this happens:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;14:21
Checkout 500 rate → 32%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI receives the alert.&lt;/p&gt;

&lt;p&gt;It uses OKF:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Checkout
   ↓
Payment API
   ↓
PostgreSQL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then it queries Prometheus:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DB connections → 98%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then Loki:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;connection pool exhausted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then Jenkins:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Payment API deployed 4 minutes ago
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And finally the runbook:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DB connection pool exhaustion
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI can produce:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Likely cause:&lt;/strong&gt; Payment API deployment caused database connection-pool exhaustion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Evidence:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Checkout 500s increased at 14:21.&lt;/li&gt;
&lt;li&gt;Payment API was deployed at 14:17.&lt;/li&gt;
&lt;li&gt;PostgreSQL connection utilization is 98%.&lt;/li&gt;
&lt;li&gt;Loki shows connection-pool exhaustion.&lt;/li&gt;
&lt;li&gt;The affected dependency chain is Checkout → Payment API → PostgreSQL.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Recommended next step:&lt;/strong&gt; Follow &lt;code&gt;DB Connection Pool Runbook&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's a much more useful SRE assistant than a chatbot that simply searches documentation.&lt;/p&gt;




&lt;h1&gt;
  
  
  9. Where RAG fits
&lt;/h1&gt;

&lt;p&gt;Now imagine you have 10,000 historical incident reports.&lt;/p&gt;

&lt;p&gt;You probably don't want to manually create OKF relationships for every sentence.&lt;/p&gt;

&lt;p&gt;This is where RAG becomes useful.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 AI SRE Assistant
                        │
             ┌──────────┴──────────┐
             ↓                     ↓
            OKF                    RAG
             │                     │
       Structured              Historical
       knowledge               documents
             │                     │
       "How things             "What happened
        connect"                 before?"
             │                     │
             └──────────┬──────────┘
                        ↓
                  Live systems
              Prometheus / Loki
                        ↓
                       LLM
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So you might use:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OKF for stable operational knowledge&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Services
Dependencies
Owners
Architecture
Runbooks
Standards
Definitions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;RAG for large, changing collections&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident reports
Postmortems
Tickets
Slack discussions
Documentation
Historical investigations
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Observability for current state&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Prometheus
Loki
Grafana
Kubernetes
Jenkins
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  10. The practical SRE adoption path
&lt;/h1&gt;

&lt;p&gt;Don't try to build everything on day one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Phase 1 — Start small
&lt;/h3&gt;

&lt;p&gt;Create OKF for &lt;strong&gt;5–10 critical services&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Capture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Service
Owner
Dependencies
Dashboards
Alerts
Runbooks
Repository
Deployment pipeline
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  Phase 2 — Add relationships
&lt;/h3&gt;

&lt;p&gt;Start connecting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Service → Dependency
Service → Team
Service → Alert
Alert → Runbook
Service → Dashboard
Service → Repository
Service → Deployment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  Phase 3 — Connect live systems
&lt;/h3&gt;

&lt;p&gt;Add integrations with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Prometheus
Loki
Grafana
Kubernetes
Jenkins
Terraform
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  Phase 4 — Add RAG
&lt;/h3&gt;

&lt;p&gt;Use RAG for the huge amount of historical and unstructured knowledge:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Postmortems
Incident tickets
Slack
Architecture documents
Historical investigations
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  Phase 5 — Let the AI investigate
&lt;/h3&gt;

&lt;p&gt;Eventually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Alert
  ↓
OKF → Understand the system
  ↓
Prometheus → Check metrics
  ↓
Loki → Check logs
  ↓
Jenkins → Check changes
  ↓
RAG → Find similar incidents
  ↓
Runbook → Find remediation
  ↓
LLM → Explain the incident
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is where the real value starts appearing.&lt;/p&gt;




&lt;h1&gt;
  
  
  11. RAG vs OKF: the easiest mental model
&lt;/h1&gt;

&lt;p&gt;If you remember only three things:&lt;/p&gt;

&lt;h3&gt;
  
  
  RAG
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;"Find the information."&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  OKF
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;"Organize the knowledge and its relationships."&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Observability
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;"Tell me what's happening right now."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Put all three together and you get something much more interesting:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;An AI that doesn't just search your SRE documentation, but understands your infrastructure, follows the relationships between systems, and uses real-time telemetry to investigate incidents.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And that's where I think the real opportunity for &lt;strong&gt;AI-powered SRE&lt;/strong&gt; lies.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>okf</category>
      <category>rag</category>
      <category>sre</category>
    </item>
    <item>
      <title>Prompt Caching Explained Like You're Talking to a Smart Human (Not an AI Researcher)</title>
      <dc:creator>Anand Kannan</dc:creator>
      <pubDate>Wed, 20 May 2026 18:28:29 +0000</pubDate>
      <link>https://dev.to/anandindia93/prompt-caching-explained-like-youre-talking-to-a-smart-human-not-an-ai-researcher-26gf</link>
      <guid>https://dev.to/anandindia93/prompt-caching-explained-like-youre-talking-to-a-smart-human-not-an-ai-researcher-26gf</guid>
      <description>&lt;p&gt;If you've started using AI APIs for coding assistants, chatbots, agents, RAG systems, or internal copilots, you've probably heard the term:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Prompt Caching”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Most explanations online sound overly technical.&lt;/p&gt;

&lt;p&gt;So let’s explain it in the simplest possible way — while still understanding the &lt;em&gt;real engineering and cost impact&lt;/em&gt; behind it.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Real Problem AI Apps Face
&lt;/h1&gt;

&lt;p&gt;Every time you send a request to an LLM, the model has to process:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;system instructions&lt;/li&gt;
&lt;li&gt;chat history&lt;/li&gt;
&lt;li&gt;codebase context&lt;/li&gt;
&lt;li&gt;documentation&lt;/li&gt;
&lt;li&gt;logs&lt;/li&gt;
&lt;li&gt;user question&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Even if 95% of that information is identical to the previous request.&lt;/p&gt;

&lt;p&gt;That repeated processing is expensive.&lt;/p&gt;

&lt;p&gt;Especially in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;IDE copilots&lt;/li&gt;
&lt;li&gt;SRE assistants&lt;/li&gt;
&lt;li&gt;AI agents&lt;/li&gt;
&lt;li&gt;RAG pipelines&lt;/li&gt;
&lt;li&gt;enterprise chat systems&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  A Simple Tea Shop Analogy
&lt;/h1&gt;

&lt;p&gt;Imagine you go to a tea shop daily and always say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“One ginger tea, less sugar, extra hot.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The shopkeeper memorizes it after a few days.&lt;/p&gt;

&lt;p&gt;Now you just say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Same tea.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The shopkeeper doesn’t need the full instruction again.&lt;/p&gt;

&lt;p&gt;That is essentially prompt caching.&lt;/p&gt;




&lt;h1&gt;
  
  
  What Actually Happens Without Prompt Cache
&lt;/h1&gt;

&lt;p&gt;Suppose your AI coding assistant sends this every request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;System Prompt:
You are a senior SRE assistant.

Repository Structure:
- services/
- monitoring/
- infra/

Guidelines:
- Prefer safe kubectl commands
- Suggest rollback plans
- Follow GitOps practices

Previous Chat:
...
...
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Size:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;25,000 tokens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the user asks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Why is my Kubernetes pod crashing?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Another:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;300 tokens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Total request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;25,300 tokens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now imagine the user asks 20 debugging questions.&lt;/p&gt;

&lt;p&gt;Without caching:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;25,300 × 20
= 506,000 tokens processed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Huge cost.&lt;/p&gt;

&lt;p&gt;Huge latency.&lt;/p&gt;

&lt;p&gt;Huge waste.&lt;/p&gt;




&lt;h1&gt;
  
  
  What Prompt Caching Changes
&lt;/h1&gt;

&lt;p&gt;With prompt caching, the provider notices:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;“This large prefix is identical to the previous request.”
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So it reuses previously processed context.&lt;/p&gt;

&lt;p&gt;Meaning:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the model does NOT fully re-process repeated sections&lt;/li&gt;
&lt;li&gt;cached sections become cheaper&lt;/li&gt;
&lt;li&gt;responses become faster&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now only the new question changes.&lt;/p&gt;




&lt;h1&gt;
  
  
  Realistic Flow
&lt;/h1&gt;

&lt;h2&gt;
  
  
  First Request
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[Large Context]
+ New Question
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Provider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Processes everything normally
Creates cache
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Second Request
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[Same Large Context]
+ Another Question
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Provider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cache hit detected
Reuses processed prefix
Only processes delta efficiently
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  The Biggest Misunderstanding
&lt;/h1&gt;

&lt;p&gt;Many people think:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Cached tokens are free.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No.&lt;/p&gt;

&lt;p&gt;They are usually:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;discounted&lt;/li&gt;
&lt;li&gt;faster&lt;/li&gt;
&lt;li&gt;partially reused&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But not free.&lt;/p&gt;

&lt;p&gt;The infrastructure still:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;validates cache&lt;/li&gt;
&lt;li&gt;stores context&lt;/li&gt;
&lt;li&gt;reconstructs attention state&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  Real Cost Impact
&lt;/h1&gt;

&lt;p&gt;Let’s use simplified numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Without Cache
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requests&lt;/th&gt;
&lt;th&gt;Tokens Each&lt;/th&gt;
&lt;th&gt;Total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;25,300&lt;/td&gt;
&lt;td&gt;506,000&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  With Cache
&lt;/h2&gt;

&lt;p&gt;First request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;25,300
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Remaining 19 requests:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;large cached prefix reused&lt;/li&gt;
&lt;li&gt;only new tokens heavily charged&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Approx effective active processing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;25,300
+ (19 × 300)
= 31,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;506,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s not a tiny optimization.&lt;/p&gt;

&lt;p&gt;That is the difference between:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a scalable AI product&lt;/li&gt;
&lt;li&gt;and an unsustainable one&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  Why Coding Assistants Benefit Massively
&lt;/h1&gt;

&lt;p&gt;Coding tools repeatedly send:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;repo instructions&lt;/li&gt;
&lt;li&gt;architecture&lt;/li&gt;
&lt;li&gt;coding guidelines&lt;/li&gt;
&lt;li&gt;open files&lt;/li&gt;
&lt;li&gt;previous chat history&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most of this barely changes.&lt;/p&gt;

&lt;p&gt;Without prompt caching:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;token burn becomes enormous&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With prompt caching:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;only diffs matter&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is why modern AI IDEs feel dramatically faster than early-generation copilots.&lt;/p&gt;




&lt;h1&gt;
  
  
  The SRE / Observability Example
&lt;/h1&gt;

&lt;p&gt;As an SRE, imagine building an AI incident assistant.&lt;/p&gt;

&lt;p&gt;Every request includes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- Kubernetes topology
- Loki queries
- Grafana dashboards
- Service dependencies
- Runbooks
- Deployment metadata
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That alone might be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;40,000 tokens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the engineer asks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Why are pods restarting in prod?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without caching:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;40k+ tokens every query
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With caching:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Topology cached once
Only incident-specific delta processed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is where enterprises save:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;money&lt;/li&gt;
&lt;li&gt;latency&lt;/li&gt;
&lt;li&gt;compute&lt;/li&gt;
&lt;li&gt;API throughput&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  Prompt Caching vs Memory
&lt;/h1&gt;

&lt;p&gt;People confuse these two constantly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prompt Caching
&lt;/h2&gt;

&lt;p&gt;Purpose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Reduce repeated processing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Focus:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Efficiency
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Memory
&lt;/h2&gt;

&lt;p&gt;Purpose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Remember user information across sessions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Focus:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Personalization
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They are completely different systems.&lt;/p&gt;




&lt;h1&gt;
  
  
  Important Engineering Detail
&lt;/h1&gt;

&lt;p&gt;Prompt caching usually works best when:&lt;/p&gt;

&lt;p&gt;✅ Prompt prefixes remain stable&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[Same system prompt]
[Same repo instructions]
[Different question]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;p&gt;It works poorly when:&lt;/p&gt;

&lt;p&gt;❌ Entire prompt changes every request&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Completely unrelated documents each time
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Hidden Insight Most Tutorials Miss
&lt;/h1&gt;

&lt;p&gt;Prompt caching becomes exponentially valuable as context windows grow.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because modern LLM apps now send:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;100k+&lt;/li&gt;
&lt;li&gt;200k+&lt;/li&gt;
&lt;li&gt;even million-token contexts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without caching:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;costs explode&lt;/li&gt;
&lt;li&gt;latency becomes painful&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Caching is one of the major reasons large-context AI systems are economically feasible today.&lt;/p&gt;




&lt;h1&gt;
  
  
  Important Reality About GitHub Copilot
&lt;/h1&gt;

&lt;p&gt;Tools like GitHub Copilot likely use:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;internal context reuse&lt;/li&gt;
&lt;li&gt;session optimization&lt;/li&gt;
&lt;li&gt;embedding reuse&lt;/li&gt;
&lt;li&gt;smart prompt assembly&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But they usually do NOT expose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;manual cache control&lt;/li&gt;
&lt;li&gt;cache hit visibility&lt;/li&gt;
&lt;li&gt;cache TTL settings&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So users benefit indirectly.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Thought
&lt;/h1&gt;

&lt;p&gt;Prompt caching sounds like a small optimization feature.&lt;/p&gt;

&lt;p&gt;It isn’t.&lt;/p&gt;

&lt;p&gt;It is one of the foundational techniques making modern AI systems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;affordable&lt;/li&gt;
&lt;li&gt;fast&lt;/li&gt;
&lt;li&gt;scalable&lt;/li&gt;
&lt;li&gt;production-ready&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without it, many long-context AI applications would become economically impractical very quickly.&lt;/p&gt;

&lt;p&gt;Especially in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;copilots&lt;/li&gt;
&lt;li&gt;AI agents&lt;/li&gt;
&lt;li&gt;enterprise assistants&lt;/li&gt;
&lt;li&gt;observability platforms&lt;/li&gt;
&lt;li&gt;multi-turn coding workflows&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The bigger your context becomes, the more prompt caching stops being “nice to have” and becomes essential infrastructure.&lt;/p&gt;




&lt;h1&gt;
  
  
  ai #llm #opensource #programming #devops
&lt;/h1&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>performance</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>🚀 Introducing DevFocus: The Micro-SaaS App for Developers – Join the Waitlist Now! 💻</title>
      <dc:creator>Anand Kannan</dc:creator>
      <pubDate>Mon, 21 Apr 2025 17:02:17 +0000</pubDate>
      <link>https://dev.to/anandindia93/introducing-devfocus-the-micro-saas-app-for-developers-join-the-waitlist-now-3dbl</link>
      <guid>https://dev.to/anandindia93/introducing-devfocus-the-micro-saas-app-for-developers-join-the-waitlist-now-3dbl</guid>
      <description>&lt;h1&gt;
  
  
  🚀 DevFocus – A Micro-SaaS With Max Benefits
&lt;/h1&gt;

&lt;p&gt;Too many tools, too little focus?&lt;/p&gt;

&lt;p&gt;That’s what inspired us — two full-time SREs — to build &lt;strong&gt;DevFocus&lt;/strong&gt;: a tool that pulls your tasks from Slack, GitHub, Jira, Gmail, Outlook, Webex, and more into one unified dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  ✨ Why?
&lt;/h2&gt;

&lt;p&gt;In a distributed team, it's hard to communicate your current focus across tools. You’re constantly updating Jira, replying to Slack pings, writing daily standup notes, and syncing in status meetings.&lt;/p&gt;

&lt;p&gt;We asked: &lt;em&gt;What if all that happened automatically?&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ Key Features
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;🧠 &lt;strong&gt;Unified Task View&lt;/strong&gt; – Pulls from all major sources&lt;/li&gt;
&lt;li&gt;🔄 &lt;strong&gt;Cross-Tool Visibility&lt;/strong&gt; – Let your team know your focus&lt;/li&gt;
&lt;li&gt;🔒 &lt;strong&gt;Private &amp;amp; Secure&lt;/strong&gt; – OAuth2 + only what you allow&lt;/li&gt;
&lt;li&gt;💬 &lt;strong&gt;Team &amp;amp; Personal View&lt;/strong&gt; – Update-less updates!&lt;/li&gt;
&lt;li&gt;📅 &lt;strong&gt;Auto-detect Deadlines&lt;/strong&gt; – Spot urgency instantly&lt;/li&gt;
&lt;li&gt;📈 &lt;strong&gt;Less cognitive load, more flow time&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  &lt;a href="https://techsei.com/" rel="noopener noreferrer"&gt;🔗 Join the waitlist and be the first to try DevFocus!&lt;/a&gt;
&lt;/h2&gt;




&lt;p&gt;We're actively building this, and feedback from early devs is pure gold. Would love your thoughts or help testing!&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;💬 “A now-page for devs — without the manual work”&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>devops</category>
      <category>productivity</category>
      <category>microsaas</category>
    </item>
  </channel>
</rss>
