DEV Community

Anand Kannan
Anand Kannan

Posted on

RAG vs OKF: What’s the Difference and Where Should SREs Use Them?

RAG vs OKF: What’s the Difference and Where Should SREs Use Them?

If you're working in SRE, you've probably heard a lot about RAG (Retrieval-Augmented Generation) and OKF (Open Knowledge Framework).

At first, they can sound like the same thing:

"Give an AI access to our documentation and let it answer questions."

But they're actually solving different problems.

Let's understand this without the AI buzzwords.


The 10-second explanation

Think of an AI system as an SRE engineer.

RAG says:

"Before answering, search our runbooks, documentation and incident reports."

OKF says:

"Don't just give me documents. Organize knowledge, relationships and context so an AI can navigate it consistently."

So:

RAG = Retrieve relevant information

OKF = Organize knowledge and relationships for AI consumption

They aren't competitors. In a good SRE system, you can use both.


1. What is RAG?

RAG stands for:

Retrieval-Augmented Generation

The idea is surprisingly simple.

Instead of asking an LLM a question and expecting it to know everything, we first retrieve relevant information from our own systems.

For example:

User:
Why are checkout 500 errors increasing?

        ↓

Search knowledge base

        ↓

Find relevant documents

- checkout-runbook.md
- incident-2026-08.md
- nginx-errors.md

        ↓

LLM

        ↓

Answer
Enter fullscreen mode Exit fullscreen mode

The AI might respond:

"The most likely cause is database connection exhaustion. The checkout runbook recommends checking the PostgreSQL connection pool and active connections."

The important part is that the model didn't necessarily know this beforehand.

We retrieved the information and gave it to the model.


2. Why is RAG useful for SRE?

Most SRE organizations already have huge amounts of operational knowledge:

  • Runbooks
  • Incident reports
  • Postmortems
  • Architecture documents
  • Terraform repositories
  • Kubernetes manifests
  • Jenkins pipelines
  • Troubleshooting guides
  • Known issues
  • Service documentation

The problem isn't necessarily lack of knowledge.

The problem is:

Finding the right knowledge at the right time.

Imagine you're paged at 2 AM.

You see:

checkout-api
HTTP 500
Error rate: 35%
Enter fullscreen mode Exit fullscreen mode

Instead of manually searching through 200 runbooks, you ask:

"What should I check for checkout-api 500 errors?"

RAG retrieves the relevant documentation and gives you a concise answer.

That's already extremely useful.


3. But RAG has a limitation

Here's where things get interesting.

Imagine your documentation contains:

checkout-api.md

Checkout API depends on Payment API.
Enter fullscreen mode Exit fullscreen mode

And another document says:

payment-api.md

Payment API depends on PostgreSQL.
Enter fullscreen mode Exit fullscreen mode

And another says:

database-runbook.md

PostgreSQL connection exhaustion can cause payment failures.
Enter fullscreen mode Exit fullscreen mode

A RAG system can retrieve these documents.

But the relationships are mostly implicit inside the text.

The AI has to figure out:

Checkout
   ↓
Payment API
   ↓
PostgreSQL
   ↓
Connection pool
   ↓
500 errors
Enter fullscreen mode Exit fullscreen mode

OKF becomes interesting here because knowledge can be deliberately structured, linked and made machine-readable instead of being treated as unrelated chunks of text. Practical OKF implementations commonly use Markdown with metadata/frontmatter and links between concepts.


4. Think about OKF

Instead of only storing documents, we explicitly represent our operational knowledge.

For example:

Checkout Service
      │
      ├── depends_on → Payment API
      ├── deployed_by → Jenkins
      ├── monitored_by → Prometheus
      └── owned_by → Payments Team

Payment API
      │
      └── depends_on → PostgreSQL

PostgreSQL
      │
      ├── monitored_by → Prometheus
      └── has_runbook → DB-POOL-001
Enter fullscreen mode Exit fullscreen mode

Now the system doesn't just know what the documents say.

It knows how things are connected.


5. Let's actually build a tiny OKF for SRE

You don't need Kubernetes, a graph database or a huge AI platform to understand the concept.

Let's build a tiny knowledge base for one imaginary service:

checkout-api
Enter fullscreen mode Exit fullscreen mode

Our goal is to answer:

"Why is checkout-api returning 500 errors?"


Step 1: Create the knowledge structure

Start with something simple:

sre-okf/
├── index.md
├── services/
│   ├── checkout-api.md
│   └── payment-api.md
├── infrastructure/
│   └── postgres.md
├── alerts/
│   └── checkout-500.md
└── runbooks/
    └── db-connection-pool.md
Enter fullscreen mode Exit fullscreen mode

The important thing isn't the exact folder names.

The important thing is that the knowledge is structured, predictable and linked.


Step 2: Create the service definition

Create:

services/checkout-api.md
Enter fullscreen mode Exit fullscreen mode

Add:

---
type: service
id: checkout-api
owner: payments-team
environment: production
---
Enter fullscreen mode Exit fullscreen mode

Then add the actual knowledge:

# Checkout API

The Checkout API handles customer checkout requests.

## Dependencies

- [Payment API](payment-api.md)

## Monitoring

- Prometheus metric: `checkout_http_requests_total`
- Grafana dashboard: Checkout Overview

## Alerts

- [Checkout 500 Alert](../alerts/checkout-500.md)
Enter fullscreen mode Exit fullscreen mode

Notice something important.

We're not just writing prose.

We're creating machine-readable metadata + human-readable content + links.

This combination is one of the useful patterns seen in practical OKF implementations.


Step 3: Define the dependency

Create:

services/payment-api.md
Enter fullscreen mode Exit fullscreen mode
---
type: service
id: payment-api
owner: payments-team
---
Enter fullscreen mode Exit fullscreen mode
# Payment API

The Payment API processes payment transactions.

## Dependencies

- [PostgreSQL](../infrastructure/postgres.md)

## Known Failure Modes

- Database connection pool exhaustion
- Payment provider timeout
Enter fullscreen mode Exit fullscreen mode

Now we have:

Checkout API
     ↓
Payment API
     ↓
PostgreSQL
Enter fullscreen mode Exit fullscreen mode

Step 4: Define the infrastructure

Create:

infrastructure/postgres.md
Enter fullscreen mode Exit fullscreen mode
---
type: infrastructure
id: postgres-payments
technology: PostgreSQL
environment: production
---
Enter fullscreen mode Exit fullscreen mode
# PostgreSQL

Primary database for Payment API.

## Used By

- [Payment API](../services/payment-api.md)

## Important Metrics

- Active connections
- Connection pool utilization
- Query latency

## Runbook

- [Database Connection Pool Runbook](../runbooks/db-connection-pool.md)
Enter fullscreen mode Exit fullscreen mode

Now the relationship exists in both directions.

Payment API
     ↓
PostgreSQL
     ↓
DB Connection Pool Runbook
Enter fullscreen mode Exit fullscreen mode

Step 5: Define the alert

Create:

alerts/checkout-500.md
Enter fullscreen mode Exit fullscreen mode
---
type: alert
id: checkout-http-500
severity: critical
service: checkout-api
---
Enter fullscreen mode Exit fullscreen mode
# Checkout HTTP 500

Triggered when:

`checkout_http_5xx_rate > 10%`

## Affected Service

- [Checkout API](../services/checkout-api.md)

## Investigation Path

Check:

1. Checkout API
2. Payment API
3. PostgreSQL connections
4. Database connection pool

## Related Runbook

- [DB Connection Pool Runbook](../runbooks/db-connection-pool.md)
Enter fullscreen mode Exit fullscreen mode

Now an alert isn't just a string.

It has:

Alert
 ↓
Service
 ↓
Dependency
 ↓
Infrastructure
 ↓
Runbook
Enter fullscreen mode Exit fullscreen mode

Step 6: Add the runbook

Create:

runbooks/db-connection-pool.md
Enter fullscreen mode Exit fullscreen mode
---
type: runbook
id: db-connection-pool
owner: payments-team
severity: critical
---
Enter fullscreen mode Exit fullscreen mode
# Database Connection Pool Runbook

Use this runbook when database connection exhaustion is suspected.

## Check

Prometheus:

`pg_stat_activity_count`

Check connection pool utilization.

## Loki

Search:

`"connection pool exhausted"`

## Remediation

1. Confirm active connections.
2. Check for connection leaks.
3. Check recent deployments.
4. Restart affected workers if approved.
5. Escalate to the database team if the issue persists.
Enter fullscreen mode Exit fullscreen mode

Now we have a very small operational knowledge system.


Step 7: Create an index

Finally, create:

index.md
Enter fullscreen mode Exit fullscreen mode
# SRE Knowledge Base

## Services

- [Checkout API](services/checkout-api.md)
- [Payment API](services/payment-api.md)

## Infrastructure

- [PostgreSQL](infrastructure/postgres.md)

## Alerts

- [Checkout 500](alerts/checkout-500.md)

## Runbooks

- [DB Connection Pool](runbooks/db-connection-pool.md)
Enter fullscreen mode Exit fullscreen mode

The index gives an agent a predictable starting point instead of forcing it to blindly search everything.

This "index + linked knowledge" approach is used by real OKF-style implementations to make knowledge navigable by both humans and agents.


6. Now connect an AI agent

At this point, you don't necessarily need a vector database.

A very simple OKF retrieval flow can be:

User question
      ↓
Find starting concept
      ↓
Read metadata
      ↓
Follow related links
      ↓
Collect relevant context
      ↓
Send context to LLM
      ↓
Generate answer
Enter fullscreen mode Exit fullscreen mode

For example:

Question:

"Why is checkout returning 500s?"
Enter fullscreen mode Exit fullscreen mode

The agent might traverse:

checkout-500.md
      ↓
checkout-api.md
      ↓
payment-api.md
      ↓
postgres.md
      ↓
db-connection-pool.md
Enter fullscreen mode Exit fullscreen mode

That traversal path is important.

The AI can explain why those pieces of knowledge were used, rather than simply saying:

"These five chunks looked similar to your question."

Graph/linked traversal is one of the approaches demonstrated by OKF implementations, where the retrieved context can be represented as a traversal path through related knowledge files.


7. Now add your real SRE systems

This is where it becomes useful.

The OKF itself doesn't have to contain your live metrics.

Instead, it can tell the AI where and how to get them.

For example:

OKF
 │
 ├── Checkout API
 │     ├── Prometheus metrics
 │     └── Grafana dashboard
 │
 ├── Payment API
 │     └── Loki logs
 │
 └── PostgreSQL
       └── Prometheus metrics
Enter fullscreen mode Exit fullscreen mode

Then your agent can call the actual systems:

OKF
 ↓
Understand architecture
 ↓
Prometheus
 ↓
Loki
 ↓
Jenkins
 ↓
LLM
Enter fullscreen mode Exit fullscreen mode

Now you're moving from a static knowledge base toward an AI-powered SRE investigation system.


8. A real incident

Imagine this happens:

14:21
Checkout 500 rate → 32%
Enter fullscreen mode Exit fullscreen mode

The AI receives the alert.

It uses OKF:

Checkout
   ↓
Payment API
   ↓
PostgreSQL
Enter fullscreen mode Exit fullscreen mode

Then it queries Prometheus:

DB connections → 98%
Enter fullscreen mode Exit fullscreen mode

Then Loki:

connection pool exhausted
Enter fullscreen mode Exit fullscreen mode

Then Jenkins:

Payment API deployed 4 minutes ago
Enter fullscreen mode Exit fullscreen mode

And finally the runbook:

DB connection pool exhaustion
Enter fullscreen mode Exit fullscreen mode

The AI can produce:

Likely cause: Payment API deployment caused database connection-pool exhaustion.

Evidence:

  • Checkout 500s increased at 14:21.
  • Payment API was deployed at 14:17.
  • PostgreSQL connection utilization is 98%.
  • Loki shows connection-pool exhaustion.
  • The affected dependency chain is Checkout → Payment API → PostgreSQL.

Recommended next step: Follow DB Connection Pool Runbook.

That's a much more useful SRE assistant than a chatbot that simply searches documentation.


9. Where RAG fits

Now imagine you have 10,000 historical incident reports.

You probably don't want to manually create OKF relationships for every sentence.

This is where RAG becomes useful.

                 AI SRE Assistant
                        │
             ┌──────────┴──────────┐
             ↓                     ↓
            OKF                    RAG
             │                     │
       Structured              Historical
       knowledge               documents
             │                     │
       "How things             "What happened
        connect"                 before?"
             │                     │
             └──────────┬──────────┘
                        ↓
                  Live systems
              Prometheus / Loki
                        ↓
                       LLM
Enter fullscreen mode Exit fullscreen mode

So you might use:

OKF for stable operational knowledge

Services
Dependencies
Owners
Architecture
Runbooks
Standards
Definitions
Enter fullscreen mode Exit fullscreen mode

RAG for large, changing collections

Incident reports
Postmortems
Tickets
Slack discussions
Documentation
Historical investigations
Enter fullscreen mode Exit fullscreen mode

Observability for current state

Prometheus
Loki
Grafana
Kubernetes
Jenkins
Enter fullscreen mode Exit fullscreen mode

10. The practical SRE adoption path

Don't try to build everything on day one.

Phase 1 — Start small

Create OKF for 5–10 critical services.

Capture:

Service
Owner
Dependencies
Dashboards
Alerts
Runbooks
Repository
Deployment pipeline
Enter fullscreen mode Exit fullscreen mode

Phase 2 — Add relationships

Start connecting:

Service → Dependency
Service → Team
Service → Alert
Alert → Runbook
Service → Dashboard
Service → Repository
Service → Deployment
Enter fullscreen mode Exit fullscreen mode

Phase 3 — Connect live systems

Add integrations with:

Prometheus
Loki
Grafana
Kubernetes
Jenkins
Terraform
Enter fullscreen mode Exit fullscreen mode

Phase 4 — Add RAG

Use RAG for the huge amount of historical and unstructured knowledge:

Postmortems
Incident tickets
Slack
Architecture documents
Historical investigations
Enter fullscreen mode Exit fullscreen mode

Phase 5 — Let the AI investigate

Eventually:

Alert
  ↓
OKF → Understand the system
  ↓
Prometheus → Check metrics
  ↓
Loki → Check logs
  ↓
Jenkins → Check changes
  ↓
RAG → Find similar incidents
  ↓
Runbook → Find remediation
  ↓
LLM → Explain the incident
Enter fullscreen mode Exit fullscreen mode

That is where the real value starts appearing.


11. RAG vs OKF: the easiest mental model

If you remember only three things:

RAG

"Find the information."

OKF

"Organize the knowledge and its relationships."

Observability

"Tell me what's happening right now."

Put all three together and you get something much more interesting:

An AI that doesn't just search your SRE documentation, but understands your infrastructure, follows the relationships between systems, and uses real-time telemetry to investigate incidents.

And that's where I think the real opportunity for AI-powered SRE lies.

Top comments (0)