RAG vs OKF: What’s the Difference and Where Should SREs Use Them?
If you're working in SRE, you've probably heard a lot about RAG (Retrieval-Augmented Generation) and OKF (Open Knowledge Framework).
At first, they can sound like the same thing:
"Give an AI access to our documentation and let it answer questions."
But they're actually solving different problems.
Let's understand this without the AI buzzwords.
The 10-second explanation
Think of an AI system as an SRE engineer.
RAG says:
"Before answering, search our runbooks, documentation and incident reports."
OKF says:
"Don't just give me documents. Organize knowledge, relationships and context so an AI can navigate it consistently."
So:
RAG = Retrieve relevant information
OKF = Organize knowledge and relationships for AI consumption
They aren't competitors. In a good SRE system, you can use both.
1. What is RAG?
RAG stands for:
Retrieval-Augmented Generation
The idea is surprisingly simple.
Instead of asking an LLM a question and expecting it to know everything, we first retrieve relevant information from our own systems.
For example:
User:
Why are checkout 500 errors increasing?
↓
Search knowledge base
↓
Find relevant documents
- checkout-runbook.md
- incident-2026-08.md
- nginx-errors.md
↓
LLM
↓
Answer
The AI might respond:
"The most likely cause is database connection exhaustion. The checkout runbook recommends checking the PostgreSQL connection pool and active connections."
The important part is that the model didn't necessarily know this beforehand.
We retrieved the information and gave it to the model.
2. Why is RAG useful for SRE?
Most SRE organizations already have huge amounts of operational knowledge:
- Runbooks
- Incident reports
- Postmortems
- Architecture documents
- Terraform repositories
- Kubernetes manifests
- Jenkins pipelines
- Troubleshooting guides
- Known issues
- Service documentation
The problem isn't necessarily lack of knowledge.
The problem is:
Finding the right knowledge at the right time.
Imagine you're paged at 2 AM.
You see:
checkout-api
HTTP 500
Error rate: 35%
Instead of manually searching through 200 runbooks, you ask:
"What should I check for checkout-api 500 errors?"
RAG retrieves the relevant documentation and gives you a concise answer.
That's already extremely useful.
3. But RAG has a limitation
Here's where things get interesting.
Imagine your documentation contains:
checkout-api.md
Checkout API depends on Payment API.
And another document says:
payment-api.md
Payment API depends on PostgreSQL.
And another says:
database-runbook.md
PostgreSQL connection exhaustion can cause payment failures.
A RAG system can retrieve these documents.
But the relationships are mostly implicit inside the text.
The AI has to figure out:
Checkout
↓
Payment API
↓
PostgreSQL
↓
Connection pool
↓
500 errors
OKF becomes interesting here because knowledge can be deliberately structured, linked and made machine-readable instead of being treated as unrelated chunks of text. Practical OKF implementations commonly use Markdown with metadata/frontmatter and links between concepts.
4. Think about OKF
Instead of only storing documents, we explicitly represent our operational knowledge.
For example:
Checkout Service
│
├── depends_on → Payment API
├── deployed_by → Jenkins
├── monitored_by → Prometheus
└── owned_by → Payments Team
Payment API
│
└── depends_on → PostgreSQL
PostgreSQL
│
├── monitored_by → Prometheus
└── has_runbook → DB-POOL-001
Now the system doesn't just know what the documents say.
It knows how things are connected.
5. Let's actually build a tiny OKF for SRE
You don't need Kubernetes, a graph database or a huge AI platform to understand the concept.
Let's build a tiny knowledge base for one imaginary service:
checkout-api
Our goal is to answer:
"Why is checkout-api returning 500 errors?"
Step 1: Create the knowledge structure
Start with something simple:
sre-okf/
├── index.md
├── services/
│ ├── checkout-api.md
│ └── payment-api.md
├── infrastructure/
│ └── postgres.md
├── alerts/
│ └── checkout-500.md
└── runbooks/
└── db-connection-pool.md
The important thing isn't the exact folder names.
The important thing is that the knowledge is structured, predictable and linked.
Step 2: Create the service definition
Create:
services/checkout-api.md
Add:
---
type: service
id: checkout-api
owner: payments-team
environment: production
---
Then add the actual knowledge:
# Checkout API
The Checkout API handles customer checkout requests.
## Dependencies
- [Payment API](payment-api.md)
## Monitoring
- Prometheus metric: `checkout_http_requests_total`
- Grafana dashboard: Checkout Overview
## Alerts
- [Checkout 500 Alert](../alerts/checkout-500.md)
Notice something important.
We're not just writing prose.
We're creating machine-readable metadata + human-readable content + links.
This combination is one of the useful patterns seen in practical OKF implementations.
Step 3: Define the dependency
Create:
services/payment-api.md
---
type: service
id: payment-api
owner: payments-team
---
# Payment API
The Payment API processes payment transactions.
## Dependencies
- [PostgreSQL](../infrastructure/postgres.md)
## Known Failure Modes
- Database connection pool exhaustion
- Payment provider timeout
Now we have:
Checkout API
↓
Payment API
↓
PostgreSQL
Step 4: Define the infrastructure
Create:
infrastructure/postgres.md
---
type: infrastructure
id: postgres-payments
technology: PostgreSQL
environment: production
---
# PostgreSQL
Primary database for Payment API.
## Used By
- [Payment API](../services/payment-api.md)
## Important Metrics
- Active connections
- Connection pool utilization
- Query latency
## Runbook
- [Database Connection Pool Runbook](../runbooks/db-connection-pool.md)
Now the relationship exists in both directions.
Payment API
↓
PostgreSQL
↓
DB Connection Pool Runbook
Step 5: Define the alert
Create:
alerts/checkout-500.md
---
type: alert
id: checkout-http-500
severity: critical
service: checkout-api
---
# Checkout HTTP 500
Triggered when:
`checkout_http_5xx_rate > 10%`
## Affected Service
- [Checkout API](../services/checkout-api.md)
## Investigation Path
Check:
1. Checkout API
2. Payment API
3. PostgreSQL connections
4. Database connection pool
## Related Runbook
- [DB Connection Pool Runbook](../runbooks/db-connection-pool.md)
Now an alert isn't just a string.
It has:
Alert
↓
Service
↓
Dependency
↓
Infrastructure
↓
Runbook
Step 6: Add the runbook
Create:
runbooks/db-connection-pool.md
---
type: runbook
id: db-connection-pool
owner: payments-team
severity: critical
---
# Database Connection Pool Runbook
Use this runbook when database connection exhaustion is suspected.
## Check
Prometheus:
`pg_stat_activity_count`
Check connection pool utilization.
## Loki
Search:
`"connection pool exhausted"`
## Remediation
1. Confirm active connections.
2. Check for connection leaks.
3. Check recent deployments.
4. Restart affected workers if approved.
5. Escalate to the database team if the issue persists.
Now we have a very small operational knowledge system.
Step 7: Create an index
Finally, create:
index.md
# SRE Knowledge Base
## Services
- [Checkout API](services/checkout-api.md)
- [Payment API](services/payment-api.md)
## Infrastructure
- [PostgreSQL](infrastructure/postgres.md)
## Alerts
- [Checkout 500](alerts/checkout-500.md)
## Runbooks
- [DB Connection Pool](runbooks/db-connection-pool.md)
The index gives an agent a predictable starting point instead of forcing it to blindly search everything.
This "index + linked knowledge" approach is used by real OKF-style implementations to make knowledge navigable by both humans and agents.
6. Now connect an AI agent
At this point, you don't necessarily need a vector database.
A very simple OKF retrieval flow can be:
User question
↓
Find starting concept
↓
Read metadata
↓
Follow related links
↓
Collect relevant context
↓
Send context to LLM
↓
Generate answer
For example:
Question:
"Why is checkout returning 500s?"
The agent might traverse:
checkout-500.md
↓
checkout-api.md
↓
payment-api.md
↓
postgres.md
↓
db-connection-pool.md
That traversal path is important.
The AI can explain why those pieces of knowledge were used, rather than simply saying:
"These five chunks looked similar to your question."
Graph/linked traversal is one of the approaches demonstrated by OKF implementations, where the retrieved context can be represented as a traversal path through related knowledge files.
7. Now add your real SRE systems
This is where it becomes useful.
The OKF itself doesn't have to contain your live metrics.
Instead, it can tell the AI where and how to get them.
For example:
OKF
│
├── Checkout API
│ ├── Prometheus metrics
│ └── Grafana dashboard
│
├── Payment API
│ └── Loki logs
│
└── PostgreSQL
└── Prometheus metrics
Then your agent can call the actual systems:
OKF
↓
Understand architecture
↓
Prometheus
↓
Loki
↓
Jenkins
↓
LLM
Now you're moving from a static knowledge base toward an AI-powered SRE investigation system.
8. A real incident
Imagine this happens:
14:21
Checkout 500 rate → 32%
The AI receives the alert.
It uses OKF:
Checkout
↓
Payment API
↓
PostgreSQL
Then it queries Prometheus:
DB connections → 98%
Then Loki:
connection pool exhausted
Then Jenkins:
Payment API deployed 4 minutes ago
And finally the runbook:
DB connection pool exhaustion
The AI can produce:
Likely cause: Payment API deployment caused database connection-pool exhaustion.
Evidence:
- Checkout 500s increased at 14:21.
- Payment API was deployed at 14:17.
- PostgreSQL connection utilization is 98%.
- Loki shows connection-pool exhaustion.
- The affected dependency chain is Checkout → Payment API → PostgreSQL.
Recommended next step: Follow
DB Connection Pool Runbook.
That's a much more useful SRE assistant than a chatbot that simply searches documentation.
9. Where RAG fits
Now imagine you have 10,000 historical incident reports.
You probably don't want to manually create OKF relationships for every sentence.
This is where RAG becomes useful.
AI SRE Assistant
│
┌──────────┴──────────┐
↓ ↓
OKF RAG
│ │
Structured Historical
knowledge documents
│ │
"How things "What happened
connect" before?"
│ │
└──────────┬──────────┘
↓
Live systems
Prometheus / Loki
↓
LLM
So you might use:
OKF for stable operational knowledge
Services
Dependencies
Owners
Architecture
Runbooks
Standards
Definitions
RAG for large, changing collections
Incident reports
Postmortems
Tickets
Slack discussions
Documentation
Historical investigations
Observability for current state
Prometheus
Loki
Grafana
Kubernetes
Jenkins
10. The practical SRE adoption path
Don't try to build everything on day one.
Phase 1 — Start small
Create OKF for 5–10 critical services.
Capture:
Service
Owner
Dependencies
Dashboards
Alerts
Runbooks
Repository
Deployment pipeline
Phase 2 — Add relationships
Start connecting:
Service → Dependency
Service → Team
Service → Alert
Alert → Runbook
Service → Dashboard
Service → Repository
Service → Deployment
Phase 3 — Connect live systems
Add integrations with:
Prometheus
Loki
Grafana
Kubernetes
Jenkins
Terraform
Phase 4 — Add RAG
Use RAG for the huge amount of historical and unstructured knowledge:
Postmortems
Incident tickets
Slack
Architecture documents
Historical investigations
Phase 5 — Let the AI investigate
Eventually:
Alert
↓
OKF → Understand the system
↓
Prometheus → Check metrics
↓
Loki → Check logs
↓
Jenkins → Check changes
↓
RAG → Find similar incidents
↓
Runbook → Find remediation
↓
LLM → Explain the incident
That is where the real value starts appearing.
11. RAG vs OKF: the easiest mental model
If you remember only three things:
RAG
"Find the information."
OKF
"Organize the knowledge and its relationships."
Observability
"Tell me what's happening right now."
Put all three together and you get something much more interesting:
An AI that doesn't just search your SRE documentation, but understands your infrastructure, follows the relationships between systems, and uses real-time telemetry to investigate incidents.
And that's where I think the real opportunity for AI-powered SRE lies.
Top comments (0)