DEV Community

#rag

Retrieval augmented generation, or RAG, is an architectural approach that can improve the efficacy of large language model (LLM) applications by leveraging custom data.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
RAG in 8 Layers: The Production Mental Model Most Tutorials Skip

RAG in 8 Layers: The Production Mental Model Most Tutorials Skip

Comments
16 min read
TradeMemory An AI-Powered Persistence Layer for Disciplined Trading

TradeMemory An AI-Powered Persistence Layer for Disciplined Trading

1
Comments
3 min read
The Data Ingestion Pipeline Nobody Designs Well Until Production Breaks It

The Data Ingestion Pipeline Nobody Designs Well Until Production Breaks It

Comments
5 min read
TradeMemory An AI-Powered Persistence Layer for Disciplined Trading

TradeMemory An AI-Powered Persistence Layer for Disciplined Trading

Comments
3 min read
Citation-Guard: Production RAG Patterns for Regulated Fintech

Citation-Guard: Production RAG Patterns for Regulated Fintech

Comments
4 min read
From Documents to Intelligent Answers: Building a RAG Agent from Scratch & Lessons Learned

From Documents to Intelligent Answers: Building a RAG Agent from Scratch & Lessons Learned

5
Comments 1
2 min read
Local LLM Deployment, Agent Handbook, & LLM Cost Reduction: Applied AI Workflows

Local LLM Deployment, Agent Handbook, & LLM Cost Reduction: Applied AI Workflows

1
Comments 1
3 min read
What a Production RAG System Actually Looks Like After 18 Months

What a Production RAG System Actually Looks Like After 18 Months

1
Comments 3
6 min read
Fine-Tuning and RAG: What a Dozen Failed Experiments Taught Me

Fine-Tuning and RAG: What a Dozen Failed Experiments Taught Me

1
Comments
7 min read
TradeMemory

TradeMemory

1
Comments
3 min read
I Built a Production RAG System on My M1 Mac for $0

I Built a Production RAG System on My M1 Mac for $0

Comments
3 min read
Reciprocal Rerank Fusion (RRF): The Simple, Powerful Way to Combine Keyword + Semantic Search in RAG

Reciprocal Rerank Fusion (RRF): The Simple, Powerful Way to Combine Keyword + Semantic Search in RAG

2
Comments
3 min read
The Prefill Wall: Why MTP's 2 Barely Moves Long-Context Latency (Qwen3.6-27B, RTX 3090)

The Prefill Wall: Why MTP's 2 Barely Moves Long-Context Latency (Qwen3.6-27B, RTX 3090)

Comments
4 min read
Building a Robust RAG Pipeline Architecture for Production

Building a Robust RAG Pipeline Architecture for Production

Comments
7 min read
Practical RAG chunking strategies for reliable RAG pipelines

Practical RAG chunking strategies for reliable RAG pipelines

Comments
7 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.