Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
llm
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Building a Persistent Knowledge Base RAG System with FastAPI, llama.cpp, Chroma, and Open WebUI
navid mirnouri
navid mirnouri
navid mirnouri
Follow
Apr 30
Building a Persistent Knowledge Base RAG System with FastAPI, llama.cpp, Chroma, and Open WebUI
#
ai
#
llm
#
programming
#
python
1
 reaction
Comments
Add Comment
7 min read
The database is where AI agents in production get weird
Luca Sannitu
Luca Sannitu
Luca Sannitu
Follow
Apr 30
The database is where AI agents in production get weird
#
ai
#
database
#
sql
#
llm
Comments
Add Comment
2 min read
How to Reduce Token Usage in OpenCode with Dynamic Context Pruning (DCP)
piratesshield
piratesshield
piratesshield
Follow
Apr 30
How to Reduce Token Usage in OpenCode with Dynamic Context Pruning (DCP)
#
ai
#
llm
#
productivity
#
softwaredevelopment
1
 reaction
Comments
Add Comment
2 min read
A Three-Layer Memory Architecture for LLMs (Redis + Postgres + Vector) MCP
jinho von choi
jinho von choi
jinho von choi
Follow
Mar 27
A Three-Layer Memory Architecture for LLMs (Redis + Postgres + Vector) MCP
#
mcp
#
ai
#
llm
#
rag
Comments
Add Comment
2 min read
I was worried about the lack of security in shared .cursorrules, so I built a static analyzer to audit them.
Hugo Damion
Hugo Damion
Hugo Damion
Follow
Mar 27
I was worried about the lack of security in shared .cursorrules, so I built a static analyzer to audit them.
#
cursor
#
llm
#
claudecode
#
webdev
1
 reaction
Comments
Add Comment
1 min read
When Your AI Elaborates, It Forgets to Count
Kuro
Kuro
Kuro
Follow
Mar 27
When Your AI Elaborates, It Forgets to Count
#
ai
#
programming
#
llm
#
debugging
Comments
Add Comment
2 min read
GPU-Accelerated LLMs: Serving at 1M Tok/s, Voxtral TTS, & 4-bit Weight Quantization
soy
soy
soy
Follow
Mar 27
GPU-Accelerated LLMs: Serving at 1M Tok/s, Voxtral TTS, & 4-bit Weight Quantization
#
ai
#
machinelearning
#
llm
Comments
Add Comment
3 min read
Cencori: A Serverless Infrastructure Layer for Secure and Scalable AI Applications
Ladipo Samuel
Ladipo Samuel
Ladipo Samuel
Follow
Apr 30
Cencori: A Serverless Infrastructure Layer for Secure and Scalable AI Applications
#
ai
#
llm
#
security
#
serverless
2
 reactions
Comments
Add Comment
5 min read
CrewAI vs LangGraph vs AutoGen: Which Multi-Agent Framework Should You Use in 2026?
Rishabh Sethia
Rishabh Sethia
Rishabh Sethia
Follow
Apr 30
CrewAI vs LangGraph vs AutoGen: Which Multi-Agent Framework Should You Use in 2026?
#
agents
#
ai
#
automation
#
llm
1
 reaction
Comments
1
 comment
9 min read
Can your AI model smell bullsh1t? BullshitBench has the receipts.
Andrew Kew
Andrew Kew
Andrew Kew
Follow
Apr 30
Can your AI model smell bullsh1t? BullshitBench has the receipts.
#
discuss
#
ai
#
llm
#
machinelearning
Comments
Add Comment
3 min read
How We Use LLM Agents + CRM APIs to Auto-Generate Contextual Follow-Up Emails
SpurIQ Engineering
SpurIQ Engineering
SpurIQ Engineering
Follow
Apr 30
How We Use LLM Agents + CRM APIs to Auto-Generate Contextual Follow-Up Emails
#
ai
#
python
#
llm
#
api
1
 reaction
Comments
Add Comment
6 min read
Structured Outputs Are the Contract Your AI Agent Is Missing
Sitaram Srivatsavai
Sitaram Srivatsavai
Sitaram Srivatsavai
Follow
Mar 26
Structured Outputs Are the Contract Your AI Agent Is Missing
#
ai
#
json
#
architecture
#
llm
3
 reactions
Comments
Add Comment
5 min read
Hermes Agent Memory System: How Persistent AI Memory Actually Works
Rost
Rost
Rost
Follow
Apr 30
Hermes Agent Memory System: How Persistent AI Memory Actually Works
#
hermes
#
architecture
#
selfhosting
#
llm
2
 reactions
Comments
Add Comment
15 min read
Ollama in Docker Compose with GPU and Persistent Model Storage
Rost
Rost
Rost
Follow
Mar 27
Ollama in Docker Compose with GPU and Persistent Model Storage
#
selfhosting
#
llm
#
ollama
#
devops
1
 reaction
Comments
Add Comment
10 min read
The Routing Pattern: How Smart Teams Actually Use Fast vs Capable Models
Aamer Mihaysi
Aamer Mihaysi
Aamer Mihaysi
Follow
Mar 27
The Routing Pattern: How Smart Teams Actually Use Fast vs Capable Models
#
ai
#
agents
#
llm
Comments
Add Comment
2 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account