DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Building a Persistent Knowledge Base RAG System with FastAPI, llama.cpp, Chroma, and Open WebUI

Building a Persistent Knowledge Base RAG System with FastAPI, llama.cpp, Chroma, and Open WebUI

1
Comments
7 min read
The database is where AI agents in production get weird

The database is where AI agents in production get weird

Comments
2 min read
How to Reduce Token Usage in OpenCode with Dynamic Context Pruning (DCP)

How to Reduce Token Usage in OpenCode with Dynamic Context Pruning (DCP)

1
Comments
2 min read
A Three-Layer Memory Architecture for LLMs (Redis + Postgres + Vector) MCP

A Three-Layer Memory Architecture for LLMs (Redis + Postgres + Vector) MCP

Comments
2 min read
I was worried about the lack of security in shared .cursorrules, so I built a static analyzer to audit them.

I was worried about the lack of security in shared .cursorrules, so I built a static analyzer to audit them.

1
Comments
1 min read
When Your AI Elaborates, It Forgets to Count

When Your AI Elaborates, It Forgets to Count

Comments
2 min read
GPU-Accelerated LLMs: Serving at 1M Tok/s, Voxtral TTS, & 4-bit Weight Quantization

GPU-Accelerated LLMs: Serving at 1M Tok/s, Voxtral TTS, & 4-bit Weight Quantization

Comments
3 min read
Cencori: A Serverless Infrastructure Layer for Secure and Scalable AI Applications

Cencori: A Serverless Infrastructure Layer for Secure and Scalable AI Applications

2
Comments
5 min read
CrewAI vs LangGraph vs AutoGen: Which Multi-Agent Framework Should You Use in 2026?

CrewAI vs LangGraph vs AutoGen: Which Multi-Agent Framework Should You Use in 2026?

1
Comments 1
9 min read
Can your AI model smell bullsh1t? BullshitBench has the receipts.

Can your AI model smell bullsh1t? BullshitBench has the receipts.

Comments
3 min read
How We Use LLM Agents + CRM APIs to Auto-Generate Contextual Follow-Up Emails

How We Use LLM Agents + CRM APIs to Auto-Generate Contextual Follow-Up Emails

1
Comments
6 min read
Structured Outputs Are the Contract Your AI Agent Is Missing

Structured Outputs Are the Contract Your AI Agent Is Missing

3
Comments
5 min read
Hermes Agent Memory System: How Persistent AI Memory Actually Works

Hermes Agent Memory System: How Persistent AI Memory Actually Works

2
Comments
15 min read
Ollama in Docker Compose with GPU and Persistent Model Storage

Ollama in Docker Compose with GPU and Persistent Model Storage

1
Comments
10 min read
The Routing Pattern: How Smart Teams Actually Use Fast vs Capable Models

The Routing Pattern: How Smart Teams Actually Use Fast vs Capable Models

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.