DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Retrieval-Augmented Self-Recall: The RAG Problem Nobody Talks About

Why similarity thresholds fail across models

Retrieval-Augmented Self-Recall: The RAG Problem Nobody Talks About

9
Comments 14
5 min read
VoxCPM2 TTS, AI Cost Optimization, and HF Hub CLI for Open Models

VoxCPM2 TTS, AI Cost Optimization, and HF Hub CLI for Open Models

Comments
4 min read
Fable 5 refusals are 200 OK — your error handler misses them

Fable 5 refusals are 200 OK — your error handler misses them

1
Comments
7 min read
Building an AI Agent That Knows When Not to Guess (Qwen + MCP)

Uncertainty as a first-class output

Building an AI Agent That Knows When Not to Guess (Qwen + MCP)

29
Comments 30
3 min read
Stateful provider fallback for LLM pipelines: an FSM pattern

Stateful provider fallback for LLM pipelines: an FSM pattern

6
Comments 2
3 min read
How I caught the voice-agent failures my eval dashboard kept missing

How I caught the voice-agent failures my eval dashboard kept missing

1
Comments 1
5 min read
Your AI agent is a stack of files. muster 1.0.0 tests all of them.

Your AI agent is a stack of files. muster 1.0.0 tests all of them.

Comments
4 min read
Block the Merge if the Model Isn't Ready": Shifting Local AI Evaluations Left with CI Gates

Block the Merge if the Model Isn't Ready": Shifting Local AI Evaluations Left with CI Gates

Comments
1 min read
Claude AI by Anthropic: Key Features That Set This Model Apart in 2024 [EN]

Claude AI by Anthropic: Key Features That Set This Model Apart in 2024 [EN]

Comments
4 min read
Fable 5 Banned: What Happens When Your AI Governance Lives Inside the Model

Fable 5 Banned: What Happens When Your AI Governance Lives Inside the Model

Comments 1
7 min read
Stop hand-picking an LLM per request: a practical case for auto-routing

Stop hand-picking an LLM per request: a practical case for auto-routing

Comments
3 min read
One base_url for GPT, Claude, and Gemini: cutting three SDKs down to one

One base_url for GPT, Claude, and Gemini: cutting three SDKs down to one

Comments
3 min read
Build the Reply Loop: Receive, Think, Respond

Build the Reply Loop: Receive, Think, Respond

Comments
5 min read
Stop Blaming the Model. Your Latency Budget Is Probably Broken.

Stop Blaming the Model. Your Latency Budget Is Probably Broken.

Comments
3 min read
Claude Is Your Insider Threat Now - Notes from Dan Tentler's Security Fest 2026 Talk

Claude Is Your Insider Threat Now - Notes from Dan Tentler's Security Fest 2026 Talk

1
Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.