DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Sonnet hallucinated. My agent stored it as fact.

Sonnet hallucinated. My agent stored it as fact.

3
Comments 45
3 min read
TurboQuant on a MacBook Pro, part 2: perplexity, KL divergence, and asymmetric K/V on M5 Max

TurboQuant on a MacBook Pro, part 2: perplexity, KL divergence, and asymmetric K/V on M5 Max

Comments
8 min read
llama.cpp b9455 Finally Caught vLLM: 70t/s on 2x3090 Qwen 27B UQ8

llama.cpp b9455 Finally Caught vLLM: 70t/s on 2x3090 Qwen 27B UQ8

Comments 1
3 min read
Why I'm Building a Local-First AI Coding Workspace (And How Behavioral Routing Makes It Work)

Why I'm Building a Local-First AI Coding Workspace (And How Behavioral Routing Makes It Work)

Comments
6 min read
We Had LLMs Hallucinating Legal URLs in Production — Here's What We Tried

We Had LLMs Hallucinating Legal URLs in Production — Here's What We Tried

2
Comments 2
4 min read
hat Makes a Good SFT Sample (And Why Most Synthetic Datasets Get It Wrong)

hat Makes a Good SFT Sample (And Why Most Synthetic Datasets Get It Wrong)

Comments 2
4 min read
Prompt Caching Works. Your Prompt Assembly Code Does Not.

Prompt Caching Works. Your Prompt Assembly Code Does Not.

Comments
4 min read
Opus 4.7 vs GLM 5.1: is mixing models worth it?

Opus 4.7 vs GLM 5.1: is mixing models worth it?

Comments
13 min read
How to track LLM costs per customer in production

How to track LLM costs per customer in production

17
Comments 2
8 min read
Upgrading Kiwi-chan’s Brain: Pushing a 30GB "Frankenstein" GPU Rig to the Limit with Qwen 3.6-35B-A3B

Upgrading Kiwi-chan’s Brain: Pushing a 30GB "Frankenstein" GPU Rig to the Limit with Qwen 3.6-35B-A3B

Comments
4 min read
Mistral Medium 3.5 GGUF, FlashQLA Boost for Qwen, & Ollama Playground

Mistral Medium 3.5 GGUF, FlashQLA Boost for Qwen, & Ollama Playground

Comments
3 min read
When the Reranker Hurts: Recall@5 Cases Where Two-Stage Retrieval Loses to One

When the Reranker Hurts: Recall@5 Cases Where Two-Stage Retrieval Loses to One

Comments
7 min read
Why Strict JSON Mode Doesn't Stop Hallucinated Tool Calls

Why Strict JSON Mode Doesn't Stop Hallucinated Tool Calls

Comments
7 min read
Every LLM Eval Library Has the Same Bug: Stochastic Judges Used as Deterministic Oracles

Every LLM Eval Library Has the Same Bug: Stochastic Judges Used as Deterministic Oracles

Comments
7 min read
Local AI Accessibility, JetBrains’ 2026 IDE Plans, and Agentic Architecture Pitfalls

Local AI Accessibility, JetBrains’ 2026 IDE Plans, and Agentic Architecture Pitfalls

Comments
2 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.