DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Prompt Caching: What Belongs in the Cacheable Prefix, What Kills Hit Rate

Prompt Caching: What Belongs in the Cacheable Prefix, What Kills Hit Rate

Comments
8 min read
HyDE, Multi-Query, Decomposition: Which Query Rewrite Actually Moves Recall?

HyDE, Multi-Query, Decomposition: Which Query Rewrite Actually Moves Recall?

Comments
10 min read
Your P99 Latency Lies. Streaming Users Feel TTFT.

Your P99 Latency Lies. Streaming Users Feel TTFT.

Comments
8 min read
Sub-Agents vs One Big Agent: 4 Signals That Decide It

Sub-Agents vs One Big Agent: 4 Signals That Decide It

Comments
9 min read
Stop Keying Agent Replay on Tool Names. Use Fingerprints.

Stop Keying Agent Replay on Tool Names. Use Fingerprints.

Comments
8 min read
Stop Using JSON Mode for Structured Output. XML Tags Win 4 of 5 Cases.

Stop Using JSON Mode for Structured Output. XML Tags Win 4 of 5 Cases.

Comments
7 min read
Eval-Driven Canary: Shipping Prompt Changes Behind a Quality Gate

Eval-Driven Canary: Shipping Prompt Changes Behind a Quality Gate

Comments
9 min read
Few-Shot Examples Are Eating Your Tokens. Here's the Cull Test.

Few-Shot Examples Are Eating Your Tokens. Here's the Cull Test.

Comments
8 min read
3 Dashboards Every LLM Team Needs: Cost, Quality, Latency, Wired End-to-End

3 Dashboards Every LLM Team Needs: Cost, Quality, Latency, Wired End-to-End

Comments
10 min read
The 3 RAG Citation Patterns: One Regulators Accept, One Users Read, One Nobody Should Ship

The 3 RAG Citation Patterns: One Regulators Accept, One Users Read, One Nobody Should Ship

Comments
10 min read
Goal Completion Verification: The Step 90% of Agents Skip

Goal Completion Verification: The Step 90% of Agents Skip

Comments
9 min read
Self-Consistency at N=5 With Sonnet Beats One Opus Call on 3 Task Types

Self-Consistency at N=5 With Sonnet Beats One Opus Call on 3 Task Types

Comments
8 min read
Diffusion Language Models: How NVIDIA Nemotron-Labs Diffusion Shatters the Autoregressive Speed Ceiling

Diffusion Language Models: How NVIDIA Nemotron-Labs Diffusion Shatters the Autoregressive Speed Ceiling

Comments
18 min read
Best AI Agent Security & Guardrails Tools in 2026: LLM Guard vs NeMo vs Guardrails AI

Best AI Agent Security & Guardrails Tools in 2026: LLM Guard vs NeMo vs Guardrails AI

1
Comments 2
3 min read
Built a Predictive Incident Response Agent with LLMs and Vector Memory

Built a Predictive Incident Response Agent with LLMs and Vector Memory

Comments
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.