DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
My Tool-Calling Loop Worked Fine, Until Compliance Wanted a Second Model to Check It

My Tool-Calling Loop Worked Fine, Until Compliance Wanted a Second Model to Check It

2
Comments 1
4 min read
What I learned trying to benchmark local LLMs honestly

What I learned trying to benchmark local LLMs honestly

Comments 2
4 min read
The night an uncapped prompt turned into a bill

The night an uncapped prompt turned into a bill

Comments
1 min read
Can a Cheap Model Beat a Frontier Model? Rebuilding Recursive Language Models with Codex

Can a Cheap Model Beat a Frontier Model? Rebuilding Recursive Language Models with Codex

2
Comments
6 min read
RAG vs MAG: Two Paths to Smarter AI Memory

RAG vs MAG: Two Paths to Smarter AI Memory

Comments
3 min read
llmperf Is Archived: Alternatives for LLM Benchmarking

llmperf Is Archived: Alternatives for LLM Benchmarking

Comments
3 min read
Measuring LLM Prefix Caching: The Cache Hit Rate Metric

Measuring LLM Prefix Caching: The Cache Hit Rate Metric

Comments
6 min read
LLM Narrative Engines, Part 6: Runtime Loop and Branching

LLM Narrative Engines, Part 6: Runtime Loop and Branching

Comments
10 min read
Beyond Size: The Three Pillars of Test-Time Scaling in Large Language Models

Beyond Size: The Three Pillars of Test-Time Scaling in Large Language Models

Comments
5 min read
Beyond RAG: Building an AI Coding Agent with Planning, Tool Execution, and ReAct Reasoning

Beyond RAG: Building an AI Coding Agent with Planning, Tool Execution, and ReAct Reasoning

Comments
3 min read
LLMs on Consumer Hardware — Part 1: The Stack and First Benchmarks

LLMs on Consumer Hardware — Part 1: The Stack and First Benchmarks

Comments
3 min read
PassiveDx: The Body's API

PassiveDx: The Body's API

1
Comments
8 min read
Stop Guessing: A Repeatable Harness for Comparing Free LLM Endpoints on Your Actual Tasks

Stop Guessing: A Repeatable Harness for Comparing Free LLM Endpoints on Your Actual Tasks

1
Comments
4 min read
Vector RAG can't fix long-context state tracking (33 runs, zero variance)

Vector RAG can't fix long-context state tracking (33 runs, zero variance)

1
Comments
2 min read
Treat Prompt Changes Like Schema Migrations: A Free, Diffable LLM Smoke Test

Treat Prompt Changes Like Schema Migrations: A Free, Diffable LLM Smoke Test

Comments
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.