DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Trying "DFlash," a Diffusion-Model Approach to Parallel Draft-Token Generation, on Gemma

Trying "DFlash," a Diffusion-Model Approach to Parallel Draft-Token Generation, on Gemma

Comments
17 min read
How many tools should an MCP server have?

How many tools should an MCP server have?

Comments
5 min read
AI Agent Testing: Why a 77% Pass Rate Can Mean 53% in Production

AI Agent Testing: Why a 77% Pass Rate Can Mean 53% in Production

Comments
7 min read
I built a sub-4ms JIT CSS framework with native Glassmorphism and built-in LLM training files

I built a sub-4ms JIT CSS framework with native Glassmorphism and built-in LLM training files

Comments
4 min read
Just Train More: Measuring the Exchange Rate

Just Train More: Measuring the Exchange Rate

Comments
5 min read
Sampling Parameters as Substances: An Honest Analogy

Sampling Parameters as Substances: An Honest Analogy

Comments
2 min read
Ollama's Responses API accepts previous_response_id, returns 200, and forgets the whole conversation

Ollama's Responses API accepts previous_response_id, returns 200, and forgets the whole conversation

1
Comments 3
6 min read
Best LLM for Coding in 2026: Claude Opus 4.8 vs GPT-5.5 vs Gemini 3.1 Pro (With Enterprise Governance Guide)

Best LLM for Coding in 2026: Claude Opus 4.8 vs GPT-5.5 vs Gemini 3.1 Pro (With Enterprise Governance Guide)

Comments
7 min read
The Model Is Not the Product. Here's What Actually Is.

The Model Is Not the Product. Here's What Actually Is.

Comments
4 min read
My Extraction Score Was 0.08 and the Model Was Innocent: Rebuilding the Ruler

My Extraction Score Was 0.08 and the Model Was Innocent: Rebuilding the Ruler

8
Comments 1
8 min read
OpenRouter Fusion: escalate hard prompts to a panel; keep the policy in your repo

OpenRouter Fusion: escalate hard prompts to a panel; keep the policy in your repo

1
Comments
6 min read
From Claude Project to Hybrid AI Agent: Lessons from a Real-World Content Workflow

From Claude Project to Hybrid AI Agent: Lessons from a Real-World Content Workflow

Comments
6 min read
Why Local LLMs Don't Need C++ or Python: Building a 15MB Native AOT Inference Engine in .NET 10

Why Local LLMs Don't Need C++ or Python: Building a 15MB Native AOT Inference Engine in .NET 10

1
Comments 5
5 min read
I pay an LLM to approve bad reviews

I pay an LLM to approve bad reviews

Comments
5 min read
Your LLM cost estimate is wrong above 200,000 tokens

Your LLM cost estimate is wrong above 200,000 tokens

Comments
2 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.