DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Self-hosting a lite agent backend on one TPU: Gemma 4 E2B + vLLM on a v5e-1

Self-hosting a lite agent backend on one TPU: Gemma 4 E2B + vLLM on a v5e-1

16
Comments 1
21 min read
Beyond The Single-Agent Ceiling: Scale Out With MCP Agent Teams

Beyond The Single-Agent Ceiling: Scale Out With MCP Agent Teams

2
Comments
17 min read
What Anthropic’s J-lens teaches us about debugging large language models

What Anthropic’s J-lens teaches us about debugging large language models

Comments
5 min read
RAG vs. Direct Context: I Tested Both on Real Documents, Here's What Broke

RAG vs. Direct Context: I Tested Both on Real Documents, Here's What Broke

11
Comments 1
5 min read
The LLM Waterfall Pattern: Never Let a Rate Limit Kill Your Workflow

The LLM Waterfall Pattern: Never Let a Rate Limit Kill Your Workflow

Comments
5 min read
What LLM Cost Calculators Get Wrong

What LLM Cost Calculators Get Wrong

3
Comments 11
8 min read
Stop Treating LLM Prompts Like Magic Spells

Stop Treating LLM Prompts Like Magic Spells

Comments 1
3 min read
A package.lock for the prompts hiding in your codebase

A package.lock for the prompts hiding in your codebase

6
Comments
5 min read
Claude Now Puts an Invisible Watermark on Everything It Writes - Including Your Code

Claude Now Puts an Invisible Watermark on Everything It Writes - Including Your Code

2
Comments
1 min read
Unlocking Atomic AI: How SKILL.md and the 5,776-Module Skill Registry are Engineering the Future of Developer Workflows

Unlocking Atomic AI: How SKILL.md and the 5,776-Module Skill Registry are Engineering the Future of Developer Workflows

Comments
4 min read
# Semantic Caching in Enterprise RAG: Production Architectures for Faster, Lower-Cost LLM Systems

# Semantic Caching in Enterprise RAG: Production Architectures for Faster, Lower-Cost LLM Systems

Comments
12 min read
AI Agents Need Runtime State Checks, Not Just Better Prompts

AI Agents Need Runtime State Checks, Not Just Better Prompts

Comments
4 min read
Are You Benchmarking the Model—or the Harness?

Are You Benchmarking the Model—or the Harness?

2
Comments 2
8 min read
Compact Design: a JSON language models can write and Figma can import

Compact Design: a JSON language models can write and Figma can import

Comments 1
3 min read
One Leg Can Raise an Objection. It Can't Settle One.

One Leg Can Raise an Objection. It Can't Settle One.

Comments 10
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.