DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
The MCP Tax Hit 42,000 Tokens on a Single Server. Here's What I Did About It.

The MCP Tax Hit 42,000 Tokens on a Single Server. Here's What I Did About It.

Comments
4 min read
Building Maestro AI: Routing LLM Calls So Your Agent Doesn't Burn Sonnet on Summaries

Building Maestro AI: Routing LLM Calls So Your Agent Doesn't Burn Sonnet on Summaries

Comments 2
7 min read
A Better LLM Judge? The Rubric Made My Small Model Worse

A Better LLM Judge? The Rubric Made My Small Model Worse

Comments
5 min read
LLM-as-a-Judge: I Built One From Scratch, Then Checked It Against Humans

LLM-as-a-Judge: I Built One From Scratch, Then Checked It Against Humans

Comments
4 min read
How Modern Transformer Blocks Work — From RMSNorm to MoE

How Modern Transformer Blocks Work — From RMSNorm to MoE

Comments
5 min read
One Agent or Five? What I Learned Running a Team of AI Coders

One Agent or Five? What I Learned Running a Team of AI Coders

Comments
4 min read
My AI Agent Read 56 KB to Answer One Question. I Made It Stop.

My AI Agent Read 56 KB to Answer One Question. I Made It Stop.

Comments
4 min read
How to switch AI models without rewriting your app

How to switch AI models without rewriting your app

Comments
3 min read
I Built Byte Because OpenWebUI Kept Breaking

I Built Byte Because OpenWebUI Kept Breaking

5
Comments
1 min read
From Transformer to ChatGPT: How One Paper Changed AI Engineering Forever

From Transformer to ChatGPT: How One Paper Changed AI Engineering Forever

1
Comments
3 min read
Stop Hardcoding Your Agent Workflows (or Don't): A Dev's Guide to Supervisor Delegation

Stop Hardcoding Your Agent Workflows (or Don't): A Dev's Guide to Supervisor Delegation

Comments
3 min read
Stop pasting your API keys into ChatGPT: a safer way to feed a codebase to an LLM

Stop pasting your API keys into ChatGPT: a safer way to feed a codebase to an LLM

Comments
2 min read
"LLM Inference Optimization: The Line Item That Decides If Your AI Ships"

"LLM Inference Optimization: The Line Item That Decides If Your AI Ships"

Comments
2 min read
Resurrecting Kepler: Getting Modern LLMs Running on a GTX 770 (Kernel 7.x)

Resurrecting Kepler: Getting Modern LLMs Running on a GTX 770 (Kernel 7.x)

1
Comments
4 min read
Caching LLM responses is just content addressing

Caching LLM responses is just content addressing

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.