DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
I Built an Agent That Reads the News For Me, So I Stopped Opening Five Apps Every Morning

I Built an Agent That Reads the News For Me, So I Stopped Opening Five Apps Every Morning

Comments
3 min read
"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window

"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window

Comments
10 min read
Why Open-Weight Models Are Closing the Gap with Closed Models

Why Open-Weight Models Are Closing the Gap with Closed Models

Comments
2 min read
Every Agent Platform Needs a Front Door: A Production YARP Gateway with Runtime Config

Every Agent Platform Needs a Front Door: A Production YARP Gateway with Runtime Config

Comments
15 min read
The Vanished Constraint: Debugging a Coding Agent That Forgot Its Own Decision

The Vanished Constraint: Debugging a Coding Agent That Forgot Its Own Decision

Comments 1
4 min read
An AI Agent Recommended a Malware Package. A Human Caught It. Next Time You Might Not Be So Lucky

An AI Agent Recommended a Malware Package. A Human Caught It. Next Time You Might Not Be So Lucky

1
Comments
6 min read
What will this actually cost per month? A method for pricing an LLM workload before you commit

What will this actually cost per month? A method for pricing an LLM workload before you commit

1
Comments
6 min read
Why Your AI Coding Agent Should Never See Your .env

Why Your AI Coding Agent Should Never See Your .env

Comments
2 min read
I'm using GPT Sol and Claude Opus for free — pi-coding-agent + OmniRoute

I'm using GPT Sol and Claude Opus for free — pi-coding-agent + OmniRoute

1
Comments
3 min read
The Day a Support Ticket Overrode My System Prompt

The Day a Support Ticket Overrode My System Prompt

1
Comments
5 min read
7 Lessons from Building Agentic AI in Production

7 Lessons from Building Agentic AI in Production

Comments
2 min read
Routing by task difficulty: the numbers that changed how our AI company spends on models

Routing by task difficulty: the numbers that changed how our AI company spends on models

Comments
5 min read
Free AI Servers Fail Quietly: A Reproducible Test for MonkeyCode's Free Tier

Free AI Servers Fail Quietly: A Reproducible Test for MonkeyCode's Free Tier

Comments
5 min read
The Eval Passed and Production Still Broke: A Token-Counting Retrospective

The Eval Passed and Production Still Broke: A Token-Counting Retrospective

Comments
4 min read
The Free-Tier Trap: A Framework for Choosing Model Access

The Free-Tier Trap: A Framework for Choosing Model Access

Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.