DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Ollama 0.30 GPU Boost: Faster local Qwen inference on NVIDIA

Ollama 0.30 GPU Boost: Faster local Qwen inference on NVIDIA

Comments
1 min read
Making a fleet of self-hosted LLM agents trustworthy

Making a fleet of self-hosted LLM agents trustworthy

1
Comments
6 min read
Local LLM vs Claude: Benchmarking qwen3-coder:30b as a Production Agent Backend

Local LLM vs Claude: Benchmarking qwen3-coder:30b as a Production Agent Backend

Comments 1
5 min read
Why We Added Rate Limits Between AI Agents

Why We Added Rate Limits Between AI Agents

Comments
3 min read
Your schema validation passes and the agent still picks the wrong tool. The bug is semantic.

Your schema validation passes and the agent still picks the wrong tool. The bug is semantic.

Comments
2 min read
My Agent's Memory File Wasn't Wrong. It Was Just Six Weeks Stale.

My Agent's Memory File Wasn't Wrong. It Was Just Six Weeks Stale.

Comments
4 min read
LLM-as-Judge Shouldn't Aggregate Scores: Binary Checks as Evidence, One Holistic Verdict

LLM-as-Judge Shouldn't Aggregate Scores: Binary Checks as Evidence, One Holistic Verdict

Comments
12 min read
Building an MCP Server in Python — Architecture, FastMCP, and Production Code

Building an MCP Server in Python — Architecture, FastMCP, and Production Code

2
Comments
8 min read
The Prefill Wall: Why MTP's 2 Barely Moves Long-Context Latency (Qwen3.6-27B, RTX 3090)

The Prefill Wall: Why MTP's 2 Barely Moves Long-Context Latency (Qwen3.6-27B, RTX 3090)

Comments
4 min read
My Home AI's First Reply Took Four Minutes. Now It Takes Eleven Seconds.

My Home AI's First Reply Took Four Minutes. Now It Takes Eleven Seconds.

2
Comments
4 min read
7 things I learned trying to stop LLM API bills from silently exploding

7 things I learned trying to stop LLM API bills from silently exploding

4
Comments 11
3 min read
The OWASP Agentic Top 10, explained for practitioners

The OWASP Agentic Top 10, explained for practitioners

1
Comments
4 min read
How to Build AI Agents from Scratch with FastAPI

How to Build AI Agents from Scratch with FastAPI

Comments
8 min read
I Built an AI Agent That Writes Tests, Finds Bugs, and Opens PRs — Autonomously

I Built an AI Agent That Writes Tests, Finds Bugs, and Opens PRs — Autonomously

Comments 1
5 min read
Kubernetes in LLMOps (Part 2): GPU Efficiency, Cost Engineering, and Real-World Failure Modes

Kubernetes in LLMOps (Part 2): GPU Efficiency, Cost Engineering, and Real-World Failure Modes

1
Comments 1
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.