DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Which Is to Be Master? Language, Authority and LLMs

Which Is to Be Master? Language, Authority and LLMs

Comments 1
14 min read
My Agent Reported an Audit as Passed - 68% Self-Reported, 0 Verified Failures

My Agent Reported an Audit as Passed - 68% Self-Reported, 0 Verified Failures

Comments 7
13 min read
Making a fleet of self-hosted LLM agents trustworthy

Making a fleet of self-hosted LLM agents trustworthy

1
Comments
6 min read
Gemma 4 QAT on 10GB Laptop: Local AI with 6.7GB VRAM

Gemma 4 QAT on 10GB Laptop: Local AI with 6.7GB VRAM

Comments
1 min read
Ollama 0.30 GPU Boost: Faster local Qwen inference on NVIDIA

Ollama 0.30 GPU Boost: Faster local Qwen inference on NVIDIA

Comments
1 min read
Why We Added Rate Limits Between AI Agents

Why We Added Rate Limits Between AI Agents

Comments
3 min read
Local LLM vs Claude: Benchmarking qwen3-coder:30b as a Production Agent Backend

Local LLM vs Claude: Benchmarking qwen3-coder:30b as a Production Agent Backend

Comments 1
5 min read
Your schema validation passes and the agent still picks the wrong tool. The bug is semantic.

Your schema validation passes and the agent still picks the wrong tool. The bug is semantic.

Comments
2 min read
Building an MCP Server in Python — Architecture, FastMCP, and Production Code

Building an MCP Server in Python — Architecture, FastMCP, and Production Code

2
Comments
8 min read
The Prefill Wall: Why MTP's 2 Barely Moves Long-Context Latency (Qwen3.6-27B, RTX 3090)

The Prefill Wall: Why MTP's 2 Barely Moves Long-Context Latency (Qwen3.6-27B, RTX 3090)

Comments
4 min read
I Built an AI Agent That Writes Tests, Finds Bugs, and Opens PRs — Autonomously

I Built an AI Agent That Writes Tests, Finds Bugs, and Opens PRs — Autonomously

Comments 1
5 min read
Kubernetes in LLMOps (Part 2): GPU Efficiency, Cost Engineering, and Real-World Failure Modes

Kubernetes in LLMOps (Part 2): GPU Efficiency, Cost Engineering, and Real-World Failure Modes

1
Comments 1
5 min read
I was fine-tuning a language model on a new language. The loss was perfect. It spoke Chinese.

I was fine-tuning a language model on a new language. The loss was perfect. It spoke Chinese.

Comments
3 min read
Model routing by task type: the savings math, the classifier overhead, and the A/B that proves it

Model routing by task type: the savings math, the classifier overhead, and the A/B that proves it

Comments
12 min read
Context Compaction Visualizer: See Exactly What Your AI Agent Forgot Before It Costs You

Context Compaction Visualizer: See Exactly What Your AI Agent Forgot Before It Costs You

8
Comments 4
7 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.