DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Flash Attention: what it does and why it matters

Flash Attention: what it does and why it matters

Comments
8 min read
Which Is to Be Master? Language, Authority and LLMs

Which Is to Be Master? Language, Authority and LLMs

3
Comments 1
14 min read
My Agent Reported an Audit as Passed - 68% Self-Reported, 0 Verified Failures

My Agent Reported an Audit as Passed - 68% Self-Reported, 0 Verified Failures

Comments 7
13 min read
Gemma 4 QAT on 10GB Laptop: Local AI with 6.7GB VRAM

Gemma 4 QAT on 10GB Laptop: Local AI with 6.7GB VRAM

Comments
1 min read
Making a fleet of self-hosted LLM agents trustworthy

Making a fleet of self-hosted LLM agents trustworthy

1
Comments
6 min read
Ollama 0.30 GPU Boost: Faster local Qwen inference on NVIDIA

Ollama 0.30 GPU Boost: Faster local Qwen inference on NVIDIA

Comments
1 min read
Local LLM vs Claude: Benchmarking qwen3-coder:30b as a Production Agent Backend

Local LLM vs Claude: Benchmarking qwen3-coder:30b as a Production Agent Backend

Comments 1
5 min read
Why We Added Rate Limits Between AI Agents

Why We Added Rate Limits Between AI Agents

Comments
3 min read
Your schema validation passes and the agent still picks the wrong tool. The bug is semantic.

Your schema validation passes and the agent still picks the wrong tool. The bug is semantic.

Comments
2 min read
My Agent's Memory File Wasn't Wrong. It Was Just Six Weeks Stale.

My Agent's Memory File Wasn't Wrong. It Was Just Six Weeks Stale.

Comments
4 min read
LLM-as-Judge Shouldn't Aggregate Scores: Binary Checks as Evidence, One Holistic Verdict

LLM-as-Judge Shouldn't Aggregate Scores: Binary Checks as Evidence, One Holistic Verdict

Comments
12 min read
Building an MCP Server in Python — Architecture, FastMCP, and Production Code

Building an MCP Server in Python — Architecture, FastMCP, and Production Code

2
Comments
8 min read
The Prefill Wall: Why MTP's 2 Barely Moves Long-Context Latency (Qwen3.6-27B, RTX 3090)

The Prefill Wall: Why MTP's 2 Barely Moves Long-Context Latency (Qwen3.6-27B, RTX 3090)

Comments
4 min read
7 things I learned trying to stop LLM API bills from silently exploding

7 things I learned trying to stop LLM API bills from silently exploding

4
Comments 11
3 min read
I Built an AI Agent That Writes Tests, Finds Bugs, and Opens PRs — Autonomously

I Built an AI Agent That Writes Tests, Finds Bugs, and Opens PRs — Autonomously

Comments 1
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.