DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Why My Local Coding Agent Could Act but Couldn't Finish

Why My Local Coding Agent Could Act but Couldn't Finish

1
Comments 2
11 min read
Why KV Cache Matters — How MQA, GQA, and MLA Make LLM Inference Faster

Why KV Cache Matters — How MQA, GQA, and MLA Make LLM Inference Faster

Comments
5 min read
Write returned success. The file was never there.

Write returned success. The file was never there.

Comments
2 min read
The 50% Context Tax: Why Your AI Agent's Million-Token Window Is Burning Money

The 50% Context Tax: Why Your AI Agent's Million-Token Window Is Burning Money

Comments
4 min read
smolagents replays its whole memory every step: the O(n ) token bill nobody mentions

smolagents replays its whole memory every step: the O(n ) token bill nobody mentions

2
Comments 3
6 min read
A local model opened 41 of our pull requests in five weeks. The model is the least interesting part.

A local model opened 41 of our pull requests in five weeks. The model is the least interesting part.

Comments
10 min read
Your First LLM API on Kubernetes: From Model to Curl Request

Your First LLM API on Kubernetes: From Model to Curl Request

Comments
10 min read
LangChain, LangGraph, LangSmith, Langflow... What's the Difference? (2026 Developer's Map)

LangChain, LangGraph, LangSmith, Langflow... What's the Difference? (2026 Developer's Map)

Comments 3
8 min read
LLM-as-a-Judge Is Too Expensive to Be the Default

LLM-as-a-Judge Is Too Expensive to Be the Default

4
Comments 6
4 min read
Prompt Caching vs Fine-Tuning: Cost-Effective LLM Strategies

Prompt Caching vs Fine-Tuning: Cost-Effective LLM Strategies

1
Comments
3 min read
Defending LLM Agents from Gradient-Based Adversarial Attacks: A VRF+LoRA Approach

Defending LLM Agents from Gradient-Based Adversarial Attacks: A VRF+LoRA Approach

Comments
6 min read
Async LLM inference in CI: stop build workers blocking on slow jobs

Async LLM inference in CI: stop build workers blocking on slow jobs

Comments
4 min read
When AI-Generated SQL Becomes Untrustworthy: How to Restore Confidence in Our Data

When AI-Generated SQL Becomes Untrustworthy: How to Restore Confidence in Our Data

5
Comments
8 min read
GLM-5.2 open agent benchmark: 22% Less Tool Failure

GLM-5.2 open agent benchmark: 22% Less Tool Failure

1
Comments
8 min read
Adversarial Comments Are Now a Vulnerability Detection Bypass Technique

Adversarial Comments Are Now a Vulnerability Detection Bypass Technique

1
Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.