DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
An AI on our team faked a tool result. Here's the detector we shipped.

An AI on our team faked a tool result. Here's the detector we shipped.

Comments 13
8 min read
GLM 5.2 isn't free: not even my US$4,000 Spark can run it

GLM 5.2 isn't free: not even my US$4,000 Spark can run it

3
Comments
6 min read
Your AGENTS.md is valid. Your agent still breaks the rules.

Your AGENTS.md is valid. Your agent still breaks the rules.

Comments 3
6 min read
CacheWeaver Reorders RAG Evidence for Prefix-Cache Reuse: Prefix-Cache-Aware Evidence Reordering

CacheWeaver Reorders RAG Evidence for Prefix-Cache Reuse: Prefix-Cache-Aware Evidence Reordering

1
Comments 3
7 min read
I Built an AI Agent That Gets Curious On Its Own

I Built an AI Agent That Gets Curious On Its Own

3
Comments 6
4 min read
My AI Wrote Code That Passed Every Test and Was Still Wrong

My AI Wrote Code That Passed Every Test and Was Still Wrong

Comments 2
4 min read
I tested 5 LLMs for prompt-injection leaks. Same code, 0% to 90%.

I tested 5 LLMs for prompt-injection leaks. Same code, 0% to 90%.

Comments
3 min read
I don't trust the LLM to classify my email. So I don't let it.

LLM as feature scorer, code as decider

I don't trust the LLM to classify my email. So I don't let it.

15
Comments 25
5 min read
The stale context problem: why your AI doesn't know what time it is

The stale context problem: why your AI doesn't know what time it is

2
Comments 2
3 min read
Has anyone else seen prompt caching break because of UUIDs/timestamps near the front?

Has anyone else seen prompt caching break because of UUIDs/timestamps near the front?

1
Comments 2
1 min read
I Installed 50 AI Agent Skills Blindly. Here's the 3-Question Filter I Use Now.

I Installed 50 AI Agent Skills Blindly. Here's the 3-Question Filter I Use Now.

Comments
3 min read
One Agent or Five? What I Learned Running a Team of AI Coders

One Agent or Five? What I Learned Running a Team of AI Coders

Comments
4 min read
Serving cheap when two models agree: a measured cost lever

Serving cheap when two models agree: a measured cost lever

3
Comments 2
5 min read
My AI Agent Read 56 KB to Answer One Question. I Made It Stop.

My AI Agent Read 56 KB to Answer One Question. I Made It Stop.

Comments
4 min read
CAPE - Collaborative Agents Prompt Engineering

CAPE - Collaborative Agents Prompt Engineering

4
Comments
10 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.