DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
The End of Undetectable AI Text? Claude’s New Watermark Explained

Separating AI provenance from detection myths

The End of Undetectable AI Text? Claude’s New Watermark Explained

159
Comments 105
6 min read
Direct LLM API Failure Boundaries: SaaS Token Cost and Fallback

Direct LLM API Failure Boundaries: SaaS Token Cost and Fallback

Comments
6 min read
Most AI Second Opinions Are Fake. I Built a Two-LLM Review Engine to Prove It.

Tested on 70 real PRs for just $0.53 total cost

Most AI Second Opinions Are Fake. I Built a Two-LLM Review Engine to Prove It.

17
Comments 12
11 min read
What a Malicious Ollama Model Can Actually Do to Your Host, and How to Sandbox /api/pull

What a Malicious Ollama Model Can Actually Do to Your Host, and How to Sandbox /api/pull

Comments
13 min read
Your LLM cannot compute a birth chart (and it won't tell you)

Your LLM cannot compute a birth chart (and it won't tell you)

1
Comments
4 min read
Changing the AI engine moved 3 of 10 results. Changing the question moved 10 of 10.

Changing the AI engine moved 3 of 10 results. Changing the question moved 10 of 10.

Comments
7 min read
Qwen3 235B Tops the Agentic Benchmark - What That Test Actually Measures

Qwen3 235B Tops the Agentic Benchmark - What That Test Actually Measures

Comments
2 min read
Buyer context moved 6 of the top 10 results on one market. I ran it again on an unrelated market and it moved 4.

Buyer context moved 6 of the top 10 results on one market. I ran it again on an unrelated market and it moved 4.

Comments
7 min read
Your models agreed with each other. They were agreeing with themselves.

Your models agreed with each other. They were agreeing with themselves.

Comments
13 min read
Your Sandbox Has a Hole in It, and the AI Agent Found It

Your Sandbox Has a Hole in It, and the AI Agent Found It

Comments
4 min read
You asked how I know what the model did. Here are the answers, with numbers.

You asked how I know what the model did. Here are the answers, with numbers.

Comments
6 min read
Human Oversight of AI Agents Failed 33% of the Time in Testing

Human Oversight of AI Agents Failed 33% of the Time in Testing

Comments
2 min read
Why agent loops need a coat, not another harness framework to conform to

Why agent loops need a coat, not another harness framework to conform to

1
Comments
5 min read
Why AI Agents Say “Done” When the Task Actually Failed

Why AI Agents Say “Done” When the Task Actually Failed

6
Comments
2 min read
Can a Mac mini run Kimi K3? I did the memory math

Can a Mac mini run Kimi K3? I did the memory math

3
Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.