DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Claude Code Now Runs Subagents in the Background by Default — What Actually Changed

Claude Code Now Runs Subagents in the Background by Default — What Actually Changed

1
Comments
3 min read
Changing the AI engine moved 3 of 10 results. Changing the question moved 10 of 10.

Changing the AI engine moved 3 of 10 results. Changing the question moved 10 of 10.

Comments
7 min read
Qwen3 235B Tops the Agentic Benchmark - What That Test Actually Measures

Qwen3 235B Tops the Agentic Benchmark - What That Test Actually Measures

Comments
2 min read
Buyer context moved 6 of the top 10 results on one market. I ran it again on an unrelated market and it moved 4.

Buyer context moved 6 of the top 10 results on one market. I ran it again on an unrelated market and it moved 4.

Comments
7 min read
Your Sandbox Has a Hole in It, and the AI Agent Found It

Your Sandbox Has a Hole in It, and the AI Agent Found It

Comments
4 min read
You asked how I know what the model did. Here are the answers, with numbers.

You asked how I know what the model did. Here are the answers, with numbers.

Comments
6 min read
Human Oversight of AI Agents Failed 33% of the Time in Testing

Human Oversight of AI Agents Failed 33% of the Time in Testing

Comments
2 min read
Why AI Agents Say “Done” When the Task Actually Failed

Why AI Agents Say “Done” When the Task Actually Failed

6
Comments
2 min read
Can a Mac mini run Kimi K3? I did the memory math

Can a Mac mini run Kimi K3? I did the memory math

3
Comments
3 min read
Teaching an Audio Model More About Barbados

Teaching an Audio Model More About Barbados

Comments
15 min read
Upgrading the judge ends one score series and starts another

Upgrading the judge ends one score series and starts another

5
Comments
9 min read
# I benchmarked AI agent memory in 2026 — and the numbers tell a different story than the marketing

# I benchmarked AI agent memory in 2026 — and the numbers tell a different story than the marketing

Comments
4 min read
What Advisory Rules Actually Do in an Agent Loop

What Advisory Rules Actually Do in an Agent Loop

Comments
7 min read
How Claude Marks AI-Generated Content?

How Claude Marks AI-Generated Content?

Comments 1
9 min read
Stop Building AI Wrappers. Start Building AI Systems

Stop Building AI Wrappers. Start Building AI Systems

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.