DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
AI Agents Have Hands. Here Is the Off Switch.

AI Agents Have Hands. Here Is the Off Switch.

Comments
5 min read
The golden set stopped catching regressions the day traffic changed

The golden set stopped catching regressions the day traffic changed

2
Comments 4
6 min read
Model Buzz Roundup: Week of June 3, 2026

Model Buzz Roundup: Week of June 3, 2026

Comments
10 min read
How I Cut My LLM API Bill by 80% With a Simple Router

How I Cut My LLM API Bill by 80% With a Simple Router

Comments 5
3 min read
Our AI agents fabricated "done" five times in 17 days. Here is what actually reduced it.

External validation beats prompt rules

Our AI agents fabricated "done" five times in 17 days. Here is what actually reduced it.

8
Comments 53
6 min read
I built a Rust entropy monitor to route LLM inference — here's what the benchmark showed

I built a Rust entropy monitor to route LLM inference — here's what the benchmark showed

3
Comments 1
2 min read
ChatGPT's Biggest Upgrade Ever: What Developers Actually Need to Know [June 2026]

ChatGPT's Biggest Upgrade Ever: What Developers Actually Need to Know [June 2026]

Comments
8 min read
rag-explained-how-it-works

rag-explained-how-it-works

Comments
5 min read
agents-concepts-principles-patterns

agents-concepts-principles-patterns

Comments
5 min read
Validate your Pydantic schema before the LLM call, not after.

Validate your Pydantic schema before the LLM call, not after.

Comments
1 min read
How much does Claude Code actually cost per session? I did the math

How much does Claude Code actually cost per session? I did the math

1
Comments
4 min read
I Benchmarked 6 Prompting Strategies on Two Models. The Winner Changes Depending on Which Model You Ask.

I Benchmarked 6 Prompting Strategies on Two Models. The Winner Changes Depending on Which Model You Ask.

3
Comments 2
4 min read
LLM Wiki: A Smarter Alternative to RAG

LLM Wiki: A Smarter Alternative to RAG

Comments
4 min read
Use Claude long enough and you'll end up with Karpathy's LLM Wiki without doing much.

Use Claude long enough and you'll end up with Karpathy's LLM Wiki without doing much.

Comments
5 min read
AutoLab Benchmarks Frontier Agents on Long-Horizon R&D Tasks: Iterative Experiment-Loop Evaluation

AutoLab Benchmarks Frontier Agents on Long-Horizon R&D Tasks: Iterative Experiment-Loop Evaluation

Comments
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.