DEV Community

#evaluation

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Stop Leaving Findings in the Judge: The Ratchet That Turns Opinions Into Gates

Stop Leaving Findings in the Judge: The Ratchet That Turns Opinions Into Gates

Comments
4 min read
The Cold-Start Problem for Agent Evals: What to Gate on Day One With Zero Labeled Data

The Cold-Start Problem for Agent Evals: What to Gate on Day One With Zero Labeled Data

1
Comments 2
4 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

1
Comments
6 min read
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs

PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs

Comments
3 min read
LLM-as-Judge Shouldn't Aggregate Scores: Binary Checks as Evidence, One Holistic Verdict

LLM-as-Judge Shouldn't Aggregate Scores: Binary Checks as Evidence, One Holistic Verdict

Comments
12 min read
The Evaluation Debt You Don't Know You Have: Why Agent Evals Fail in Production

The Evaluation Debt You Don't Know You Have: Why Agent Evals Fail in Production

1
Comments 2
7 min read
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Comments
7 min read
An LLM judge is a biased instrument, not a measurement

An LLM judge is a biased instrument, not a measurement

1
Comments 1
6 min read
Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One

Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One

2
Comments 3
5 min read
Evaluating LLM Apps in Java

Evaluating LLM Apps in Java

Comments
10 min read
Evaluating LLM Apps in Python

Evaluating LLM Apps in Python

Comments
9 min read
Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference

Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference

1
Comments
5 min read
Your AI judge might be reliable — and still be wrong

Your AI judge might be reliable — and still be wrong

Comments
3 min read
Reliable, and still wrong

Reliable, and still wrong

Comments
3 min read
Give Your Agent a Type Signature: Contract-First Output Beats a Smarter Judge

Give Your Agent a Type Signature: Contract-First Output Beats a Smarter Judge

1
Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.