DEV Community

#evaluation

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement

Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement

1
Comments 1
6 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Comments
7 min read
Your Agent's Context Window Overflowed and It Answered Anyway

Your Agent's Context Window Overflowed and It Answered Anyway

2
Comments
4 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

1
Comments 2
7 min read
How EvalPort's Grader System Works: 11 Types for LLM Evaluation

How EvalPort's Grader System Works: 11 Types for LLM Evaluation

Comments
2 min read
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

Comments 2
7 min read
OpenEval: Why LLM Evaluation Needs a Standard Format

OpenEval: Why LLM Evaluation Needs a Standard Format

Comments
1 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

1
Comments
6 min read
Your Golden Dataset Is Rotting: The Eval Oracle Nobody Re-Validates

Your Golden Dataset Is Rotting: The Eval Oracle Nobody Re-Validates

2
Comments 1
5 min read
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs

PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs

Comments
3 min read
Your reasoning model isn't dumb. Your parser is throwing away its best answers.

Your reasoning model isn't dumb. Your parser is throwing away its best answers.

1
Comments 2
4 min read
I Built an Agent Evaluation Harness for Local AI — What Most People Get Wrong

I Built an Agent Evaluation Harness for Local AI — What Most People Get Wrong

Comments 3
13 min read
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Comments
7 min read
Evaluating LLM Apps in Python

Evaluating LLM Apps in Python

Comments
9 min read
Evaluating LLM Apps in Java

Evaluating LLM Apps in Java

Comments
10 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.