Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 12
Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement
#
ai
#
evaluation
#
dotnet
#
testing
1
 reaction
Comments
1
 comment
6 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
Sanjeev Kumar
Sanjeev Kumar
Sanjeev Kumar
Follow
Aug 12
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
#
ai
#
llm
#
evaluation
Comments
Add Comment
7 min read
Your Agent's Context Window Overflowed and It Answered Anyway
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 12
Your Agent's Context Window Overflowed and It Answered Anyway
#
ai
#
agents
#
observability
#
evaluation
2
 reactions
Comments
Add Comment
4 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
Sanjeev Kumar
Sanjeev Kumar
Sanjeev Kumar
Follow
Aug 12
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
#
ai
#
llm
#
evaluation
1
 reaction
Comments
2
 comments
7 min read
How EvalPort's Grader System Works: 11 Types for LLM Evaluation
Adha AK
Adha AK
Adha AK
Follow
Aug 4
How EvalPort's Grader System Works: 11 Types for LLM Evaluation
#
llm
#
evaluation
#
testing
#
opensource
Comments
Add Comment
2 min read
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother
Xinyang Wu
Xinyang Wu
Xinyang Wu
Follow
Aug 3
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother
#
rag
#
llm
#
embeddings
#
evaluation
Comments
2
 comments
7 min read
OpenEval: Why LLM Evaluation Needs a Standard Format
Adha AK
Adha AK
Adha AK
Follow
Jul 30
OpenEval: Why LLM Evaluation Needs a Standard Format
#
llm
#
evaluation
#
ai
#
testing
Comments
Add Comment
1 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
Maya Andersson
Maya Andersson
Maya Andersson
Follow
Jul 21
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
#
statistics
#
machinelearning
#
datascience
#
evaluation
1
 reaction
Comments
Add Comment
6 min read
Your Golden Dataset Is Rotting: The Eval Oracle Nobody Re-Validates
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 9
Your Golden Dataset Is Rotting: The Eval Oracle Nobody Re-Validates
#
ai
#
agents
#
evaluation
#
observability
2
 reactions
Comments
1
 comment
5 min read
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 15
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs
#
llms
#
codegeneration
#
selfrepair
#
evaluation
Comments
Add Comment
3 min read
Your reasoning model isn't dumb. Your parser is throwing away its best answers.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Aug 7
Your reasoning model isn't dumb. Your parser is throwing away its best answers.
#
machinelearning
#
llm
#
evaluation
#
ai
1
 reaction
Comments
2
 comments
4 min read
I Built an Agent Evaluation Harness for Local AI — What Most People Get Wrong
shakti tiwari
shakti tiwari
shakti tiwari
Follow
Aug 5
I Built an Agent Evaluation Harness for Local AI — What Most People Get Wrong
#
aiagents
#
evaluation
#
localai
#
testing
Comments
3
 comments
13 min read
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.
Muhammed Rasin O M
Muhammed Rasin O M
Muhammed Rasin O M
Follow
Jul 10
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.
#
evaluation
#
dataagents
#
benchmarks
#
syntheticdata
Comments
Add Comment
7 min read
Evaluating LLM Apps in Python
Puneet Gupta
Puneet Gupta
Puneet Gupta
Follow
Jul 5
Evaluating LLM Apps in Python
#
python
#
ai
#
llm
#
evaluation
Comments
Add Comment
9 min read
Evaluating LLM Apps in Java
Puneet Gupta
Puneet Gupta
Puneet Gupta
Follow
Jul 5
Evaluating LLM Apps in Java
#
java
#
ai
#
llm
#
evaluation
Comments
Add Comment
10 min read
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account