Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Stop Leaving Findings in the Judge: The Ratchet That Turns Opinions Into Gates
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 24
Stop Leaving Findings in the Judge: The Ratchet That Turns Opinions Into Gates
#
ai
#
agents
#
evaluation
#
observability
Comments
Add Comment
4 min read
The Cold-Start Problem for Agent Evals: What to Gate on Day One With Zero Labeled Data
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 22
The Cold-Start Problem for Agent Evals: What to Gate on Day One With Zero Labeled Data
#
ai
#
agents
#
evaluation
#
typescript
1
 reaction
Comments
2
 comments
4 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
Maya Andersson
Maya Andersson
Maya Andersson
Follow
Jul 21
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
#
statistics
#
machinelearning
#
datascience
#
evaluation
1
 reaction
Comments
Add Comment
6 min read
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 15
PoPE: Placebo-Controlled Evaluation Challenges Error-Conditioned Self-Repair in Small Code LLMs
#
llms
#
codegeneration
#
selfrepair
#
evaluation
Comments
Add Comment
3 min read
LLM-as-Judge Shouldn't Aggregate Scores: Binary Checks as Evidence, One Holistic Verdict
Tatsuya Shimomoto
Tatsuya Shimomoto
Tatsuya Shimomoto
Follow
Jul 14
LLM-as-Judge Shouldn't Aggregate Scores: Binary Checks as Evidence, One Holistic Verdict
#
llm
#
promptengineering
#
evaluation
#
claudecode
Comments
Add Comment
12 min read
The Evaluation Debt You Don't Know You Have: Why Agent Evals Fail in Production
Paul Twist
Paul Twist
Paul Twist
Follow
Jul 13
The Evaluation Debt You Don't Know You Have: Why Agent Evals Fail in Production
#
agents
#
ai
#
evaluation
#
infrastructure
1
 reaction
Comments
2
 comments
7 min read
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.
Muhammed Rasin O M
Muhammed Rasin O M
Muhammed Rasin O M
Follow
Jul 10
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.
#
evaluation
#
dataagents
#
benchmarks
#
syntheticdata
Comments
Add Comment
7 min read
An LLM judge is a biased instrument, not a measurement
Maya Andersson
Maya Andersson
Maya Andersson
Follow
Jul 22
An LLM judge is a biased instrument, not a measurement
#
llm
#
evaluation
#
statistics
#
ai
1
 reaction
Comments
1
 comment
6 min read
Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 19
Stop Judging Every Run: Eval Sampling Is a Budget Decision, Not a Coverage One
#
ai
#
agents
#
evaluation
#
observability
2
 reactions
Comments
3
 comments
5 min read
Evaluating LLM Apps in Java
Puneet Gupta
Puneet Gupta
Puneet Gupta
Follow
Jul 5
Evaluating LLM Apps in Java
#
java
#
ai
#
llm
#
evaluation
Comments
Add Comment
10 min read
Evaluating LLM Apps in Python
Puneet Gupta
Puneet Gupta
Puneet Gupta
Follow
Jul 5
Evaluating LLM Apps in Python
#
python
#
ai
#
llm
#
evaluation
Comments
Add Comment
9 min read
Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jul 2
Short-Circuit Your Agent Evals: Tier Order Is a Latency Budget, Not a Preference
#
ai
#
agents
#
evaluation
#
typescript
1
 reaction
Comments
Add Comment
5 min read
Your AI judge might be reliable — and still be wrong
Breach Protocol
Breach Protocol
Breach Protocol
Follow
Jul 1
Your AI judge might be reliable — and still be wrong
#
evaluation
#
llmjudges
#
rlhf
#
methodology
Comments
Add Comment
3 min read
Reliable, and still wrong
Breach Protocol
Breach Protocol
Breach Protocol
Follow
Jul 1
Reliable, and still wrong
#
evaluation
#
llmasjudge
#
benchmarks
Comments
Add Comment
3 min read
Give Your Agent a Type Signature: Contract-First Output Beats a Smarter Judge
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Jun 29
Give Your Agent a Type Signature: Contract-First Output Beats a Smarter Judge
#
ai
#
agents
#
evaluation
#
typescript
1
 reaction
Comments
Add Comment
4 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account