Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
← All Trends
Rigorous Coding-Agent Benchmarking & Scoring
13 posts in this trend in the last 7 days
•
Active 23 minutes ago
Freeze the Manifest Before the Agent Leaderboard
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 17
Freeze the Manifest Before the Agent Leaderboard
#
ai
#
testing
#
python
#
opensource
Comments
Add Comment
8 min read
Calibrate the Judge Before the Agent Score
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 19
Calibrate the Judge Before the Agent Score
#
ai
#
testing
#
python
#
programming
Comments
Add Comment
7 min read
Agent Scores Without a Null Pack Are Marketing
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 18
Agent Scores Without a Null Pack Are Marketing
#
ai
#
testing
#
python
#
productivity
Comments
Add Comment
6 min read
Sign the Metric Before You Publish an Agent Score
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 20
Sign the Metric Before You Publish an Agent Score
#
ai
#
testing
#
python
#
agents
1
reaction
Comments
1
comment
6 min read
Freeze a Holdout Before You Quote a Coding-Agent Score
Casey Zhang
Casey Zhang
Casey Zhang
Follow
Sep 17
Freeze a Holdout Before You Quote a Coding-Agent Score
#
ai
#
python
#
testing
#
programming
Comments
Add Comment
9 min read
Control Deltas Turn Agent Scores Into Evidence
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 21
Control Deltas Turn Agent Scores Into Evidence
#
ai
#
python
#
testing
#
productivity
Comments
Add Comment
6 min read
Stratify the Task Pack Before Averaging Agent Scores
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 16
Stratify the Task Pack Before Averaging Agent Scores
#
ai
#
python
#
testing
#
agents
1
reaction
Comments
Add Comment
7 min read
Don't Average Pass Rates Across Unequal Token Budgets
Casey Zhang
Casey Zhang
Casey Zhang
Follow
Sep 19
Don't Average Pass Rates Across Unequal Token Budgets
#
ai
#
testing
#
python
#
productivity
1
reaction
Comments
Add Comment
7 min read
Label Infra Failures Before You Rank Coding Agents
Casey Zhang
Casey Zhang
Casey Zhang
Follow
Sep 16
Label Infra Failures Before You Rank Coding Agents
#
ai
#
testing
#
python
#
productivity
Comments
Add Comment
7 min read
Score Agent Patches on a Frozen Surface. Ledger the Flakes.
Finley Zhou
Finley Zhou
Finley Zhou
Follow
Sep 21
Score Agent Patches on a Frozen Surface. Ledger the Flakes.
#
testing
#
python
#
ai
#
devops
Comments
1
comment
8 min read
I Counted Drops as Wrongs. The Chart Was Theater.
Jordan Liu
Jordan Liu
Jordan Liu
Follow
Sep 21
I Counted Drops as Wrongs. The Chart Was Theater.
#
ai
#
python
#
testing
#
debugging
Comments
Add Comment
7 min read
Treat the Grader as Code, Not a Hidden Prompt
Dakota Ma
Dakota Ma
Dakota Ma
Follow
Sep 16
Treat the Grader as Code, Not a Hidden Prompt
#
ai
#
python
#
testing
#
opensource
Comments
Add Comment
6 min read
If the Patch Authored the Test, Score the Overlap
Finley Zhou
Finley Zhou
Finley Zhou
Follow
Sep 17
If the Patch Authored the Test, Score the Overlap
#
testing
#
python
#
ai
#
devops
Comments
1
comment
6 min read
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account