Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
← All Trends
Rigorous Measurement Protocols for AI Agent Scoring
12 posts in this trend in the last 7 days
•
Active 34 minutes ago
Sign the Metric Before You Publish an Agent Score
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 20
Sign the Metric Before You Publish an Agent Score
#
ai
#
testing
#
python
#
agents
1
reaction
Comments
1
comment
6 min read
Hash the Task Pack Before Ranking Coding Agents
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 23
Hash the Task Pack Before Ranking Coding Agents
#
ai
#
testing
#
python
#
productivity
Comments
Add Comment
6 min read
Freeze the Manifest Before the Agent Leaderboard
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 17
Freeze the Manifest Before the Agent Leaderboard
#
ai
#
testing
#
python
#
opensource
Comments
Add Comment
8 min read
Calibrate the Judge Before the Agent Score
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 19
Calibrate the Judge Before the Agent Score
#
ai
#
testing
#
python
#
programming
Comments
Add Comment
7 min read
Agent Scores Without a Null Pack Are Marketing
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 18
Agent Scores Without a Null Pack Are Marketing
#
ai
#
testing
#
python
#
productivity
Comments
Add Comment
6 min read
Control Deltas Turn Agent Scores Into Evidence
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 21
Control Deltas Turn Agent Scores Into Evidence
#
ai
#
python
#
testing
#
productivity
Comments
Add Comment
6 min read
Freeze a Holdout Before You Quote a Coding-Agent Score
Casey Zhang
Casey Zhang
Casey Zhang
Follow
Sep 17
Freeze a Holdout Before You Quote a Coding-Agent Score
#
ai
#
python
#
testing
#
programming
Comments
Add Comment
9 min read
Replay Fixtures Separate Agent Scores From API Weather
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 22
Replay Fixtures Separate Agent Scores From API Weather
#
ai
#
testing
#
python
#
performance
Comments
Add Comment
7 min read
Don't Average Pass Rates Across Unequal Token Budgets
Casey Zhang
Casey Zhang
Casey Zhang
Follow
Sep 19
Don't Average Pass Rates Across Unequal Token Budgets
#
ai
#
testing
#
python
#
productivity
1
reaction
Comments
Add Comment
7 min read
Score Agent Patches on a Frozen Surface. Ledger the Flakes.
Finley Zhou
Finley Zhou
Finley Zhou
Follow
Sep 21
Score Agent Patches on a Frozen Surface. Ledger the Flakes.
#
testing
#
python
#
ai
#
devops
Comments
1
comment
8 min read
I Counted Drops as Wrongs. The Chart Was Theater.
Jordan Liu
Jordan Liu
Jordan Liu
Follow
Sep 21
I Counted Drops as Wrongs. The Chart Was Theater.
#
ai
#
python
#
testing
#
debugging
Comments
Add Comment
7 min read
If the Patch Authored the Test, Score the Overlap
Finley Zhou
Finley Zhou
Finley Zhou
Follow
Sep 17
If the Patch Authored the Test, Score the Overlap
#
testing
#
python
#
ai
#
devops
Comments
1
comment
6 min read
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account