Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
← All Trends
Rigorous AI Agent Evaluation & Score Locking
15 posts in this trend in the last 7 days
•
Active 31 minutes ago
Sign the Metric Before You Publish an Agent Score
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 20
Sign the Metric Before You Publish an Agent Score
#
ai
#
testing
#
python
#
agents
1
reaction
Comments
1
comment
6 min read
Freeze the Manifest Before the Agent Leaderboard
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 17
Freeze the Manifest Before the Agent Leaderboard
#
ai
#
testing
#
python
#
opensource
Comments
Add Comment
8 min read
Hash the Task Pack Before Ranking Coding Agents
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 23
Hash the Task Pack Before Ranking Coding Agents
#
ai
#
testing
#
python
#
productivity
Comments
Add Comment
6 min read
Agent Scores Without a Null Pack Are Marketing
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 18
Agent Scores Without a Null Pack Are Marketing
#
ai
#
testing
#
python
#
productivity
Comments
Add Comment
6 min read
Calibrate the Judge Before the Agent Score
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 19
Calibrate the Judge Before the Agent Score
#
ai
#
testing
#
python
#
programming
Comments
Add Comment
7 min read
Control Deltas Turn Agent Scores Into Evidence
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 21
Control Deltas Turn Agent Scores Into Evidence
#
ai
#
python
#
testing
#
productivity
Comments
Add Comment
6 min read
Freeze a Holdout Before You Quote a Coding-Agent Score
Casey Zhang
Casey Zhang
Casey Zhang
Follow
Sep 17
Freeze a Holdout Before You Quote a Coding-Agent Score
#
ai
#
python
#
testing
#
programming
Comments
Add Comment
9 min read
Don't Average Pass Rates Across Unequal Token Budgets
Casey Zhang
Casey Zhang
Casey Zhang
Follow
Sep 19
Don't Average Pass Rates Across Unequal Token Budgets
#
ai
#
testing
#
python
#
productivity
1
reaction
Comments
Add Comment
7 min read
Replay Fixtures Separate Agent Scores From API Weather
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 22
Replay Fixtures Separate Agent Scores From API Weather
#
ai
#
testing
#
python
#
performance
Comments
Add Comment
7 min read
Score Agent Patches on a Frozen Surface. Ledger the Flakes.
Finley Zhou
Finley Zhou
Finley Zhou
Follow
Sep 21
Score Agent Patches on a Frozen Surface. Ledger the Flakes.
#
testing
#
python
#
ai
#
devops
Comments
1
comment
8 min read
I Counted Drops as Wrongs. The Chart Was Theater.
Jordan Liu
Jordan Liu
Jordan Liu
Follow
Sep 21
I Counted Drops as Wrongs. The Chart Was Theater.
#
ai
#
python
#
testing
#
debugging
Comments
Add Comment
7 min read
Seal the Proof Before the Agent Starts
Harper Zhu
Harper Zhu
Harper Zhu
Follow
Sep 17
Seal the Proof Before the Agent Starts
#
ai
#
testing
#
git
#
productivity
Comments
Add Comment
6 min read
A Human Suite Is Not an Agent Harness: A Myth-Busting FAQ
Sam Yang
Sam Yang
Sam Yang
Follow
Sep 20
A Human Suite Is Not an Agent Harness: A Myth-Busting FAQ
#
testing
#
ai
#
debugging
#
productivity
1
reaction
Comments
1
comment
7 min read
If the Patch Authored the Test, Score the Overlap
Finley Zhou
Finley Zhou
Finley Zhou
Follow
Sep 17
If the Patch Authored the Test, Score the Overlap
#
testing
#
python
#
ai
#
devops
Comments
1
comment
6 min read
A Courtesy Slot Is Not a Lab: A Myth-Busting FAQ
Sam Yang
Sam Yang
Sam Yang
Follow
Sep 17
A Courtesy Slot Is Not a Lab: A Myth-Busting FAQ
#
ai
#
testing
#
programming
#
productivity
Comments
Add Comment
8 min read
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account