Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
← All Trends
Rigorous Measurement Protocols for AI Agent Scoring
12 posts in this trend in the last 7 days
•
Active about 3 hours ago
Sign the Metric Before You Publish an Agent Score
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 20
Sign the Metric Before You Publish an Agent Score
#
ai
#
testing
#
python
#
agents
1
reaction
Comments
1
comment
6 min read
Hash the Task Pack Before Ranking Coding Agents
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 23
Hash the Task Pack Before Ranking Coding Agents
#
ai
#
testing
#
python
#
productivity
Comments
Add Comment
6 min read
Freeze the Manifest Before the Agent Leaderboard
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 17
Freeze the Manifest Before the Agent Leaderboard
#
ai
#
testing
#
python
#
opensource
Comments
Add Comment
8 min read
Calibrate the Judge Before the Agent Score
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 19
Calibrate the Judge Before the Agent Score
#
ai
#
testing
#
python
#
programming
Comments
Add Comment
7 min read
Agent Scores Without a Null Pack Are Marketing
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 18
Agent Scores Without a Null Pack Are Marketing
#
ai
#
testing
#
python
#
productivity
Comments
Add Comment
6 min read
Control Deltas Turn Agent Scores Into Evidence
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 21
Control Deltas Turn Agent Scores Into Evidence
#
ai
#
python
#
testing
#
productivity
Comments
Add Comment
6 min read
Freeze a Holdout Before You Quote a Coding-Agent Score
Casey Zhang
Casey Zhang
Casey Zhang
Follow
Sep 17
Freeze a Holdout Before You Quote a Coding-Agent Score
#
ai
#
python
#
testing
#
programming
Comments
Add Comment
9 min read
Replay Fixtures Separate Agent Scores From API Weather
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 22
Replay Fixtures Separate Agent Scores From API Weather
#
ai
#
testing
#
python
#
performance
Comments
Add Comment
7 min read
Don't Average Pass Rates Across Unequal Token Budgets
Casey Zhang
Casey Zhang
Casey Zhang
Follow
Sep 19
Don't Average Pass Rates Across Unequal Token Budgets
#
ai
#
testing
#
python
#
productivity
1
reaction
Comments
Add Comment
7 min read
Score Agent Patches on a Frozen Surface. Ledger the Flakes.
Finley Zhou
Finley Zhou
Finley Zhou
Follow
Sep 21
Score Agent Patches on a Frozen Surface. Ledger the Flakes.
#
testing
#
python
#
ai
#
devops
Comments
1
comment
8 min read
I Counted Drops as Wrongs. The Chart Was Theater.
Jordan Liu
Jordan Liu
Jordan Liu
Follow
Sep 21
I Counted Drops as Wrongs. The Chart Was Theater.
#
ai
#
python
#
testing
#
debugging
Comments
Add Comment
7 min read
If the Patch Authored the Test, Score the Overlap
Finley Zhou
Finley Zhou
Finley Zhou
Follow
Sep 17
If the Patch Authored the Test, Score the Overlap
#
testing
#
python
#
ai
#
devops
Comments
1
comment
6 min read
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account