Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
Leaderboard Forensics Series' Articles
Back to Ward Ed's Series
When a 0.4-Point Lead Means Nothing: Reading Open-Model Leaderboard Margins Like a Statistician
Ward Ed
Ward Ed
Ward Ed
Follow
Oct 6
When a 0.4-Point Lead Means Nothing: Reading Open-Model Leaderboard Margins Like a Statistician
#
huggingface
#
llm
#
benchmark
#
machinelearning
Comments
Add Comment
7 min read
Benchmark Contamination 101: How Train/Test Overlap Inflates Leaderboard Scores (and How to Catch It)
Ward Ed
Ward Ed
Ward Ed
Follow
Oct 6
Benchmark Contamination 101: How Train/Test Overlap Inflates Leaderboard Scores (and How to Catch It)
#
machinelearning
#
llm
#
benchmark
#
datascience
Comments
Add Comment
7 min read
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account