Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
benchmark
Follow
Hide
Posts
Left menu
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Who Leads Hugging Face's Official Benchmarks? A Measurement of Concentration, Gaps and Hidden Entries
QuantID
QuantID
QuantID
Follow
Oct 5
Who Leads Hugging Face's Official Benchmarks? A Measurement of Concentration, Gaps and Hidden Entries
#
datascience
#
huggingface
#
benchmark
#
statistics
Comments
Add Comment
11 min read
How Hugging Face Official Benchmark Leaderboards Actually Work: .eval_results YAML, the base_model Filter, and the 30% the Default View Hides
Ward Ed
Ward Ed
Ward Ed
Follow
Oct 5
How Hugging Face Official Benchmark Leaderboards Actually Work: .eval_results YAML, the base_model Filter, and the 30% the Default View Hides
#
huggingface
#
llm
#
benchmark
#
machinelearning
Comments
Add Comment
10 min read
AI Bias Detection: HY3 vs Nemotron 3 Ultra on the Bias Stereotypes Benchmark
RESK
RESK
RESK
Follow
Oct 5
AI Bias Detection: HY3 vs Nemotron 3 Ultra on the Bias Stereotypes Benchmark
#
ai
#
bias
#
fairness
#
benchmark
Comments
Add Comment
3 min read
Hugging Face now has 48 official benchmarks. Here is what the map looks like
ai maya
ai maya
ai maya
Follow
Oct 4
Hugging Face now has 48 official benchmarks. Here is what the map looks like
#
ai
#
llm
#
machinelearning
#
benchmark
Comments
Add Comment
3 min read
Hugging Face official benchmarks: the complete list (48) and how their leaderboards work
AI OpenFree
AI OpenFree
AI OpenFree
Follow
Oct 4
Hugging Face official benchmarks: the complete list (48) and how their leaderboards work
#
huggingface
#
llm
#
benchmark
#
machinelearning
Comments
Add Comment
4 min read
Gemini 4 Argon ขึ้นที่ 1 Arena AI แล้ว แต่มีตัวเลขสามตัวที่บทความต้นทางไม่ได้บอก
Nokka
Nokka
Nokka
Follow
Oct 2
Gemini 4 Argon ขึ้นที่ 1 Arena AI แล้ว แต่มีตัวเลขสามตัวที่บทความต้นทางไม่ได้บอก
#
thai
#
ai
#
google
#
benchmark
Comments
Add Comment
3 min read
โมเดล 27B ย่อ 4 บิต ชนะโมเดลคลาวด์ในโจทย์เดียว แต่ทำไมยังไม่ควรเชื่อ
Nokka
Nokka
Nokka
Follow
Oct 2
โมเดล 27B ย่อ 4 บิต ชนะโมเดลคลาวด์ในโจทย์เดียว แต่ทำไมยังไม่ควรเชื่อ
#
thai
#
localllama
#
ai
#
benchmark
Comments
Add Comment
2 min read
Understanding Tokens per Second: A Practical Benchmark Guide
qjlsh
qjlsh
qjlsh
Follow
Oct 2
Understanding Tokens per Second: A Practical Benchmark Guide
#
tokens
#
benchmark
#
abotwrotethis
#
ai
Comments
Add Comment
5 min read
How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark
RESK
RESK
RESK
Follow
Oct 2
How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark
#
llm
#
evaluation
#
benchmark
#
reverseengineering
2
reactions
Comments
Add Comment
4 min read
We recompute TypeSafe's 444x claim — here's what we found
chunxiaoxx
chunxiaoxx
chunxiaoxx
Follow
Sep 29
We recompute TypeSafe's 444x claim — here's what we found
#
ai
#
llm
#
benchmark
#
verification
Comments
Add Comment
1 min read
Verification as Protocol: We Test AI Agents' Memory — Our Grader Failed First
chunxiaoxx
chunxiaoxx
chunxiaoxx
Follow
Sep 29
Verification as Protocol: We Test AI Agents' Memory — Our Grader Failed First
#
ai
#
agents
#
benchmark
#
verification
Comments
Add Comment
1 min read
One Hung API Call Used to Kill My 1,000-Run Benchmark. Here's the Fix.
Debashish Ghosal
Debashish Ghosal
Debashish Ghosal
Follow
Sep 26
One Hung API Call Used to Kill My 1,000-Run Benchmark. Here's the Fix.
#
ai
#
llm
#
programming
#
benchmark
7
reactions
Comments
Add Comment
3 min read
Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs
RESK
RESK
RESK
Follow
Sep 30
Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs
#
llm
#
benchmark
#
politics
#
evaluation
1
reaction
Comments
Add Comment
4 min read
In Search of the Best A.I. Literary Translator – Detailed
Bozhidar Valchev
Bozhidar Valchev
Bozhidar Valchev
Follow
Oct 1
In Search of the Best A.I. Literary Translator – Detailed
#
ai
#
llm
#
translation
#
benchmark
Comments
Add Comment
15 min read
Claude Opus 5.5: 40% cheaper, frontier-grade performance
The Dev Signal
The Dev Signal
The Dev Signal
Follow
Sep 24
Claude Opus 5.5: 40% cheaper, frontier-grade performance
#
ai
#
devtools
#
programming
#
benchmark
Comments
Add Comment
5 min read
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account