DEV Community

#benchmark

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Who Leads Hugging Face's Official Benchmarks? A Measurement of Concentration, Gaps and Hidden Entries

Who Leads Hugging Face's Official Benchmarks? A Measurement of Concentration, Gaps and Hidden Entries

Comments
11 min read
How Hugging Face Official Benchmark Leaderboards Actually Work: .eval_results YAML, the base_model Filter, and the 30% the Default View Hides

How Hugging Face Official Benchmark Leaderboards Actually Work: .eval_results YAML, the base_model Filter, and the 30% the Default View Hides

Comments
10 min read
AI Bias Detection: HY3 vs Nemotron 3 Ultra on the Bias Stereotypes Benchmark

AI Bias Detection: HY3 vs Nemotron 3 Ultra on the Bias Stereotypes Benchmark

Comments
3 min read
Hugging Face now has 48 official benchmarks. Here is what the map looks like

Hugging Face now has 48 official benchmarks. Here is what the map looks like

Comments
3 min read
Hugging Face official benchmarks: the complete list (48) and how their leaderboards work

Hugging Face official benchmarks: the complete list (48) and how their leaderboards work

Comments
4 min read
Gemini 4 Argon ขึ้นที่ 1 Arena AI แล้ว แต่มีตัวเลขสามตัวที่บทความต้นทางไม่ได้บอก

Gemini 4 Argon ขึ้นที่ 1 Arena AI แล้ว แต่มีตัวเลขสามตัวที่บทความต้นทางไม่ได้บอก

Comments
3 min read
โมเดล 27B ย่อ 4 บิต ชนะโมเดลคลาวด์ในโจทย์เดียว แต่ทำไมยังไม่ควรเชื่อ

โมเดล 27B ย่อ 4 บิต ชนะโมเดลคลาวด์ในโจทย์เดียว แต่ทำไมยังไม่ควรเชื่อ

Comments
2 min read
Understanding Tokens per Second: A Practical Benchmark Guide

Understanding Tokens per Second: A Practical Benchmark Guide

Comments
5 min read
How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark

How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark

2
Comments
4 min read
We recompute TypeSafe's 444x claim — here's what we found

We recompute TypeSafe's 444x claim — here's what we found

Comments
1 min read
Verification as Protocol: We Test AI Agents' Memory — Our Grader Failed First

Verification as Protocol: We Test AI Agents' Memory — Our Grader Failed First

Comments
1 min read
One Hung API Call Used to Kill My 1,000-Run Benchmark. Here's the Fix.

One Hung API Call Used to Kill My 1,000-Run Benchmark. Here's the Fix.

7
Comments
3 min read
Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs

Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs

1
Comments
4 min read
In Search of the Best A.I. Literary Translator – Detailed

In Search of the Best A.I. Literary Translator – Detailed

Comments
15 min read
Claude Opus 5.5: 40% cheaper, frontier-grade performance

Claude Opus 5.5: 40% cheaper, frontier-grade performance

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.