DEV Community

#benchmarks

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Frontier Models Hit a Wall on Research Thinking

Frontier Models Hit a Wall on Research Thinking

Comments
2 min read
Ranking Language Models by How Well They Spot Liars

Ranking Language Models by How Well They Spot Liars

Comments
9 min read
The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It

The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It

1
Comments
4 min read
Labs are ditching factual knowledge for reasoning speed

Labs are ditching factual knowledge for reasoning speed

Comments
2 min read
Why AI Benchmarks Mean Less Than You Think

Why AI Benchmarks Mean Less Than You Think

Comments
6 min read
An AI Capture-the-Flag Tournament: What the Scoreboard Counted

An AI Capture-the-Flag Tournament: What the Scoreboard Counted

Comments
6 min read
Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.

Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.

1
Comments
8 min read
Meta's Muse Code Clears 59% on Deep Software Engineering

Meta's Muse Code Clears 59% on Deep Software Engineering

Comments
2 min read
Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)

Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)

Comments
9 min read
SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses

SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses

Comments
4 min read
UniClawBench: A New Benchmark for Proactive AI Agents in Real-World Scenarios

UniClawBench: A New Benchmark for Proactive AI Agents in Real-World Scenarios

Comments
3 min read
AI News Roundup: Grok 4.5 Hits Tesla, Perplexity's Orchestrator Beats Opus, and Meta Undercuts Pricing

AI News Roundup: Grok 4.5 Hits Tesla, Perplexity's Orchestrator Beats Opus, and Meta Undercuts Pricing

Comments
2 min read
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Comments
7 min read
We hit 99.95% on the LoCoMo memory benchmark. Here's the catch, and why it still matters.

We hit 99.95% on the LoCoMo memory benchmark. Here's the catch, and why it still matters.

6
Comments 1
4 min read
Agent Leaderboards Measure Score. We Added Price.

Agent Leaderboards Measure Score. We Added Price.

Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.