DEV Community

#benchmark

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Glasshouse v0.1 Is Out: A Memory Benchmark for AI Systems

Glasshouse v0.1 Is Out: A Memory Benchmark for AI Systems

7
Comments 1
2 min read
Onboarding benchmarks from real data across 464 SaaS products: median tour completion is 29%, 1-2 step tours complete at 73%, 9+ step tours at 8%

Onboarding benchmarks from real data across 464 SaaS products: median tour completion is 29%, 1-2 step tours complete at 73%, 9+ step tours at 8%

Comments 1
2 min read
How LLM Evaluation Actually Works: Inside the Satellite Geo QCM Leaderboard

How LLM Evaluation Actually Works: Inside the Satellite Geo QCM Leaderboard

1
Comments
4 min read
The 5 Walls Between a 3M req/s HTTP Benchmark and Production

The 5 Walls Between a 3M req/s HTTP Benchmark and Production

Comments
7 min read
netcup VPS 1000 G12 benchmarked: how fast is it really?

netcup VPS 1000 G12 benchmarked: how fast is it really?

Comments
7 min read
A code review benchmark that isn't the vendor ranking itself

A code review benchmark that isn't the vendor ranking itself

Comments
4 min read
Enterprise Vector Database 2026: Qdrant vs Milvus vs pgvector vs Pinecone

Enterprise Vector Database 2026: Qdrant vs Milvus vs pgvector vs Pinecone

Comments
23 min read
You can read Claude Code's whole harness now. That's what every benchmark score throws away

You can read Claude Code's whole harness now. That's what every benchmark score throws away

Comments
5 min read
TypeSafe’s JEV Model: Is It Really 193x Faster and 444x Cheaper?

TypeSafe’s JEV Model: Is It Really 193x Faster and 444x Cheaper?

2
Comments 1
6 min read
JetBrains Ranked AI Agents on Real Kotlin Projects. The Token Column Is the Real Story.

JetBrains Ranked AI Agents on Real Kotlin Projects. The Token Column Is the Real Story.

1
Comments
7 min read
Two "Codex CLI" models on the same benchmark: the harness hides the model

Two "Codex CLI" models on the same benchmark: the harness hides the model

Comments
2 min read
How Fast Can Neovim Start? Benchmarking Popular Distros

How Fast Can Neovim Start? Benchmarking Popular Distros

Comments
6 min read
Why I’m Building a New AI Memory Benchmark (And Why the Existing Ones Fall Short)

Why I’m Building a New AI Memory Benchmark (And Why the Existing Ones Fall Short)

2
Comments
2 min read
Elysia 2 vs NestJS 12: Runtime +64.6%, Framework +10.6%

Elysia 2 vs NestJS 12: Runtime +64.6%, Framework +10.6%

Comments
14 min read
เมื่อ Benchmark โกหกคุณ, SWE-Bench ProMax กับคะแนนจริงที่โมเดลเก่งสุดทำได้แค่ 41.2%

เมื่อ Benchmark โกหกคุณ, SWE-Bench ProMax กับคะแนนจริงที่โมเดลเก่งสุดทำได้แค่ 41.2%

Comments
2 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.