DEV Community

#benchmark

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
OpenRouter vs Vercel vs LLMGateway Performance

OpenRouter vs Vercel vs LLMGateway Performance

Comments
6 min read
MCPMark v2: InsForge on Sonnet 4.6

MCPMark v2: InsForge on Sonnet 4.6

2
Comments 2
3 min read
One RTX 5090 vs a 12-GPU Cluster — Benchmarking a Decade of GPUs on the Same Go Proof

One RTX 5090 vs a 12-GPU Cluster — Benchmarking a Decade of GPUs on the Same Go Proof

Comments
4 min read
SDABench: A New Benchmark for Evaluating LLMs in Scientific Discovery

SDABench: A New Benchmark for Evaluating LLMs in Scientific Discovery

Comments
4 min read
Can You Beat an LLM? Building Humans vs. Humanity's Last Exam

Can You Beat an LLM? Building Humans vs. Humanity's Last Exam

11
Comments
4 min read
Model Showdown Round 9: Qwen 3.6 27B vs Qwen 3.6 35B-A3B vs Qwythos-9B vs GLM-4.7-Flash vs Nemotron-3-Nano

Model Showdown Round 9: Qwen 3.6 27B vs Qwen 3.6 35B-A3B vs Qwythos-9B vs GLM-4.7-Flash vs Nemotron-3-Nano

Comments
14 min read
DeepSeek vs GLM vs Qwen: Which Free LLM API is Best for Your Project?

DeepSeek vs GLM vs Qwen: Which Free LLM API is Best for Your Project?

Comments
4 min read
AdvancedMathBench: A New Benchmark for LLM Advanced Mathematical Reasoning

AdvancedMathBench: A New Benchmark for LLM Advanced Mathematical Reasoning

Comments
3 min read
TurboQuant, Four Months Later: Chasing Google's 6x VRAM Claim Into the Wild

TurboQuant, Four Months Later: Chasing Google's 6x VRAM Claim Into the Wild

Comments
6 min read
Your agent's memory remembers what you chose. Does it remember what you rejected?

Your agent's memory remembers what you chose. Does it remember what you rejected?

3
Comments
5 min read
Which LLM should I actually code with? I built a small benchmark to find out

Which LLM should I actually code with? I built a small benchmark to find out

Comments
2 min read
I Benchmarked 42 Compression Formats Spanning Four Decades. Here's What to Actually Use.

I Benchmarked 42 Compression Formats Spanning Four Decades. Here's What to Actually Use.

Comments
5 min read
ComfyUI, Lemonade, and LocalAI: Scouting the Next Wave of Homelab AI Tools

ComfyUI, Lemonade, and LocalAI: Scouting the Next Wave of Homelab AI Tools

Comments
7 min read
AI Coding Tools Benchmark 2026: Cursor vs Copilot vs Windsurf vs Claude Code

AI Coding Tools Benchmark 2026: Cursor vs Copilot vs Windsurf vs Claude Code

1
Comments
5 min read
The Same RTX 5090, but the GPU Sat Idle — a CPU-Bound Go Solver and the Case for L2 Cache

The Same RTX 5090, but the GPU Sat Idle — a CPU-Bound Go Solver and the Case for L2 Cache

Comments
6 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.