DEV Community

Foundations Series' Articles

Back to Ofri Peretz's Series
The Base Rate Problem: Why 95% Precision Means Nothing Without Context

The Base Rate Problem: Why 95% Precision Means Nothing Without Context

Comments
7 min read
Bias in Measurement: The Silent Failure Mode in Benchmark Design
Cover image for Bias in Measurement: The Silent Failure Mode in Benchmark Design

Bias in Measurement: The Silent Failure Mode in Benchmark Design

1
Comments
8 min read
Goodhart's Law in Benchmarking: When the Metric Becomes the Target

Goodhart's Law in Benchmarking: When the Metric Becomes the Target

Comments
7 min read
Inter-Rater Agreement and Cohen's Kappa: When Your Labels Are Opinions

Inter-Rater Agreement and Cohen's Kappa: When Your Labels Are Opinions

Comments
7 min read
The Confusion Matrix: What TP, FP, FN, and TN Actually Mean

The Confusion Matrix: What TP, FP, FN, and TN Actually Mean

Comments
8 min read
Composite Scores and Weighting: Your 'Overall Score' Is an Editorial

Composite Scores and Weighting: Your 'Overall Score' Is an Editorial

Comments
8 min read
CVSS Scores Explained: The Number Measures Severity, Not Risk

CVSS Scores Explained: The Number Measures Severity, Not Risk

Comments
8 min read
The CWE Taxonomy, Explained: A Name Is Not a Verdict

The CWE Taxonomy, Explained: A Name Is Not a Verdict

Comments
8 min read
Ground Truth in Security Testing: Who Decides What's Vulnerable?

Ground Truth in Security Testing: Who Decides What's Vulnerable?

Comments
8 min read
OWASP Top 10, Explained: An Address System, Not a Severity Scale

OWASP Top 10, Explained: An Address System, Not a Severity Scale

Comments
8 min read
Ranking vs. Measuring: What a Leaderboard Throws Away

Ranking vs. Measuring: What a Leaderboard Throws Away

Comments
8 min read
Reproducibility vs Replicability: Why Benchmark Numbers Need Both

Reproducibility vs Replicability: Why Benchmark Numbers Need Both

Comments
7 min read
Statistical Significance and p-Values: What p > 0.05 Actually Means

Statistical Significance and p-Values: What p > 0.05 Actually Means

Comments
8 min read
Taint vs. Heuristic Detection: The Difference Between a Proof and a Hunch

Taint vs. Heuristic Detection: The Difference Between a Proof and a Hunch

1
Comments 1
7 min read
Precision, Recall, and F1 for Static Analysis: Same Score, Opposite Tools
Cover image for Precision, Recall, and F1 for Static Analysis: Same Score, Opposite Tools

Precision, Recall, and F1 for Static Analysis: Same Score, Opposite Tools

Comments
8 min read
Proxy Metrics: The Number You Optimize Is Not the Thing You Want

Proxy Metrics: The Number You Optimize Is Not the Thing You Want

Comments
8 min read
Sample Size and Statistical Power: What a Small Sample Can and Cannot Tell You

Sample Size and Statistical Power: What a Small Sample Can and Cannot Tell You

Comments
8 min read
Static Analysis vs. SAST vs. Linting: The Taxonomy That Matters for Security Teams

Static Analysis vs. SAST vs. Linting: The Taxonomy That Matters for Security Teams

Comments
8 min read
Valid vs. Reliable Metrics: Consistent Numbers Can Still Be Wrong

Valid vs. Reliable Metrics: Consistent Numbers Can Still Be Wrong

1
Comments
7 min read