DEV Community

Foundations Series' Articles

Back to Ofri Peretz's Series
The Base Rate Problem: Why 95% Precision Means Nothing Without Context
Cover image for The Base Rate Problem: Why 95% Precision Means Nothing Without Context

The Base Rate Problem: Why 95% Precision Means Nothing Without Context

Comments
7 min read
Bias in Measurement: The Silent Failure Mode in Benchmark Design
Cover image for Bias in Measurement: The Silent Failure Mode in Benchmark Design

Bias in Measurement: The Silent Failure Mode in Benchmark Design

1
Comments
8 min read
Goodhart's Law in Benchmarking: When the Metric Becomes the Target
Cover image for Goodhart's Law in Benchmarking: When the Metric Becomes the Target

Goodhart's Law in Benchmarking: When the Metric Becomes the Target

Comments
7 min read
Inter-Rater Agreement and Cohen's Kappa: When Your Labels Are Opinions
Cover image for Inter-Rater Agreement and Cohen's Kappa: When Your Labels Are Opinions

Inter-Rater Agreement and Cohen's Kappa: When Your Labels Are Opinions

Comments
7 min read
The Confusion Matrix: What TP, FP, FN, and TN Actually Mean
Cover image for The Confusion Matrix: What TP, FP, FN, and TN Actually Mean

The Confusion Matrix: What TP, FP, FN, and TN Actually Mean

Comments
8 min read
Composite Scores and Weighting: Your 'Overall Score' Is an Editorial
Cover image for Composite Scores and Weighting: Your 'Overall Score' Is an Editorial

Composite Scores and Weighting: Your 'Overall Score' Is an Editorial

Comments
8 min read
CVSS Scores Explained: The Number Measures Severity, Not Risk
Cover image for CVSS Scores Explained: The Number Measures Severity, Not Risk

CVSS Scores Explained: The Number Measures Severity, Not Risk

Comments
8 min read
The CWE Taxonomy, Explained: A Name Is Not a Verdict
Cover image for The CWE Taxonomy, Explained: A Name Is Not a Verdict

The CWE Taxonomy, Explained: A Name Is Not a Verdict

Comments
8 min read
Ground Truth in Security Testing: Who Decides What's Vulnerable?
Cover image for Ground Truth in Security Testing: Who Decides What's Vulnerable?

Ground Truth in Security Testing: Who Decides What's Vulnerable?

Comments
8 min read
OWASP Top 10, Explained: An Address System, Not a Severity Scale
Cover image for OWASP Top 10, Explained: An Address System, Not a Severity Scale

OWASP Top 10, Explained: An Address System, Not a Severity Scale

Comments
8 min read
Ranking vs. Measuring: What a Leaderboard Throws Away
Cover image for Ranking vs. Measuring: What a Leaderboard Throws Away

Ranking vs. Measuring: What a Leaderboard Throws Away

Comments
8 min read
Reproducibility vs Replicability: Why Benchmark Numbers Need Both
Cover image for Reproducibility vs Replicability: Why Benchmark Numbers Need Both

Reproducibility vs Replicability: Why Benchmark Numbers Need Both

Comments
7 min read
Statistical Significance and p-Values: What p > 0.05 Actually Means
Cover image for Statistical Significance and p-Values: What p > 0.05 Actually Means

Statistical Significance and p-Values: What p > 0.05 Actually Means

Comments
8 min read
Taint vs. Heuristic Detection: The Difference Between a Proof and a Hunch
Cover image for Taint vs. Heuristic Detection: The Difference Between a Proof and a Hunch

Taint vs. Heuristic Detection: The Difference Between a Proof and a Hunch

1
Comments 2
7 min read
Precision, Recall, and F1 for Static Analysis: Same Score, Opposite Tools
Cover image for Precision, Recall, and F1 for Static Analysis: Same Score, Opposite Tools

Precision, Recall, and F1 for Static Analysis: Same Score, Opposite Tools

Comments
8 min read
Proxy Metrics: The Number You Optimize Is Not the Thing You Want
Cover image for Proxy Metrics: The Number You Optimize Is Not the Thing You Want

Proxy Metrics: The Number You Optimize Is Not the Thing You Want

Comments
8 min read
Sample Size and Statistical Power: What a Small Sample Can and Cannot Tell You
Cover image for Sample Size and Statistical Power: What a Small Sample Can and Cannot Tell You

Sample Size and Statistical Power: What a Small Sample Can and Cannot Tell You

1
Comments
8 min read
Static Analysis vs. SAST vs. Linting: The Taxonomy That Matters for Security Teams
Cover image for Static Analysis vs. SAST vs. Linting: The Taxonomy That Matters for Security Teams

Static Analysis vs. SAST vs. Linting: The Taxonomy That Matters for Security Teams

Comments
8 min read
Valid vs. Reliable Metrics: Consistent Numbers Can Still Be Wrong
Cover image for Valid vs. Reliable Metrics: Consistent Numbers Can Still Be Wrong

Valid vs. Reliable Metrics: Consistent Numbers Can Still Be Wrong

1
Comments
7 min read
Nobody Writes Bad Crypto. They Write Correct Crypto at Four Layers.
Cover image for Nobody Writes Bad Crypto. They Write Correct Crypto at Four Layers.

Nobody Writes Bad Crypto. They Write Correct Crypto at Four Layers.

Comments
4 min read
innerHTML Has Five Doors. Most Reviews Only Watch One.
Cover image for innerHTML Has Five Doors. Most Reviews Only Watch One.

innerHTML Has Five Doors. Most Reviews Only Watch One.

2
Comments 3
5 min read