DEV Community

compilersutra
compilersutra

Posted on Originally published at compilersutra.com

Evidence Levels: What a Single JSON Can and Cannot Claim

Why this lesson exists

Teams often shout "this runs 10 % faster" after a single csperf run. The claim is true only if the evidence is strong enough to survive noise, variance, and the temptation to cherry‑pick.

Recap — where we are in the series

Ep 1: single‑shot timings mislead. Ep 2: warm‑up and repeat are essential. Ep 3: screenshots are not artifacts. Ep 4: metadata is needed for cross‑machine claims. Ep 5: observatory commands give quick evidence. Ep 6: doctor validates toolchain. Ep 7: quickstart gets you a reproducible artifact in 30 s. Ep 8: JSON is the source of truth. Ep 9: backends matter. Ep 10: stability metrics expose volatility. Ep 11: warm‑up impact quantified. Ep 12: locality measured. Ep 13: O‑3 myth debunked. Ep 14: Linux perf counters validated. Ep 15: noise sources isolated.

The misconception

A single JSON file looks like a definitive benchmark, but without context it can be misleading. People read the mean and declare a win, ignoring min/max, stdev, or the fact that the run was a warm‑up.

What problem csperf solves (this episode's slice)

csperf attaches evidence levels to each metric: raw, sanitized, validated. By inspecting these levels you can decide whether a claim is statistically sound, machine‑aware, or comparable.

Mental model

Evidence Level What it guarantees Typical use case
raw The raw number from the profiler Quick sanity check
sanitized Outliers removed, min/max reported Reporting to stakeholders
validated Cross‑checked against metadata and noise model Publishing a paper

Lab: install and first commands

# Ensure csperf is in your path
source .venv/bin/activate
# Run the tiled matrix multiply workload
csperf run \
  --input examples/cpp/tiled_matmul.cpp \
  --backend cpu \
  --warmup-runs 1 \
  --repeat-runs 5 \
  --output results/matmul.json
# Inspect the artifact
csperf profile results/matmul.json
Enter fullscreen mode Exit fullscreen mode

Lab: what we ran on this machine

Hostname: f4c59d864117
CPU: AMD Ryzen 7 9700X 8‑core, 16 threads
OS: Ubuntu 24.04 (7.0.0‑34‑generic)

Results (real numbers only)

Metric Value Min Max Stdev
execution_time_ms 5.135 5.123 5.143 0.0076
cpu_cycles 2 401 953 – – –
ipc 2.1918 – – –

How to read the artifacts

Open results/matmul.json:

{
  "metrics": {
    "execution_time_ms": 5.135,
    "execution_time_summary_ms": {
      "count": 5,
      "min": 5.123,
      "max": 5.143,
      "mean": 5.135,
      "median": 5.137,
      "stdev": 0.007616
    },
    ...
  }
}
Enter fullscreen mode Exit fullscreen mode

The execution_time_summary_ms block gives the evidence level: sanitized because it reports min/max and stdev. If you see only a single number, the evidence is raw.

Common mistakes (teacher checklist)

  1. Ignoring min/max – a single outlier can skew the mean.
  2. Treating warm‑up as a measurement – always separate warm‑up runs.
  3. Assuming the JSON is the final word – check the metadata section for CPU, OS, and compiler.
  4. Comparing artifacts from different machines without metadata – use csperf observatory first.

Try this next (homework)

Do this tonight — Episode 17 starts by assuming you did.

Closing

Evidence levels let you say “this run is statistically significant” rather than “this run is faster”. The next episode will show how to diff two runs without spreadsheet chaos.

The series so far

  • Ep 1 – Why a Single ./a.out Time Misleads Your Performance Claims
  • Ep 2 – Warm vs Cold: Why a Single Trial Misleads Performance Claims
  • Ep 3 – Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts
  • Ep 4 – Comparing Across Machines: Why Metadata Matters in csperf
  • Ep 5 – csperf: A Lightweight Observatory for Honest Performance Tracking
  • Ep 6 – csperf Doctor: First Command to Diagnose Toolchain Issues
  • Ep 7 – csperf Quickstart: 30‑Second First Win
  • Ep 8 – Reading csperf JSON: The Real Performance Artifact
  • Ep 9 – Backends 101: Choosing the Right Measurement Surface
  • Ep 10 – Stability Metrics: Min, Max, and Standard Deviation as First-Class Citizens
  • Ep 11 – Warmup Deep Dive: What You’re Throwing Away and Why
  • Ep 12 – Row vs. Column: Measuring Locality with csperf
  • Ep 13 – Unmasking the O‑3 Myth: Real‑World Optimizer Impact with csperf
  • Ep 14 – Turning on Linux perf Counters with csperf (honestly)
  • Ep 15 – Noise Sources: Affinity, Frequency, and Background Load — How to isolate and quantify non‑compiler variance with csperf
  • Ep 16 – Evidence Levels: What a Single JSON Can and Cannot Claim (this article)

Top comments (0)