Why this lesson exists
When you look at a single mean you miss the story of variability. A program that sometimes runs in 5 ms and sometimes in 10 ms can still have a mean of 7.5 ms – a perfectly plausible number that hides a performance cliff.
Recap — where we are in the series
- Ep 1: Why a Single ./a.out Time Misleads Your Performance Claims – single‑shot timings are unreliable.
- Ep 2: Warm vs Cold: Why a Single Trial Misleads Performance Claims – warm‑up and repeat runs give reproducible data.
- Ep 3: Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts – artifacts beat screenshots.
- Ep 4: Comparing Across Machines: Why Metadata Matters in csperf – metadata enables honest cross‑machine comparison.
- Ep 5: csperf: A Lightweight Observatory for Honest Performance Tracking – observatory commands give quick evidence.
- Ep 6: csperf Doctor: First Command to Diagnose Toolchain Issues – validate toolchain before blaming performance.
- Ep 7: csperf Quickstart: 30‑Second First Win – first reproducible artifact in half a minute.
- Ep 8: Reading csperf JSON: The Real Performance Artifact – JSON is the source of truth.
- Ep 9: Backends 101: Choosing the Right Measurement Surface – list backends, enumerate devices, run workloads on the correct surface.
The misconception
People often report only the mean execution time. They ignore that the mean can be the same for wildly different distributions – a narrow bell curve or a long tail.
What problem csperf solves (this episode's slice)
csperf now exposes min, max, and standard deviation in the summary block, making volatility visible. You can spot outliers, decide whether to discard them, and compare stability across compiler flags.
Mental model
Think of a performance run as a distribution of samples. The mean tells you the center, but the min/max and stdev tell you the spread. A low stdev means you can trust the mean; a high stdev means you should look deeper.
Lab: install and first commands
# Install csperf if you haven’t already
pip install csperf
Lab: what we ran on this machine
We used the matrix traversal example on an AMD Ryzen 7 9700X. The command sequence:
csperf run \
--input examples/cpp/matrix_traversal.cpp \
--backend cpu \
--warmup-runs 3 \
--repeat-runs 10 \
--output results/stable.json
csperf profile results/stable.json
Results (real numbers only)
| Metric | Value |
|---|---|
| execution_time_ms | 5.1458 |
| execution_time_summary_ms.min | 5.132 |
| execution_time_summary_ms.max | 5.162 |
| execution_time_summary_ms.mean | 5.1458 |
| execution_time_summary_ms.stdev | 0.010358 |
| cpu_cycles | 7 317 425 |
| ipc | 1.68653 |
How to read the artifacts
Open results/stable.json. The execution_time_summary_ms object contains count, min, max, mean, median, and stdev. The metrics section shows raw counters. The artifacts field lists CSV/XLSX for spreadsheet analysis.
Common mistakes (teacher checklist)
- Ignoring stdev – assume mean is enough.
- Using too few repeats – 10 runs is a minimum for stable stats.
- Discarding warm‑ups – warm‑up runs are intentional discards.
-
Comparing across machines without metadata – always check the
hardwaresection.
Try this next (homework)
Run the same workload with --repeat-runs 20 and compare the stdev. Does it shrink? Document the change.
Do this tonight — Episode 11 starts by assuming you did.
Closing
Stability metrics let you see whether a mean is trustworthy. In the next episode we’ll dig into what warm‑up runs actually discard and why that matters.
The series so far
- Ep 1: Why a Single ./a.out Time Misleads Your Performance Claims
- Ep 2: Warm vs Cold: Why a Single Trial Misleads Performance Claims
- Ep 3: Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts
- Ep 4: Comparing Across Machines: Why Metadata Matters in csperf
- Ep 5: csperf: A Lightweight Observatory for Honest Performance Tracking
- Ep 6: csperf Doctor: First Command to Diagnose Toolchain Issues
- Ep 7: csperf Quickstart: 30‑Second First Win
- Ep 8: Reading csperf JSON: The Real Performance Artifact
- Ep 9: Backends 101: Choosing the Right Measurement Surface
- Ep 10: Stability Metrics: Min, Max, and Standard Deviation as First-Class Citizens
Top comments (0)