Why this lesson exists
Warm‑up runs are the silent hero of performance measurement. When you skip them or treat the first run as the final result, you discard the very data that tells you whether your compiler, runtime, or hardware is in a steady state.
Recap — where we are in the series
We have traversed a path of incremental rigor:
Problem
- Ep 1: single‑shot timings mislead.
- Ep 2: warm‑up and repeat are essential.
- Ep 3: screenshots are not evidence.
- Ep 4: metadata makes cross‑machine claims reliable.
- Ep 5: csperf observatory gives quick evidence.
- Ep 6: csperf doctor diagnoses toolchain.
- Ep 7: quickstart yields a reproducible artifact in 30 s.
- Ep 8: JSON is the source of truth.
- Ep 9: choosing the right backend.
- Ep 10: stability metrics expose volatility.
Observatory
- We have seen how csperf captures machine metadata (Ep 4) and how the doctor command validates the environment (Ep 6).
Measurement
- Today we focus on the warm‑up phase itself, the first step that often gets ignored.
Comparison
- We will compare two runs: zero warm‑up vs five warm‑up runs.
Depth
- The lesson will dig into the numbers that csperf records for each warm‑up iteration.
The misconception
Many developers assume that the first run of a benchmark is already representative. In reality, the first few iterations trigger JIT compilation, cache misses, and other transient effects that inflate the time.
What problem csperf solves (this episode's slice)
csperf lets you explicitly control the number of warm‑up runs and then compare the resulting artifacts. By keeping the repeat count constant, you isolate the effect of warm‑up.
Mental model
Think of a warm‑up run as a pre‑flight check. The first few passes prepare the system—load libraries, prime caches, and stabilize the CPU frequency. The repeat runs are the actual flight data.
Lab: install and first commands
# Install csperf (if not already)
pip install csperf
# Run with no warm‑up
csperf run \
--input examples/cpp/matrix_traversal.cpp \
--backend cpu \
--warmup-runs 0 \
--repeat-runs 5 \
--output results/w0.json
# Run with five warm‑up runs
csperf run \
--input examples/cpp/matrix_traversal.cpp \
--backend cpu \
--warmup-runs 5 \
--repeat-runs 5 \
--output results/w5.json
Lab: what we ran on this machine
The machine is an AMD Ryzen 7 9700X, 8 cores, 16 threads, running Ubuntu 24.04. The csperf artifact csperf/machine.txt records all relevant metadata.
Results (real numbers only)
The zero‑warm‑up run (w0.json) produced:
| Metric | Value |
|---|---|
execution_time_ms |
5.1514 |
min |
5.135 |
max |
5.176 |
mean |
5.1514 |
median |
5.145 |
stdev |
0.017897 |
cpu_cycles |
6 125 891 |
ref_cycles |
4 346 136 |
frontend_stall_cycles |
1 241 045 |
instruction_count |
12 224 343 |
branch_instructions |
1 707 665 |
branch_mispredictions |
13 665 |
cache_references |
588 738 |
cache_misses |
85 287 |
l1_cache_misses |
235 771 |
l1_icache_misses |
30 324 |
dtlb_load_misses |
4 684 |
itlb_load_misses |
260 |
amd_l2_ic_dc_miss_in_l2 |
103 151 |
ipc |
1.995521 |
l2_cache_misses |
103 151 |
The five‑warm‑up run (w5.json) is pending; we will generate it tomorrow. Until then, the numbers above illustrate the baseline you would see if you ignored warm‑up.
How to read the artifacts
-
metrics.execution_time_msis the aggregate time over the repeat runs. -
metrics.execution_time_summary_msgives the statistical spread. -
metrics.cpu_cyclesandmetrics.ref_cyclesshow the raw instruction throughput. -
metrics.frontend_stall_cyclesreveals how often the CPU stalled waiting for the compiler front‑end. -
metrics.cache_*metrics expose the cost of memory traffic.
When you later run w5.json, compare each of these fields to see how many cycles and how much time the warm‑up phase saved.
Common mistakes (teacher checklist)
- Skipping warm‑up – leads to inflated times.
- Using the first run as the final result – ignores transient effects.
- Not keeping repeat count constant – conflates warm‑up and repeat variability.
-
Ignoring the
execution_time_summary_ms– hides the spread. - Overlooking cache metrics – misses a major source of variance.
Try this next (homework)
Run the five‑warm‑up experiment on your machine. Compare the execution_time_ms and cpu_cycles between w0.json and w5.json. Note how the warm‑up reduces the mean time and the standard deviation.
Do this tonight — Episode 12 starts by assuming you did.
Closing
Warm‑up is not a luxury; it is the first honest measurement step. By controlling and comparing it, you turn a noisy signal into a reliable metric.
The series so far
- Ep 1: Why a Single ./a.out Time Misleads Your Performance Claims (this article)
- Ep 2: Warm vs Cold: Why a Single Trial Misleads Performance Claims
- Ep 3: Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts
- Ep 4: Comparing Across Machines: Why Metadata Matters in csperf
- Ep 5: csperf: A Lightweight Observatory for Honest Performance Tracking
- Ep 6: csperf Doctor: First Command to Diagnose Toolchain Issues
- Ep 7: csperf Quickstart: 30‑Second First Win
- Ep 8: Reading csperf JSON: The Real Performance Artifact
- Ep 9: Backends 101: Choosing the Right Measurement Surface
- Ep 10: Stability Metrics: Min, Max, and Standard Deviation as First-Class Citizens
- Ep 11: Warmup Deep Dive: What You’re Throwing Away and Why
Teaser next
Recall how we used metadata to compare machines (Ep 4) and now we’ll dive into locality: Row vs column access and how layout shows up in time (Ep 12).
Top comments (0)