DEV Community

compilersutra
compilersutra

Posted on Originally published at compilersutra.com

Warmup Deep Dive: What You’re Throwing Away and Why

Why this lesson exists

Warm‑up runs are the silent hero of performance measurement. When you skip them or treat the first run as the final result, you discard the very data that tells you whether your compiler, runtime, or hardware is in a steady state.

Recap — where we are in the series

We have traversed a path of incremental rigor:

Problem

  • Ep 1: single‑shot timings mislead.
  • Ep 2: warm‑up and repeat are essential.
  • Ep 3: screenshots are not evidence.
  • Ep 4: metadata makes cross‑machine claims reliable.
  • Ep 5: csperf observatory gives quick evidence.
  • Ep 6: csperf doctor diagnoses toolchain.
  • Ep 7: quickstart yields a reproducible artifact in 30 s.
  • Ep 8: JSON is the source of truth.
  • Ep 9: choosing the right backend.
  • Ep 10: stability metrics expose volatility.

Observatory

  • We have seen how csperf captures machine metadata (Ep 4) and how the doctor command validates the environment (Ep 6).

Measurement

  • Today we focus on the warm‑up phase itself, the first step that often gets ignored.

Comparison

  • We will compare two runs: zero warm‑up vs five warm‑up runs.

Depth

  • The lesson will dig into the numbers that csperf records for each warm‑up iteration.

The misconception

Many developers assume that the first run of a benchmark is already representative. In reality, the first few iterations trigger JIT compilation, cache misses, and other transient effects that inflate the time.

What problem csperf solves (this episode's slice)

csperf lets you explicitly control the number of warm‑up runs and then compare the resulting artifacts. By keeping the repeat count constant, you isolate the effect of warm‑up.

Mental model

Think of a warm‑up run as a pre‑flight check. The first few passes prepare the system—load libraries, prime caches, and stabilize the CPU frequency. The repeat runs are the actual flight data.

Lab: install and first commands

# Install csperf (if not already)
pip install csperf

# Run with no warm‑up
csperf run \
  --input examples/cpp/matrix_traversal.cpp \
  --backend cpu \
  --warmup-runs 0 \
  --repeat-runs 5 \
  --output results/w0.json

# Run with five warm‑up runs
csperf run \
  --input examples/cpp/matrix_traversal.cpp \
  --backend cpu \
  --warmup-runs 5 \
  --repeat-runs 5 \
  --output results/w5.json
Enter fullscreen mode Exit fullscreen mode

Lab: what we ran on this machine

The machine is an AMD Ryzen 7 9700X, 8 cores, 16 threads, running Ubuntu 24.04. The csperf artifact csperf/machine.txt records all relevant metadata.

Results (real numbers only)

The zero‑warm‑up run (w0.json) produced:

Metric Value
execution_time_ms 5.1514
min 5.135
max 5.176
mean 5.1514
median 5.145
stdev 0.017897
cpu_cycles 6 125 891
ref_cycles 4 346 136
frontend_stall_cycles 1 241 045
instruction_count 12 224 343
branch_instructions 1 707 665
branch_mispredictions 13 665
cache_references 588 738
cache_misses 85 287
l1_cache_misses 235 771
l1_icache_misses 30 324
dtlb_load_misses 4 684
itlb_load_misses 260
amd_l2_ic_dc_miss_in_l2 103 151
ipc 1.995521
l2_cache_misses 103 151

The five‑warm‑up run (w5.json) is pending; we will generate it tomorrow. Until then, the numbers above illustrate the baseline you would see if you ignored warm‑up.

How to read the artifacts

  • metrics.execution_time_ms is the aggregate time over the repeat runs.
  • metrics.execution_time_summary_ms gives the statistical spread.
  • metrics.cpu_cycles and metrics.ref_cycles show the raw instruction throughput.
  • metrics.frontend_stall_cycles reveals how often the CPU stalled waiting for the compiler front‑end.
  • metrics.cache_* metrics expose the cost of memory traffic.

When you later run w5.json, compare each of these fields to see how many cycles and how much time the warm‑up phase saved.

Common mistakes (teacher checklist)

  1. Skipping warm‑up – leads to inflated times.
  2. Using the first run as the final result – ignores transient effects.
  3. Not keeping repeat count constant – conflates warm‑up and repeat variability.
  4. Ignoring the execution_time_summary_ms – hides the spread.
  5. Overlooking cache metrics – misses a major source of variance.

Try this next (homework)

Run the five‑warm‑up experiment on your machine. Compare the execution_time_ms and cpu_cycles between w0.json and w5.json. Note how the warm‑up reduces the mean time and the standard deviation.

Do this tonight — Episode 12 starts by assuming you did.

Closing

Warm‑up is not a luxury; it is the first honest measurement step. By controlling and comparing it, you turn a noisy signal into a reliable metric.

The series so far

  • Ep 1: Why a Single ./a.out Time Misleads Your Performance Claims (this article)
  • Ep 2: Warm vs Cold: Why a Single Trial Misleads Performance Claims
  • Ep 3: Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts
  • Ep 4: Comparing Across Machines: Why Metadata Matters in csperf
  • Ep 5: csperf: A Lightweight Observatory for Honest Performance Tracking
  • Ep 6: csperf Doctor: First Command to Diagnose Toolchain Issues
  • Ep 7: csperf Quickstart: 30‑Second First Win
  • Ep 8: Reading csperf JSON: The Real Performance Artifact
  • Ep 9: Backends 101: Choosing the Right Measurement Surface
  • Ep 10: Stability Metrics: Min, Max, and Standard Deviation as First-Class Citizens
  • Ep 11: Warmup Deep Dive: What You’re Throwing Away and Why

Teaser next

Recall how we used metadata to compare machines (Ep 4) and now we’ll dive into locality: Row vs column access and how layout shows up in time (Ep 12).

Top comments (0)