Why this lesson exists
The online compiler‑opt community still swears that -O3 is the gold standard. In practice, the extra passes can hurt or help only in narrow cases. This episode shows how to prove or disprove that myth on your own hardware.
Recap — where we are in the series
-
Ep 1 – A single
./a.outtime is misleading; use csperf for reproducible evidence. - Ep 2 – Warm‑up and repeat runs give credible performance data.
- Ep 3 – Screenshots mislead; csperf provides metadata‑rich artifacts.
- Ep 4 – Metadata‑rich artifacts enable honest cross‑machine comparison.
- Ep 5 – csperf’s observatory commands give quick, reliable evidence.
- Ep 6 – csperf doctor validates the toolchain.
- Ep 7 – Quickstart: 30‑second first win.
- Ep 8 – Read csperf JSON, the source of truth.
- Ep 9 – Choose the right backend.
- Ep 10 – Stability metrics: min, max, stdev.
- Ep 11 – Warm‑up deep dive: quantify its impact.
- Ep 12 – Row vs. column locality measured.
The misconception
-O3 is often treated as a silver bullet, but the optimizer’s extra passes can increase code size, alter instruction scheduling, and even trigger de‑optimizations. The real question is: does -O3 actually make my code faster on my CPU?
What problem csperf solves (this episode's slice)
csperf’s --diff-optimize flag runs the same source through multiple optimization levels, collects the same metrics, and produces a side‑by‑side comparison. The report shows where the optimizer gains or loses.
Mental model
source → clang frontend → LLVM IR → native code (O0/O1/O2/O3)
│ │ │
└─> metrics (time, cycles, cache misses) ──> diff report
The diff report is the artifact that tells you which level is best for your workload.
Lab: install and first commands
# Ensure csperf is up to date
pip install --upgrade csperf
# Run the matrix traversal example with a diff
csperf run \
--input examples/cpp/matrix_traversal.cpp \
--diff-optimize \
--no-perf \
--report-format html
The --no-perf flag skips Linux perf counters; we focus on wall‑time and basic metrics.
Lab: what we ran on this machine
| Level | Command | Notes |
|---|---|---|
-O0 |
clang++ -O0 |
Baseline, no optimizations |
-O1 |
clang++ -O1 |
Basic optimizations |
-O2 |
clang++ -O2 |
Aggressive optimizations |
-O3 |
clang++ -O3 |
Full optimization, vectorization |
The machine is an AMD Ryzen 7 9700X (8 cores, 16 threads) running Ubuntu 24.04.
Results (real numbers only)
The only artifact we have in this run is the O0 JSON. It reports:
execution_time_ms: 5.142533
min: 5.129
max: 5.165
mean: 5.142533
stdev: 0.010696
Artifacts for -O1, -O2, and -O3 were not captured in this particular run, so we cannot provide their numbers here.
Artifacts missing for this episode — numbers withheld.
How to read the artifacts
Open the generated HTML report. It contains a table with rows for each metric and columns for each optimization level. The diff column shows the percentage change from -O0. A negative value means the level is faster; a positive value means slower.
If you only have the JSON, you can inspect the metrics section:
{
"execution_time_ms": 5.142533,
"execution_time_summary_ms": {
"count": 15,
"min": 5.129,
"max": 5.165,
"mean": 5.142533,
"median": 5.141,
"stdev": 0.010696
}
}
Common mistakes (teacher checklist)
- Skipping warm‑up – The first run may be slower due to cache misses.
- Ignoring stdev – A low mean can hide high variance.
-
Assuming
-O3is always best – Verify with--diff-optimize. - Using the wrong backend – Ensure you run on the same device for each level.
Try this next (homework)
Run the same experiment on a different compiler (e.g., GCC 13) and compare the results. Document any differences in the diff report.
Do this tonight — Episode 14 starts by assuming you did.
Closing
The --diff-optimize flag turns a black‑box claim into a white‑box fact. By looking at the artifact, you can see exactly where the optimizer helps or hurts.
The series so far
- Ep 1: Why a Single
./a.outTime Misleads Your Performance Claims - Ep 2: Warm vs Cold: Why a Single Trial Misleads Performance Claims
- Ep 3: Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts
- Ep 4: Comparing Across Machines: Why Metadata Matters in csperf
- Ep 5: csperf: A Lightweight Observatory for Honest Performance Tracking
- Ep 6: csperf Doctor: First Command to Diagnose Toolchain Issues
- Ep 7: csperf Quickstart: 30‑Second First Win
- Ep 8: Reading csperf JSON: The Real Performance Artifact
- Ep 9: Backends 101: Choosing the Right Measurement Surface
- Ep 10: Stability Metrics: Min, Max, and Standard Deviation as First-Class Citizens
- Ep 11: Warmup Deep Dive: What You’re Throwing Away and Why
- Ep 12: Row vs. Column: Measuring Locality with csperf
- Ep 13: Unmasking the O‑3 Myth: Real‑World Optimizer Impact with csperf
Next episode hint
Turning on Linux perf counters with csperf (honestly) — perf counters need honesty
Top comments (0)