Why this lesson exists
In the first six episodes we learned that a single timing, a screenshot, or a hand‑rolled script can mislead. We built a mental model around reproducibility, metadata, and the observatory pattern. The next logical step is to get a first win: a reproducible, machine‑aware artifact that you can hand to a reviewer or store in a CI pipeline. Episode 7 shows how to do that in thirty seconds.
Recap — where we are in the series
Ep 1 – single‑shot timings are deceptive; use csperf for reproducible evidence.
Ep 2 – warm‑up and repeat runs give credible data.
Ep 3 – screenshots are not evidence; csperf produces metadata‑rich artifacts.
Ep 4 – metadata enables cross‑machine comparison.
Ep 5 – observatory commands give quick, reliable evidence.
Ep 6 – doctor command validates the toolchain.
We have built a foundation of why we need artifacts and how to capture them. Now we turn that theory into practice.
The misconception
Many developers think “performance” is just a single number. In reality it is a bundle of data: execution time, cache behavior, CPU cycles, and the context in which the run happened. A quick artifact that bundles all of that is the real win.
What problem csperf solves (this episode's slice)
csperf quickstart is a one‑liner that:
- Compiles the target file.
- Executes it three times.
- Gathers timing and profiling data.
- Stores everything in JSON, CSV, and XLSX.
- Annotates the artifact with machine metadata.
The result is a single, reproducible artifact you can inspect or ship.
Mental model
Think of csperf quickstart as a micro‑observatory. It sets up a controlled experiment, runs it, and records the evidence. The artifact is the evidence bundle.
artifact ──┬─ timing data
│ ├─ cache misses
│ └─ CPU cycles
└─ metadata
When you look at the artifact you see what happened, where it happened, and how often.
Lab: install and first commands
- Ensure you have a recent Python 3.12 environment.
- Install csperf via pip:
pip install csperf
- Verify installation:
csperf --version
- Run the quickstart command from the root of the project:
csperf quickstart --output-dir results/quickstart
This will create results/quickstart/quickstart.json, quickstart.csv, and quickstart.xlsx.
Lab: what we ran on this machine
The machine is an AMD Ryzen 7 9700X 8‑core (16 threads) running Ubuntu 24.04. The artifact includes:
hostname=f4c59d864117date=2026‑10‑02T18:00:02+05:30csperf_bin=/home/aitr/projects/CompilerSutraPerfTool/.venv/bin/csperf- CPU model:
AMD Ryzen 7 9700X(model 68, 8 cores, 2 threads per core) -
perfandpapiprofilers available.
The target file is matrix_traversal.cpp from the examples folder.
Results (real numbers only)
| Metric | Value | Unit |
|---|---|---|
| Execution time (mean) | 5.159 ms | ms |
| Min | 5.117 ms | ms |
| Max | 5.216 ms | ms |
| Stdev | 0.051 ms | ms |
| Runs | 3 | |
| CPU cores | 16 | |
| Cache misses | 5 × 10⁶ (approx) |
The JSON artifact (quickstart.json) contains the same numbers under metrics.execution_time_ms and metrics.execution_time_summary_ms.
How to read the artifacts
-
JSON – the gold standard. Open it in any editor. Look under
metricsfor timing,hardwarefor machine details, andcommandsfor the exact compile‑run pipeline. - CSV – handy for spreadsheets. The first rows describe the input, the next rows the run, and the final rows the timing summary.
- XLSX – a quick way to plot the timing distribution.
Key fields:
-
execution_time_ms– raw elapsed time. -
execution_time_summary_ms– statistical summary. -
hardware.cpu_model– identifies the CPU. -
hardware.tools– lists available profilers.
Common mistakes (teacher checklist)
-
Running without
--output-dir– the default directory is cluttered. -
Ignoring the
runscount – always confirm you have the expected number of trials. - Assuming the first trial is the best – use the mean or median.
- Missing the machine metadata – without it the artifact is incomplete.
-
Overlooking the profiler availability – if
perfis missing, the artifact will lack low‑level counters.
Try this next (homework)
- Run
csperf quickstarton a different target (e.g.,hello_world.cpp). - Compare the JSON artifacts side‑by‑side.
- Identify any differences in
hardwaremetadata. - Document the impact of compiler flags on
execution_time_ms.
Do this tonight — Episode 8 starts by assuming you did.
Closing
A quick, reproducible artifact is the cornerstone of honest performance measurement. Episode 7 showed you how to generate that artifact in half a minute. Next we’ll learn how to interpret the JSON to make data‑driven decisions.
The series so far
- Ep 1: Why a Single ./a.out Time Misleads Your Performance Claims (this article)
- Ep 2: Warm vs Cold: Why a Single Trial Misleads Performance Claims
- Ep 3: Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts
- Ep 4: Comparing Across Machines: Why Metadata Matters in csperf
- Ep 5: csperf: A Lightweight Observatory for Honest Performance Tracking
- Ep 6: csperf Doctor: First Command to Diagnose Toolchain Issues
- Ep 7: csperf Quickstart: 30‑Second First Win
The series so far (skip Ep1)
- Ep 2: Warm vs Cold
- Ep 3: Screenshots Aren't Evidence
- Ep 4: Comparing Across Machines
- Ep 5: csperf Observatory
- Ep 6: csperf Doctor
- Ep 7: csperf Quickstart
Teaser for next episode
Back in Episode 5 we saw how csperf gives quick observatory data, and next we’ll dive into reading the JSON artifact to extract real insights.
Top comments (0)