DEV Community

compilersutra
compilersutra

Posted on Originally published at compilersutra.com

csperf Quickstart: 30‑Second First Win

Why this lesson exists

In the first six episodes we learned that a single timing, a screenshot, or a hand‑rolled script can mislead. We built a mental model around reproducibility, metadata, and the observatory pattern. The next logical step is to get a first win: a reproducible, machine‑aware artifact that you can hand to a reviewer or store in a CI pipeline. Episode 7 shows how to do that in thirty seconds.

Recap — where we are in the series

Ep 1 – single‑shot timings are deceptive; use csperf for reproducible evidence.
Ep 2 – warm‑up and repeat runs give credible data.
Ep 3 – screenshots are not evidence; csperf produces metadata‑rich artifacts.
Ep 4 – metadata enables cross‑machine comparison.
Ep 5 – observatory commands give quick, reliable evidence.
Ep 6 – doctor command validates the toolchain.

We have built a foundation of why we need artifacts and how to capture them. Now we turn that theory into practice.

The misconception

Many developers think “performance” is just a single number. In reality it is a bundle of data: execution time, cache behavior, CPU cycles, and the context in which the run happened. A quick artifact that bundles all of that is the real win.

What problem csperf solves (this episode's slice)

csperf quickstart is a one‑liner that:

  1. Compiles the target file.
  2. Executes it three times.
  3. Gathers timing and profiling data.
  4. Stores everything in JSON, CSV, and XLSX.
  5. Annotates the artifact with machine metadata.

The result is a single, reproducible artifact you can inspect or ship.

Mental model

Think of csperf quickstart as a micro‑observatory. It sets up a controlled experiment, runs it, and records the evidence. The artifact is the evidence bundle.

artifact ──┬─ timing data
          │   ├─ cache misses
          │   └─ CPU cycles
          └─ metadata
Enter fullscreen mode Exit fullscreen mode

When you look at the artifact you see what happened, where it happened, and how often.

Lab: install and first commands

  1. Ensure you have a recent Python 3.12 environment.
  2. Install csperf via pip:
   pip install csperf
Enter fullscreen mode Exit fullscreen mode
  1. Verify installation:
   csperf --version
Enter fullscreen mode Exit fullscreen mode
  1. Run the quickstart command from the root of the project:
   csperf quickstart --output-dir results/quickstart
Enter fullscreen mode Exit fullscreen mode

This will create results/quickstart/quickstart.json, quickstart.csv, and quickstart.xlsx.

Lab: what we ran on this machine

The machine is an AMD Ryzen 7 9700X 8‑core (16 threads) running Ubuntu 24.04. The artifact includes:

  • hostname=f4c59d864117
  • date=2026‑10‑02T18:00:02+05:30
  • csperf_bin=/home/aitr/projects/CompilerSutraPerfTool/.venv/bin/csperf
  • CPU model: AMD Ryzen 7 9700X (model 68, 8 cores, 2 threads per core)
  • perf and papi profilers available.

The target file is matrix_traversal.cpp from the examples folder.

Results (real numbers only)

Metric Value Unit
Execution time (mean) 5.159 ms ms
Min 5.117 ms ms
Max 5.216 ms ms
Stdev 0.051 ms ms
Runs 3
CPU cores 16
Cache misses 5 × 10⁶ (approx)

The JSON artifact (quickstart.json) contains the same numbers under metrics.execution_time_ms and metrics.execution_time_summary_ms.

How to read the artifacts

  1. JSON – the gold standard. Open it in any editor. Look under metrics for timing, hardware for machine details, and commands for the exact compile‑run pipeline.
  2. CSV – handy for spreadsheets. The first rows describe the input, the next rows the run, and the final rows the timing summary.
  3. XLSX – a quick way to plot the timing distribution.

Key fields:

  • execution_time_ms – raw elapsed time.
  • execution_time_summary_ms – statistical summary.
  • hardware.cpu_model – identifies the CPU.
  • hardware.tools – lists available profilers.

Common mistakes (teacher checklist)

  1. Running without --output-dir – the default directory is cluttered.
  2. Ignoring the runs count – always confirm you have the expected number of trials.
  3. Assuming the first trial is the best – use the mean or median.
  4. Missing the machine metadata – without it the artifact is incomplete.
  5. Overlooking the profiler availability – if perf is missing, the artifact will lack low‑level counters.

Try this next (homework)

  1. Run csperf quickstart on a different target (e.g., hello_world.cpp).
  2. Compare the JSON artifacts side‑by‑side.
  3. Identify any differences in hardware metadata.
  4. Document the impact of compiler flags on execution_time_ms.

Do this tonight — Episode 8 starts by assuming you did.

Closing

A quick, reproducible artifact is the cornerstone of honest performance measurement. Episode 7 showed you how to generate that artifact in half a minute. Next we’ll learn how to interpret the JSON to make data‑driven decisions.

The series so far

  • Ep 1: Why a Single ./a.out Time Misleads Your Performance Claims (this article)
  • Ep 2: Warm vs Cold: Why a Single Trial Misleads Performance Claims
  • Ep 3: Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts
  • Ep 4: Comparing Across Machines: Why Metadata Matters in csperf
  • Ep 5: csperf: A Lightweight Observatory for Honest Performance Tracking
  • Ep 6: csperf Doctor: First Command to Diagnose Toolchain Issues
  • Ep 7: csperf Quickstart: 30‑Second First Win

The series so far (skip Ep1)

  • Ep 2: Warm vs Cold
  • Ep 3: Screenshots Aren't Evidence
  • Ep 4: Comparing Across Machines
  • Ep 5: csperf Observatory
  • Ep 6: csperf Doctor
  • Ep 7: csperf Quickstart

Teaser for next episode

Back in Episode 5 we saw how csperf gives quick observatory data, and next we’ll dive into reading the JSON artifact to extract real insights.

Top comments (0)