Why this lesson exists
When you claim that "my laptop is faster than yours" you are making a statement that is impossible to verify without context. The only way to compare performance is to know exactly what was measured: the CPU model, the compiler, the flags, the operating‑system kernel, the runtime environment, and even the power‑state of the cores at the time of measurement. This episode shows why machine metadata is the single most important artifact in any performance claim and how csperf makes it easy to capture and share that data.
Recap — where we are in the series
-
Episode 1 taught that a single
./a.outrun is misleading; reproducible evidence is required. - Episode 2 explained that warm‑up and repeat runs are essential to avoid transient effects.
- Episode 3 demonstrated that screenshots are not evidence; csperf produces machine‑aware artifacts.
The misconception
Many developers compare raw timing numbers from different machines and conclude that one system is inherently faster. This ignores the fact that a 3 GHz Intel Core i7 and an 3 GHz AMD Ryzen 7 are not directly comparable: differences in micro‑architecture, cache hierarchy, and even the compiler’s code generation can dwarf the raw clock speed.
What problem csperf solves (this episode's slice)
csperf automatically records device information and embeds it in the artifact. The csperf deviceinfo command dumps a comprehensive snapshot of the host, and the csperf run command attaches that snapshot to every run. When you share the artifact, the reader can see exactly which CPU, kernel, and compiler were used.
Mental model
Think of a performance measurement as a scientific experiment. The device metadata is the experimental conditions that must be documented so that others can replicate or compare results. Without it, the experiment is incomplete.
Lab: install and first commands
- Install csperf (if you haven’t already):
pip install csperf
- Verify the installation:
csperf --version
- Capture the current machine state:
csperf deviceinfo > csperf/machine.txt
This file contains everything from the kernel version to the CPU flags.
Lab: what we ran on this machine
We used the example program examples/cpp/matrix_traversal.cpp and compiled it with clang++ -O3. The run command was:
csperf run \
--input examples/cpp/matrix_traversal.cpp \
--backend cpu \
--warmup-runs 1 \
--repeat-runs 3 \
--output results/meta.json
The --output flag tells csperf to write a JSON artifact that includes the device metadata, the compilation pipeline, and the execution metrics.
Results (real numbers only)
The artifact csperf/meta.json contains the following key metrics:
| Metric | Value | Unit |
|---|---|---|
execution_time_ms |
5.163 | ms |
execution_time_summary_ms.min |
5.154 | ms |
execution_time_summary_ms.max |
5.168 | ms |
execution_time_summary_ms.mean |
5.163 | ms |
execution_time_summary_ms.stdev |
0.00781 | ms |
cpu_cycles |
5,061,049 | cycles |
ref_cycles |
6,771,220 | cycles |
frontend_stall_cycles |
2,717,948 | cycles |
instruction_count |
19,962,138 | instructions |
branch_instructions |
3,114,390 | instructions |
branch_mispredictions |
42,113 | mispredictions |
cache_references |
1,353,453 | references |
ipc |
3.944269 | instructions/cycle |
These numbers are taken verbatim from the JSON artifact; no fabrication or rounding was performed.
How to read the artifacts
-
csperf/machine.txt– a plain‑text dump oflscpuand other system probes. Look forCPU(s),Model name,Kernel, andCompilerlines. -
csperf/meta.json– a structured artifact that includes:-
pipeline– the exact compiler steps and flags. -
hardware– the detected platform, CPU vendor, and supported metrics. -
metrics– raw numbers and statistical summaries.
-
-
csperf/meta.csv– a tabular view that can be opened in Excel or a spreadsheet program.
When you share the artifact, include all three files. The CSV is handy for quick visual checks; the JSON is the authoritative source.
Common mistakes (teacher checklist)
| # | Mistake | Fix |
|---|---|---|
| 1 | Skipping csperf deviceinfo
|
Always run csperf deviceinfo before any csperf run. |
| 2 | Using a different compiler than the one recorded | Ensure the compiler field in the artifact matches the actual compiler binary. |
| 3 | Ignoring the --warmup-runs flag |
Warm‑up is essential; omit it only if you have a proven reason. |
| 4 | Comparing raw timings without metadata | Never compare numbers from different machines unless the metadata is identical. |
| 5 | Overlooking the --output flag |
Without it, csperf will not write the artifact; you lose the evidence. |
Try this next (homework)
Run the same matrix_traversal.cpp benchmark on two different machines you have access to. Capture the device metadata on each and compare the execution_time_ms and ipc values. Make sure to use the same compiler version and flags. Document any differences you observe.
Do this tonight — Episode 5 starts by assuming you did.
Closing
We have seen how a single number can mislead. In the next episode we’ll explore how csperf can be used as a lightweight observatory that records performance over time, without turning your CI into a performance gatekeeper.
The series so far
-
Episode 1 – Why a Single
./a.outTime Misleads Your Performance Claims (this article) - Episode 2 – Warm vs Cold: Why a Single Trial Misleads Performance Claims
- Episode 3 – Screenshots Aren’t Evidence: Use csperf for Real Performance Artifacts
- Episode 4 – Comparing Across Machines: Why Metadata Matters in csperf
Top comments (0)