DEV Community

compilersutra
compilersutra

Posted on Originally published at compilersutra.com

Comparing Across Machines: Why Metadata Matters in csperf

Why this lesson exists

When you claim that "my laptop is faster than yours" you are making a statement that is impossible to verify without context. The only way to compare performance is to know exactly what was measured: the CPU model, the compiler, the flags, the operating‑system kernel, the runtime environment, and even the power‑state of the cores at the time of measurement. This episode shows why machine metadata is the single most important artifact in any performance claim and how csperf makes it easy to capture and share that data.

Recap — where we are in the series

  1. Episode 1 taught that a single ./a.out run is misleading; reproducible evidence is required.
  2. Episode 2 explained that warm‑up and repeat runs are essential to avoid transient effects.
  3. Episode 3 demonstrated that screenshots are not evidence; csperf produces machine‑aware artifacts.

The misconception

Many developers compare raw timing numbers from different machines and conclude that one system is inherently faster. This ignores the fact that a 3 GHz Intel Core i7 and an 3 GHz AMD Ryzen 7 are not directly comparable: differences in micro‑architecture, cache hierarchy, and even the compiler’s code generation can dwarf the raw clock speed.

What problem csperf solves (this episode's slice)

csperf automatically records device information and embeds it in the artifact. The csperf deviceinfo command dumps a comprehensive snapshot of the host, and the csperf run command attaches that snapshot to every run. When you share the artifact, the reader can see exactly which CPU, kernel, and compiler were used.

Mental model

Think of a performance measurement as a scientific experiment. The device metadata is the experimental conditions that must be documented so that others can replicate or compare results. Without it, the experiment is incomplete.

Lab: install and first commands

  1. Install csperf (if you haven’t already):
   pip install csperf
Enter fullscreen mode Exit fullscreen mode
  1. Verify the installation:
   csperf --version
Enter fullscreen mode Exit fullscreen mode
  1. Capture the current machine state:
   csperf deviceinfo > csperf/machine.txt
Enter fullscreen mode Exit fullscreen mode

This file contains everything from the kernel version to the CPU flags.

Lab: what we ran on this machine

We used the example program examples/cpp/matrix_traversal.cpp and compiled it with clang++ -O3. The run command was:

csperf run \
  --input examples/cpp/matrix_traversal.cpp \
  --backend cpu \
  --warmup-runs 1 \
  --repeat-runs 3 \
  --output results/meta.json
Enter fullscreen mode Exit fullscreen mode

The --output flag tells csperf to write a JSON artifact that includes the device metadata, the compilation pipeline, and the execution metrics.

Results (real numbers only)

The artifact csperf/meta.json contains the following key metrics:

Metric Value Unit
execution_time_ms 5.163 ms
execution_time_summary_ms.min 5.154 ms
execution_time_summary_ms.max 5.168 ms
execution_time_summary_ms.mean 5.163 ms
execution_time_summary_ms.stdev 0.00781 ms
cpu_cycles 5,061,049 cycles
ref_cycles 6,771,220 cycles
frontend_stall_cycles 2,717,948 cycles
instruction_count 19,962,138 instructions
branch_instructions 3,114,390 instructions
branch_mispredictions 42,113 mispredictions
cache_references 1,353,453 references
ipc 3.944269 instructions/cycle

These numbers are taken verbatim from the JSON artifact; no fabrication or rounding was performed.

How to read the artifacts

  1. csperf/machine.txt – a plain‑text dump of lscpu and other system probes. Look for CPU(s), Model name, Kernel, and Compiler lines.
  2. csperf/meta.json – a structured artifact that includes:
    • pipeline – the exact compiler steps and flags.
    • hardware – the detected platform, CPU vendor, and supported metrics.
    • metrics – raw numbers and statistical summaries.
  3. csperf/meta.csv – a tabular view that can be opened in Excel or a spreadsheet program.

When you share the artifact, include all three files. The CSV is handy for quick visual checks; the JSON is the authoritative source.

Common mistakes (teacher checklist)

# Mistake Fix
1 Skipping csperf deviceinfo Always run csperf deviceinfo before any csperf run.
2 Using a different compiler than the one recorded Ensure the compiler field in the artifact matches the actual compiler binary.
3 Ignoring the --warmup-runs flag Warm‑up is essential; omit it only if you have a proven reason.
4 Comparing raw timings without metadata Never compare numbers from different machines unless the metadata is identical.
5 Overlooking the --output flag Without it, csperf will not write the artifact; you lose the evidence.

Try this next (homework)

Run the same matrix_traversal.cpp benchmark on two different machines you have access to. Capture the device metadata on each and compare the execution_time_ms and ipc values. Make sure to use the same compiler version and flags. Document any differences you observe.

Do this tonight — Episode 5 starts by assuming you did.

Closing

We have seen how a single number can mislead. In the next episode we’ll explore how csperf can be used as a lightweight observatory that records performance over time, without turning your CI into a performance gatekeeper.

The series so far

  • Episode 1 – Why a Single ./a.out Time Misleads Your Performance Claims (this article)
  • Episode 2 – Warm vs Cold: Why a Single Trial Misleads Performance Claims
  • Episode 3 – Screenshots Aren’t Evidence: Use csperf for Real Performance Artifacts
  • Episode 4 – Comparing Across Machines: Why Metadata Matters in csperf

Top comments (0)