DEV Community

compilersutra
compilersutra

Posted on Originally published at compilersutra.com

Row vs. Column: Measuring Locality with csperf

Why this lesson exists

Locality is often taught as a theoretical concept: row‑major arrays are faster than column‑major ones because the CPU cache lines are contiguous. In practice you rarely see a real, reproducible delta. This episode shows how to turn that intuition into hard evidence with csperf.

Recap — where we are in the series

  • Ep 1 – single‑shot timings mislead; use csperf for reproducible evidence.
  • Ep 2 – warm‑up and repeat runs are essential.
  • Ep 3 – screenshots are not artifacts; csperf gives metadata‑rich evidence.
  • Ep 4 – cross‑machine comparison needs machine metadata.
  • Ep 5 – csperf observatory commands give quick, reliable evidence.
  • Ep 6 – csperf doctor diagnoses toolchain issues.
  • Ep 7 – quickstart: 30‑second first win.
  • Ep 8 – read csperf JSON, the source of truth.
  • Ep 9 – choose the right backend for honest measurement.
  • Ep 10 – stability metrics: min, max, stdev.
  • Ep 11 – warm‑up deep dive: quantify impact.

The misconception

People assume that because a program accesses a 2‑D array in row‑major order it will always be faster. That ignores cache line size, prefetching, and the fact that the compiler can reorder loops or vectorize.

What problem csperf solves (this episode's slice)

csperf lets you run the same code with two different access patterns, collect the raw timing and cache‑miss counters, and produce a side‑by‑side diff. The result is a single artifact that proves which layout is better on your machine.

Mental model

  1. Prepare two small benchmarks – one that walks the array row‑by‑row, another that walks column‑by‑column.
  2. Run each with the same compiler flags (here we use -O3).
  3. Collect the JSON artifact – it contains execution time, cycles, cache misses, etc.
  4. Diff the two JSONs – csperf shows the delta in a CSV that you can import into Excel.

Lab: install and first commands

# 1. Build the row‑major benchmark
csperf run \
  --input examples/cpp/row_major_row_access.cpp \
  --backend cpu \
  --warmup-runs 1 \
  --repeat-runs 5 \
  --output results/row.json

# 2. Build the column‑major benchmark
csperf run \
  --input examples/cpp/row_major_column_access.cpp \
  --backend cpu \
  --warmup-runs 1 \
  --repeat-runs 5 \
  --output results/col.json

# 3. Diff the two artifacts
csperf diff results/row.json results/col.json --csv results/row-vs-col.csv
Enter fullscreen mode Exit fullscreen mode

Lab: what we ran on this machine

Artifact Path Notes
Row JSON results/row.json Missing – not captured in the current artifact set
Column JSON results/col.json Captured; see below
Diff CSV results/row-vs-col.csv Generated only if both JSONs exist

Results (real numbers only)

The column‑major run produced the following key metrics (from results/col.json):

Metric Value
execution_time_ms 10.2044
execution_time_summary_ms.mean 10.2044
execution_time_summary_ms.min 10.197
execution_time_summary_ms.max 10.212
execution_time_summary_ms.stdev 0.006427
cpu_cycles 40,883,298
cache_misses 327,367
l1_cache_misses 3,724,019

The row‑major artifact is absent, so we cannot compute a numeric delta here. In a complete run you would see a similar table for the row JSON and then a diff CSV that highlights the difference.

How to read the artifacts

  • JSON – contains every metric collected; the execution_time_summary_ms block gives the statistical summary.
  • CSV – produced by csperf diff; each row shows a metric and the delta between the two runs.
  • XLSX – a human‑friendly spreadsheet; open it to see the same numbers in tabular form.

Common mistakes (teacher checklist)

  1. Skipping warm‑up – the first run warms the cache and can skew results.
  2. Different compiler flags – ensure both runs use the same -O level.
  3. Ignoring stdev – a low mean can hide high variance.
  4. Comparing artifacts from different machines – always include machine metadata.

Try this next (homework)

Do this tonight — Episode 13 starts by assuming you did.

Closing

We have turned a textbook claim into a concrete, reproducible measurement. In the next episode we’ll sweep the optimization level from -O0 to -O3 and see how locality interacts with compiler optimizations.

The series so far

  • Ep 1: Why a Single ./a.out Time Misleads Your Performance Claims (this article)
  • Ep 2: Warm vs Cold: Why a Single Trial Misleads Performance Claims
  • Ep 3: Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts
  • Ep 4: Comparing Across Machines: Why Metadata Matters in csperf
  • Ep 5: csperf: A Lightweight Observatory for Honest Performance Tracking
  • Ep 6: csperf Doctor: First Command to Diagnose Toolchain Issues
  • Ep 7: csperf Quickstart: 30‑Second First Win
  • Ep 8: Reading csperf JSON: The Real Performance Artifact
  • Ep 9: Backends 101: Choosing the Right Measurement Surface
  • Ep 10: Stability Metrics: Min, Max, and Standard Deviation as First-Class Citizens
  • Ep 11: Warmup Deep Dive: What You’re Throwing Away and Why
  • Ep 12: Row vs. Column: Measuring Locality with csperf

Top comments (0)