DEV Community

compilersutra
compilersutra

Posted on Originally published at compilersutra.com

Backends 101: Choosing the Right Measurement Surface

Why this lesson exists

When a program runs on a GPU, the only way to know where the real cost lies is to measure that GPU. Defaulting to the CPU surface is a common mistake that can waste days of debugging. This episode shows how to use csperf to discover the right backend and switch when necessary.

Recap — where we are in the series

  • Ep 1: Why a single ./a.out time misleads – we learned that reproducible evidence is essential.
  • Ep 2: Warm vs Cold – a single trial is unreliable; we introduced warm‑up and repeat.
  • Ep 3: Screenshots aren’t evidence – csperf gives metadata‑rich artifacts.
  • Ep 4: Comparing across machines – metadata matters for cross‑machine claims.
  • Ep 5: csperf observatory – quick, reliable evidence of the toolchain.
  • Ep 6: csperf Doctor – diagnose toolchain issues.
  • Ep 7: Quickstart – 30‑second first win.
  • Ep 8: Reading csperf JSON – the source of truth.

The misconception

Many developers assume that a program’s performance is dominated by the CPU, even when the workload is offloaded to a GPU. They run csperf run --backend cpu and interpret the results as the whole story. In reality, the CPU may only be orchestrating kernel launches, while the GPU does the heavy lifting. If you never measure the GPU, you’ll never know where the bottleneck is.

What problem csperf solves (this episode's slice)

csperf gives you a simple, declarative way to list available backends, enumerate devices for a chosen backend, and run a workload on any backend. By inspecting the artifacts you can decide whether to stay on the CPU surface or switch to the GPU surface.

Mental model

  1. List backends – see what execution surfaces csperf knows.
  2. List devices – for a chosen backend, see the concrete hardware.
  3. Run on a backend – execute the workload and collect artifacts.
  4. Inspect artifacts – look at execution_time_ms and other metrics.
  5. Decide – if CPU metrics are trivial compared to GPU metrics, switch.

Lab: install and first commands

# 1. List all backends csperf knows about
csperf list-backends

# 2. For the CPU backend, list the devices that can be used
csperf list-devices --backend cpu

# 3. Run the matrix traversal example on the CPU backend
csperf run \
  --input examples/cpp/matrix_traversal.cpp \
  --backend cpu \
  --output results/cpu.json
Enter fullscreen mode Exit fullscreen mode

The commands above are the same ones you used in Ep 8 when we produced the cpu.json artifact.

Lab: what we ran on this machine

We executed the matrix traversal example on an AMD Ryzen 7 9700X machine with 16 logical CPUs. The machine’s csperf/machine.txt shows the full lscpu output. The cpu.json artifact contains three trials:

Trial Time (ms)
1 5.185
2 5.153
3 5.148

The mean execution time is 5.162 ms with a standard deviation of 0.020 ms.

Results (real numbers only)

From the cpu.json artifact:

  • execution_time_ms: 5.162
  • min_ms: 5.148
  • max_ms: 5.185
  • mean_ms: 5.162
  • stdev_ms: 0.020075

These numbers come straight from the artifact; no fabricated data.

How to read the artifacts

The JSON artifact is a self‑contained record. Key sections:

  • pipeline – shows the compiler steps and the execution backend.
  • metrics – raw counters from the profiler.
  • execution_time_summary_ms – statistical summary of the trials.

If you open cpu.json in a JSON viewer, you’ll see that the backend field is cpu. To switch to a GPU, you would change that field to gpu and rerun.

Common mistakes (teacher checklist)

  1. Assuming CPU metrics are the whole story – always check the backend field.
  2. Running only one trial – use the --repeat flag to get a statistical summary.
  3. Ignoring device enumeration – csperf list-devices tells you which GPUs are available.
  4. Overlooking the profiler – the artifact lists which profiler was used (perf/papi).

Try this next (homework)

Run the same matrix traversal example on the GPU backend (if you have one). Compare the execution_time_ms to the CPU result. Document whether the GPU dominates the cost.

Do this tonight — Episode 10 starts by assuming you did.

Closing

We’ve shown how to avoid the CPU‑only trap by inspecting the backend in the artifact. In the next episode we’ll treat stability metrics—min, max, and stdev—as first‑class citizens, giving you a richer picture of performance variance.

The series so far

  • Ep 1: Why a Single ./a.out Time Misleads Your Performance Claims (this article)
  • Ep 2: Warm vs Cold: Why a Single Trial Misleads Performance Claims
  • Ep 3: Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts
  • Ep 4: Comparing Across Machines: Why Metadata Matters in csperf
  • Ep 5: csperf: A Lightweight Observatory for Honest Performance Tracking
  • Ep 6: csperf Doctor: First Command to Diagnose Toolchain Issues
  • Ep 7: csperf Quickstart: 30‑Second First Win
  • Ep 8: Reading csperf JSON: The Real Performance Artifact
  • Ep 9: Backends 101: Choosing the Right Measurement Surface (this article)

Top comments (0)