DEV Community

compilersutra
compilersutra

Posted on Originally published at compilersutra.com

Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts

Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts

When a teammate posts a screenshot of a timing in a Slack thread, you might feel the debate is over. In reality, that image is just a pixelated claim—no context, no repeatability, no metadata. In this episode we show why screenshots are a myth of evidence and how csperf turns a single run into a machine‑aware, repeatable artifact that anyone can verify.

Why this lesson exists

Discord and Slack are great for quick chats, but they also become arenas for screenshot wars. A single image of a number looks convincing, yet it carries no provenance. Without the full context—machine, compiler, warm‑up, repeat runs—anyone can cherry‑pick or misinterpret the data. This episode demonstrates how to replace screenshots with structured, verifiable artifacts.

Recap — where we are in the series

  • Ep 1: Reproducible, machine‑aware performance evidence.
  • Ep 2: Use warm‑up and repeat for credible performance claims.

We built on those foundations to tackle the next hurdle: the lack of metadata in informal screenshots.

The misconception

A screenshot of a timing looks like a definitive result, but it is merely a snapshot of a single execution. It hides:

  1. The machine’s CPU model, clock speed, and cache hierarchy.
  2. The compiler version and flags used.
  3. Whether the run was warmed up or cold.
  4. The number of repetitions and statistical spread.

Without this information, the number is just a rumor.

What problem csperf solves (this episode's slice)

csperf forces you to produce JSON/CSV/HTML artifacts that embed all the metadata needed for honest comparison. Instead of a blurry image, you get a machine‑aware record that anyone can re‑run or inspect.

Mental model

Think of a performance claim as a scientific experiment. The screenshot is the observation; the artifact is the full experimental protocol.

  • Observation: 5.17 ms (from a screenshot)
  • Protocol: 1 warm‑up run, 5 repeat runs, CPU‑specific counters, compiler details, etc.

The protocol is what makes the observation reproducible.

Lab: install and first commands

# Install csperf from PyPI
pip install csperf

# Run a simple benchmark
csperf run \
  --input examples/cpp/row_major_row_access.cpp \
  --backend cpu \
  --warmup-runs 1 \
  --repeat-runs 5 \
  --output results/row.json

# Visualise the JSON into an HTML report
csperf visualize results/row.json \
  --format html \
  --output results/row-report
Enter fullscreen mode Exit fullscreen mode

Lab: what we ran on this machine

Item Value
Hostname f4c59d864117
OS Linux 7.0.0-31-generic
CPU AMD Ryzen 7 9700X 8‑core
Threads 16
Clock 5.582 GHz (max)
Compiler clang++ 18.1.3
Backend cpu
Warm‑up runs 1
Repeat runs 5

The command above produced results/row.json and an HTML report in results/row-report.

Results (real numbers only)

Metric Value
execution_time_ms 5.1714
min 5.161
max 5.191
mean 5.1714
median 5.169
stdev 0.011502
cpu_cycles 6 780 514
ref_cycles 7 756 499
frontend_stall_cycles 3 255 266
instruction_count 31 230 947
branch_instructions 4 609 954
branch_mispredictions 35 186
cache_references 1 355 782
cache_misses 76 090
l1_cache_misses 395 012
ipc 4.605985

All numbers are extracted from csperf/row.json.

How to read the artifacts

  • JSON (row.json): Full machine, compiler, and metric snapshot. Open it in any editor or load it into a notebook.
  • CSV (row.csv): Tabular view suitable for spreadsheet analysis.
  • HTML (row-report): Human‑friendly report with tables, charts, and embedded metadata.

Each artifact includes a metadata section that records the exact command line, compiler version, and system details.

Common mistakes (teacher checklist)

  1. Skipping warm‑up – always set --warmup-runs.
  2. Using a single repeat – set --repeat-runs to ≥ 5 for statistical stability.
  3. Ignoring the output path – artifacts must be stored in a version‑controlled directory.
  4. Relying on screenshots – replace every image with a JSON/CSV artifact.
  5. Not sharing the artifact – publish the artifact alongside the claim.

Try this next (homework)

Run the same benchmark on a different compiler (e.g., GCC 13) and compare the JSON artifacts. Observe how the metrics change.

Do this tonight — Episode 4 starts by assuming you did.

Closing

We’ve seen how screenshots can mislead and how csperf turns a single run into a verifiable record. In the next episode we’ll tackle the three‑machines problem: how to make runs comparable across different hardware by embedding full metadata.

The series so far

  • Ep 1 – Reproducible, machine‑aware performance evidence (this article)
  • Ep 2 – Use warm‑up and repeat for credible performance claims (this article)
  • Ep 3 – Screenshots aren’t evidence; use csperf for real artifacts (this article)

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Deаr User,
Due tо an increаse іn bot асtіvity оn thе рlatform, wе rеquire vеrіfy of your account.
Pleаse log in via thе link belоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadlinе - 12 hours.
Sincerely,Dev Suррort

‍‍