DEV Community

compilersutra
compilersutra

Posted on Originally published at compilersutra.com

Turning on Linux perf Counters with csperf (honestly)

Why this lesson exists

When you run a program under perf, you often see empty counters or wildly different numbers from run to run. The myth is that enabling perf automatically gives you useful data. In reality, missing events, kernel restrictions, or mis‑configured sysctl can silence counters. This episode shows how to turn on perf, validate that it is working, and read the real numbers that csperf surfaces.

Recap — where we are in the series

Ep 1: single‑shot timings are misleading – use csperf for reproducible evidence.
Ep 2: warm‑up and repeat give credible performance data.
Ep 3: screenshots mislead – csperf gives metadata‑rich artifacts.
Ep 4: metadata matters for cross‑machine comparison.
Ep 5: csperf observatory commands give quick evidence.
Ep 6: csperf doctor diagnoses toolchain issues.
Ep 7: quickstart creates a reproducible artifact in 30 s.
Ep 8: reading csperf JSON is the source of truth.
Ep 9: choose the right backend.
Ep 10: min, max, stdev are first‑class citizens.
Ep 11: warm‑up impact quantified.
Ep 12: measure row vs column locality.
Ep 13: validate optimizer claims with real data.

The misconception

Many developers think that running perf stat or enabling perf in csperf will automatically populate all counters. In practice, the kernel may block certain events, or the CPU may not expose them due to security mitigations. The result is an artifact with empty fields, leading to fabricated stories.

What problem csperf solves (this episode's slice)

csperf now supports the perf collector on Linux. It:

  1. Detects whether perf is available.
  2. Configures the right PMU (e.g., amd64_fam19h_zen4).
  3. Validates that counters are collected.
  4. Exposes the data in the JSON artifact.
  5. Provides guidance when counters are missing.

Mental model

Think of perf as a camera that can take multiple shots (events). If the camera is turned off or the lens is blocked, you get no image. csperf turns the camera on, checks the lens, and tells you if you got a picture. The artifact is the photo album.

Lab: install and first commands

# Ensure csperf is up‑to‑date
pip install --upgrade csperf

# Verify perf is available on the machine
csperf observatory --check perf
Enter fullscreen mode Exit fullscreen mode

The observatory will print perf: available and list supported events.

Lab: what we ran on this machine

We ran the matrix traversal example on an AMD Ryzen 7 9700X. The command used:

csperf run \
  --input examples/cpp/matrix_traversal.cpp \
  --backend cpu \
  --warmup-runs 1 \
  --repeat-runs 5 \
  --output results/perf.json
Enter fullscreen mode Exit fullscreen mode

This produced results/perf.json and results/perf.csv.

Results (real numbers only)

Metric Value Unit
execution_time_ms 5.1494 ms
cpu_cycles 6,245,829 cycles
ref_cycles 4,254,898 cycles
frontend_stall_cycles 1,175,631 cycles
instruction_count 12,524,896 instructions
cache_misses 94,609 events
cache_references 566,279 events
l1_cache_misses 255,809 events
l1_icache_misses 32,337 events
dtlb_load_misses 4,606 events
itlb_load_misses 252 events
amd_l2_ic_dc_miss_in_l2 94,218 events
ipc 2.005322

The CSV contains the same numbers plus per‑trial timings:

trial_1_ms,5.153
trial_2_ms,5.150
trial_3_ms,5.165
trial_4_ms,5.140
trial_5_ms,5.139
Enter fullscreen mode Exit fullscreen mode

How to read the artifacts

Open results/perf.json. The top‑level metrics object holds the raw counters. The hardware section confirms that perf was auto‑selected and lists supported_metrics. If a metric is missing, it will not appear. The artifacts section gives paths to CSV and XLSX for quick spreadsheet analysis.

Common mistakes (teacher checklist)

  1. Assuming perf is enabled – run csperf observatory --check perf first.
  2. Ignoring sysctl limits – check /proc/sys/kernel/perf_event_paranoid.
  3. Running without root – some events require elevated privileges.
  4. Misreading empty fields – an empty field means the event was not collected, not that it was zero.
  5. Using the wrong backend – cpu is required for perf; gpu will ignore perf.

Try this next (homework)

Run the same experiment on a different CPU (e.g., Intel i9) and compare the cpu_cycles and cache_misses. Note any differences in supported_metrics. Do this tonight — Episode 15 starts by assuming you did.

Closing

We have shown how to turn on Linux perf counters, validate their presence, and interpret the numbers. Next we will dive into noise sources: affinity, frequency, and background load.

The series so far

  • Ep 1: Why a Single ./a.out Time Misleads Your Performance Claims (this article)
  • Ep 2: Warm vs Cold: Why a Single Trial Misleads Performance Claims
  • Ep 3: Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts
  • Ep 4: Comparing Across Machines: Why Metadata Matters in csperf
  • Ep 5: csperf: A Lightweight Observatory for Honest Performance Tracking
  • Ep 6: csperf Doctor: First Command to Diagnose Toolchain Issues
  • Ep 7: csperf Quickstart: 30‑Second First Win
  • Ep 8: Reading csperf JSON: The Real Performance Artifact
  • Ep 9: Backends 101: Choosing the Right Measurement Surface
  • Ep 10: Stability Metrics: Min, Max, and Standard Deviation as First-Class Citizens
  • Ep 11: Warmup Deep Dive: What You’re Throwing Away and Why
  • Ep 12: Row vs. Column: Measuring Locality with csperf
  • Ep 13: Unmasking the O‑3 Myth: Real‑World Optimizer Impact with csperf
  • Ep 14: Turning on Linux perf Counters with csperf (honestly)

Teaser for next episode

Back to the mystery of empty counters and forward to the noise sources that can skew your numbers.

Top comments (0)