<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: compilersutra</title>
    <description>The latest articles on DEV Community by compilersutra (@aabhinavg).</description>
    <link>https://dev.to/aabhinavg</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1030092%2Fea9dd738-34c3-4265-9f09-1782daf49d3e.jpeg</url>
      <title>DEV Community: compilersutra</title>
      <link>https://dev.to/aabhinavg</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aabhinavg"/>
    <language>en</language>
    <item>
      <title>csperf Doctor: First Command to Diagnose Toolchain Issues</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Thu, 01 Oct 2026 12:30:08 +0000</pubDate>
      <link>https://dev.to/aabhinavg/csperf-doctor-first-command-to-diagnose-toolchain-issues-5fgk</link>
      <guid>https://dev.to/aabhinavg/csperf-doctor-first-command-to-diagnose-toolchain-issues-5fgk</guid>
      <description>&lt;h2&gt;
  
  
  Why this lesson exists
&lt;/h2&gt;

&lt;p&gt;When a csperf run fails, the default reaction is to blame the perf tool for a bad measurement. In reality, the culprit is often an incomplete or mis‑configured compiler toolchain. The &lt;code&gt;csperf doctor&lt;/code&gt; command is designed to surface those hidden problems before you even start a benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recap — where we are in the series
&lt;/h2&gt;

&lt;p&gt;Episode 1: &lt;em&gt;Why a Single ./a.out Time Misleads Your Performance Claims&lt;/em&gt; – single‑shot timings are unreliable.&lt;br&gt;
Episode 2: &lt;em&gt;Warm vs Cold: Why a Single Trial Misleads Performance Claims&lt;/em&gt; – warm‑up and repeat are essential.&lt;br&gt;
Episode 3: &lt;em&gt;Screenshots Aren’t Evidence: Use csperf for Real Performance Artifacts&lt;/em&gt; – metadata‑rich artifacts beat screenshots.&lt;br&gt;
Episode 4: &lt;em&gt;Comparing Across Machines: Why Metadata Matters in csperf&lt;/em&gt; – machine metadata enables honest cross‑machine comparison.&lt;br&gt;
Episode 5: &lt;em&gt;csperf: A Lightweight Observatory for Honest Performance Tracking&lt;/em&gt; – observatory commands give quick, reliable evidence of your compiler environment.&lt;/p&gt;
&lt;h2&gt;
  
  
  The misconception
&lt;/h2&gt;

&lt;p&gt;Many developers assume that a failing csperf run indicates a problem with the perf tool or the benchmark code. In practice, the failure is almost always due to missing compiler components, incorrect paths, or mismatched versions that prevent the toolchain from building the test harness.&lt;/p&gt;
&lt;h2&gt;
  
  
  What problem csperf solves (this episode's slice)
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;csperf doctor&lt;/code&gt; performs a quick sanity check of the compiler toolchain and the csperf installation. It verifies that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The compiler binary (clang, gcc, etc.) is reachable and executable.&lt;/li&gt;
&lt;li&gt;The required LLVM/Clang libraries are present.&lt;/li&gt;
&lt;li&gt;The csperf Python package can import its dependencies.&lt;/li&gt;
&lt;li&gt;The environment variables used by csperf are set correctly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If any of these checks fail, the output clearly indicates the missing piece, saving you from chasing perf‑related noise.&lt;/p&gt;
&lt;h2&gt;
  
  
  Mental model
&lt;/h2&gt;

&lt;p&gt;Think of &lt;code&gt;csperf doctor&lt;/code&gt; as a health checkup for your performance measurement stack. Just as a doctor checks vital signs before a physical exam, this command checks the health of your compiler environment before you run any benchmarks.&lt;/p&gt;
&lt;h2&gt;
  
  
  Lab: install and first commands
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Ensure your virtual environment is activated:
&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   &lt;span class="nb"&gt;source&lt;/span&gt; /home/aitr/projects/CompilerSutraPerfTool/.venv/bin/activate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;ol&gt;
&lt;li&gt;Run the basic doctor check:
&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   csperf doctor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;ol&gt;
&lt;li&gt;If you see any missing components, install them with:
&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   csperf doctor &lt;span class="nt"&gt;--install&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h2&gt;
  
  
  Lab: what we ran on this machine
&lt;/h2&gt;

&lt;p&gt;The machine on which we ran the lab is a Ryzen 7 9700X 8‑core (16 threads) running Ubuntu 24.04. The relevant machine metadata is captured in &lt;code&gt;csperf/machine.txt&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;hostname=f4c59d864117
CPU(s): 16
On‑line CPU(s) list: 0‑15
Vendor ID: AuthenticAMD
Model name: AMD Ryzen 7 9700X 8‑Core Processor
CPU max MHz: 5582.3008
CPU min MHz: 605.3100
BogoMIPS: 7599.98
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;csperf doctor&lt;/code&gt; command was executed from the same environment and produced a &lt;code&gt;csperf/doctor.txt&lt;/code&gt; artifact.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results (real numbers only)
&lt;/h2&gt;

&lt;p&gt;All checks in &lt;code&gt;csperf/doctor.txt&lt;/code&gt; returned &lt;code&gt;[ok]&lt;/code&gt;. No errors were reported. The machine metadata shows 16 logical CPUs and a maximum frequency of 5.58 GHz, which matches the expected performance baseline for this workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to read the artifacts
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;csperf/machine.txt&lt;/strong&gt; – contains the raw system information captured by csperf.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;csperf/file‑list.txt&lt;/strong&gt; – lists all artifacts generated during the run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;csperf/doctor.txt&lt;/strong&gt; – lists each sanity check with its status. Lines beginning with &lt;code&gt;[ok]&lt;/code&gt; indicate a passed check; &lt;code&gt;[missing]&lt;/code&gt; or &lt;code&gt;[error]&lt;/code&gt; flag a problem.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When you open &lt;code&gt;doctor.txt&lt;/code&gt;, look for sections labeled &lt;code&gt;Compiler&lt;/code&gt;, &lt;code&gt;LLVM&lt;/code&gt;, &lt;code&gt;csperf&lt;/code&gt;, and &lt;code&gt;Environment&lt;/code&gt;. Each section will have a list of checks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes (teacher checklist)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Skipping the doctor check&lt;/strong&gt; – always run &lt;code&gt;csperf doctor&lt;/code&gt; before any benchmark.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Using the wrong virtual environment&lt;/strong&gt; – ensure the csperf virtualenv is activated.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Missing system dependencies&lt;/strong&gt; – e.g., &lt;code&gt;libclang-dev&lt;/code&gt; on Debian/Ubuntu.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Incorrect PATH&lt;/strong&gt; – the compiler binary must be in the PATH.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Out‑of‑date csperf&lt;/strong&gt; – upgrade with &lt;code&gt;pip install --upgrade csperf&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Try this next (homework)
&lt;/h2&gt;

&lt;p&gt;Run &lt;code&gt;csperf doctor&lt;/code&gt; on a different machine (e.g., a cloud instance) and compare the output to the local machine. Note any differences in the checks that pass or fail. Do this tonight — Episode 7 starts by assuming you did.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;csperf doctor&lt;/code&gt; is a lightweight, zero‑cost way to catch toolchain issues before they masquerade as perf failures. By incorporating it into your workflow, you eliminate a common source of noise and make your performance claims truly honest.&lt;/p&gt;

&lt;h2&gt;
  
  
  The series so far
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Episode 1 – Why a Single ./a.out Time Misleads Your Performance Claims&lt;/li&gt;
&lt;li&gt;Episode 2 – Warm vs Cold: Why a Single Trial Misleads Performance Claims&lt;/li&gt;
&lt;li&gt;Episode 3 – Screenshots Aren’t Evidence: Use csperf for Real Performance Artifacts&lt;/li&gt;
&lt;li&gt;Episode 4 – Comparing Across Machines: Why Metadata Matters in csperf&lt;/li&gt;
&lt;li&gt;Episode 5 – csperf: A Lightweight Observatory for Honest Performance Tracking&lt;/li&gt;
&lt;li&gt;Episode 6 – csperf Doctor: First Command to Diagnose Toolchain Issues &lt;em&gt;(this article)&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Teaser
&lt;/h2&gt;

&lt;p&gt;Remember how csperf’s observatory commands give quick evidence (Episode 5), and next we’ll dive into the 30‑second quickstart win (Episode 7).&lt;/p&gt;

</description>
      <category>csperf</category>
      <category>tooling</category>
      <category>compilers</category>
      <category>linux</category>
    </item>
    <item>
      <title>csperf: A Lightweight Observatory for Honest Performance Tracking</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Tue, 29 Sep 2026 12:30:08 +0000</pubDate>
      <link>https://dev.to/aabhinavg/csperf-a-lightweight-observatory-for-honest-performance-tracking-5eaa</link>
      <guid>https://dev.to/aabhinavg/csperf-a-lightweight-observatory-for-honest-performance-tracking-5eaa</guid>
      <description>&lt;h2&gt;
  
  
  Why this lesson exists
&lt;/h2&gt;

&lt;p&gt;In the first four episodes we learned that a single timing, a single trial, or a screenshot can mislead.  We also saw how metadata lets us compare across machines.  Yet many teams still rely on bloated CI dashboards that hide noise behind fancy graphs.  This episode shows how a tiny, focused observatory—&lt;code&gt;csperf&lt;/code&gt;—can give you the same honest evidence without the overhead.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recap — where we are in the series
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ep 1&lt;/strong&gt; – &lt;em&gt;Why a Single ./a.out Time Misleads Your Performance Claims&lt;/em&gt;: single‑shot timings are unreliable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ep 2&lt;/strong&gt; – &lt;em&gt;Warm vs Cold: Why a Single Trial Misleads Performance Claims&lt;/em&gt;: warm‑up and repeat runs are essential.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ep 3&lt;/strong&gt; – &lt;em&gt;Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts&lt;/em&gt;: screenshots lack reproducibility.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ep 4&lt;/strong&gt; – &lt;em&gt;Comparing Across Machines: Why Metadata Matters in csperf&lt;/em&gt;: metadata makes cross‑machine comparison honest.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The misconception
&lt;/h2&gt;

&lt;p&gt;Some developers think a small tool like &lt;code&gt;csperf&lt;/code&gt; is just a toy or that it can replace full CI pipelines.  The truth is that &lt;code&gt;csperf&lt;/code&gt; is a &lt;em&gt;tool&lt;/em&gt;—not a replacement for CI, but a lightweight observatory that surfaces the raw evidence you need before you build a pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  What problem csperf solves (this episode's slice)
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;csperf&lt;/code&gt; gives you:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A quick sanity check of your environment (&lt;code&gt;csperf doctor&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;A list of backends you can run (&lt;code&gt;csperf list-backends&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;A machine snapshot (&lt;code&gt;csperf machine.txt&lt;/code&gt;) that you can attach to any measurement.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;These are the first steps in a disciplined performance workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mental model
&lt;/h2&gt;

&lt;p&gt;Think of &lt;code&gt;csperf&lt;/code&gt; as a &lt;em&gt;weather station&lt;/em&gt; for your compiler experiments.  It records the &lt;em&gt;temperature&lt;/em&gt; (machine state) and &lt;em&gt;wind speed&lt;/em&gt; (CPU frequency) before you take a &lt;em&gt;rain gauge&lt;/em&gt; (actual measurement).  Without the station you can’t tell if a sudden drop in performance is due to a storm or a change in the weather.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lab: install and first commands
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Help&lt;/strong&gt; – see what the tool can do:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   csperf &lt;span class="nt"&gt;--help&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Output shows available sub‑commands and flags.&lt;/em&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;List backends&lt;/strong&gt; – discover what you can benchmark:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   csperf list-backends
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The tool writes &lt;code&gt;csperf/backends.txt&lt;/code&gt; and prints the names on stdout.&lt;/em&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Doctor&lt;/strong&gt; – run a sanity check of the environment:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   csperf doctor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The tool writes &lt;code&gt;csperf/doctor.txt&lt;/code&gt; and exits with 0 if everything looks good.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Lab: what we ran on this machine
&lt;/h2&gt;

&lt;p&gt;We executed the three commands above on a machine with the following snapshot (excerpt from &lt;code&gt;csperf/machine.txt&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;hostname=f4c59d864117
CPU(s)=16
Thread(s) per core=2
Core(s) per socket=8
Socket(s)=1
Frequency boost=enabled
CPU scaling MHz=74%
CPU max MHz=5582.3008
CPU min MHz=605.3100
BogoMIPS=7599.98
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The artifact files were written to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/home/aitr/compilersutra/jenkins_pipeline/artifacts/build-23/csperf/backends.txt
/home/aitr/compilersutra/jenkins_pipeline/artifacts/build-23/csperf/doctor.txt
/home/aitr/compilersutra/jenkins_pipeline/artifacts/build-23/csperf/machine.txt
/home/aitr/compilersutra/jenkins_pipeline/artifacts/build-23/csperf/file-list.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Results (real numbers only)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Machine snapshot&lt;/strong&gt; shows 16 logical CPUs, 8 cores, 2 threads per core, and a 74 % scaling factor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Doctor&lt;/strong&gt; produced no errors; the exit code was 0.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backends&lt;/strong&gt; list contains the backends available on this host (the exact names are in &lt;code&gt;backends.txt&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These numbers are the foundation for any subsequent measurement.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to read the artifacts
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;What it contains&lt;/th&gt;
&lt;th&gt;How to use it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;machine.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Full machine metadata&lt;/td&gt;
&lt;td&gt;Attach to any measurement to contextualize results&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;doctor.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Diagnostic log&lt;/td&gt;
&lt;td&gt;Verify environment before measurement&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;backends.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;List of compiler backends&lt;/td&gt;
&lt;td&gt;Pick a backend for &lt;code&gt;csperf run&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;file-list.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inventory of all artifacts&lt;/td&gt;
&lt;td&gt;Useful for reproducibility scripts&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Open each file in a text editor or &lt;code&gt;less&lt;/code&gt; to inspect the raw data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes (teacher checklist)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Skipping &lt;code&gt;csperf doctor&lt;/code&gt;&lt;/strong&gt; – you’ll get noisy data if the environment is mis‑configured.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assuming &lt;code&gt;csperf list-backends&lt;/code&gt; output is exhaustive&lt;/strong&gt; – it only lists what the tool can find locally.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ignoring the machine snapshot&lt;/strong&gt; – without it you can’t compare across runs or machines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treating &lt;code&gt;csperf&lt;/code&gt; as a full CI replacement&lt;/strong&gt; – it is an observatory, not a pipeline.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Try this next (homework)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Run &lt;code&gt;csperf doctor&lt;/code&gt; again and inspect &lt;code&gt;doctor.txt&lt;/code&gt; for any warnings.&lt;/li&gt;
&lt;li&gt;Use &lt;code&gt;csperf list-backends&lt;/code&gt; to pick a backend you want to benchmark.&lt;/li&gt;
&lt;li&gt;Prepare a simple &lt;code&gt;hello.c&lt;/code&gt; program and run a measurement with &lt;code&gt;csperf run --backend &amp;lt;name&amp;gt; hello.c&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Do this tonight — Episode 6 starts by assuming you did.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;We have moved from single‑shot timings to a lightweight observatory that records the environment.  In the next episode we’ll dive into the first measurement command and see how &lt;code&gt;csperf doctor&lt;/code&gt; ensures the data you collect is trustworthy.&lt;/p&gt;

&lt;h2&gt;
  
  
  The series so far
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ep 1&lt;/strong&gt; – Why a Single ./a.out Time Misleads Your Performance Claims &lt;em&gt;(this article)&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ep 2&lt;/strong&gt; – Warm vs Cold: Why a Single Trial Misleads Performance Claims&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ep 3&lt;/strong&gt; – Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ep 4&lt;/strong&gt; – Comparing Across Machines: Why Metadata Matters in csperf&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ep 5&lt;/strong&gt; – csperf: A Lightweight Observatory for Honest Performance Tracking&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>csperf</category>
      <category>compilers</category>
      <category>performance</category>
      <category>tooling</category>
    </item>
    <item>
      <title>Comparing Across Machines: Why Metadata Matters in csperf</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Mon, 28 Sep 2026 12:30:09 +0000</pubDate>
      <link>https://dev.to/aabhinavg/comparing-across-machines-why-metadata-matters-in-csperf-51ld</link>
      <guid>https://dev.to/aabhinavg/comparing-across-machines-why-metadata-matters-in-csperf-51ld</guid>
      <description>&lt;h2&gt;
  
  
  Why this lesson exists
&lt;/h2&gt;

&lt;p&gt;When you claim that &lt;em&gt;"my laptop is faster than yours"&lt;/em&gt; you are making a statement that is impossible to verify without context.  The only way to compare performance is to know &lt;strong&gt;exactly&lt;/strong&gt; what was measured: the CPU model, the compiler, the flags, the operating‑system kernel, the runtime environment, and even the power‑state of the cores at the time of measurement.  This episode shows why machine metadata is the single most important artifact in any performance claim and how csperf makes it easy to capture and share that data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recap — where we are in the series
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Episode 1&lt;/strong&gt; taught that a single &lt;code&gt;./a.out&lt;/code&gt; run is misleading; reproducible evidence is required.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Episode 2&lt;/strong&gt; explained that warm‑up and repeat runs are essential to avoid transient effects.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Episode 3&lt;/strong&gt; demonstrated that screenshots are not evidence; csperf produces machine‑aware artifacts.
&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The misconception
&lt;/h2&gt;

&lt;p&gt;Many developers compare raw timing numbers from different machines and conclude that one system is inherently faster.  This ignores the fact that a 3 GHz Intel Core i7 and an 3 GHz AMD Ryzen 7 are not directly comparable: differences in micro‑architecture, cache hierarchy, and even the compiler’s code generation can dwarf the raw clock speed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What problem csperf solves (this episode's slice)
&lt;/h2&gt;

&lt;p&gt;csperf automatically records &lt;strong&gt;device information&lt;/strong&gt; and embeds it in the artifact.  The &lt;code&gt;csperf deviceinfo&lt;/code&gt; command dumps a comprehensive snapshot of the host, and the &lt;code&gt;csperf run&lt;/code&gt; command attaches that snapshot to every run.  When you share the artifact, the reader can see &lt;em&gt;exactly&lt;/em&gt; which CPU, kernel, and compiler were used.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mental model
&lt;/h2&gt;

&lt;p&gt;Think of a performance measurement as a &lt;em&gt;scientific experiment&lt;/em&gt;.  The &lt;strong&gt;device metadata&lt;/strong&gt; is the &lt;em&gt;experimental conditions&lt;/em&gt; that must be documented so that others can replicate or compare results.  Without it, the experiment is incomplete.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lab: install and first commands
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Install csperf (if you haven’t already):
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   pip &lt;span class="nb"&gt;install &lt;/span&gt;csperf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;Verify the installation:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   csperf &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;Capture the current machine state:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   csperf deviceinfo &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; csperf/machine.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This file contains everything from the kernel version to the CPU flags.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lab: what we ran on this machine
&lt;/h2&gt;

&lt;p&gt;We used the example program &lt;code&gt;examples/cpp/matrix_traversal.cpp&lt;/code&gt; and compiled it with &lt;code&gt;clang++ -O3&lt;/code&gt;.  The run command was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;csperf run &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--input&lt;/span&gt; examples/cpp/matrix_traversal.cpp &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--backend&lt;/span&gt; cpu &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--warmup-runs&lt;/span&gt; 1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--repeat-runs&lt;/span&gt; 3 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt; results/meta.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;--output&lt;/code&gt; flag tells csperf to write a JSON artifact that includes the device metadata, the compilation pipeline, and the execution metrics.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results (real numbers only)
&lt;/h2&gt;

&lt;p&gt;The artifact &lt;code&gt;csperf/meta.json&lt;/code&gt; contains the following key metrics:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;Unit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;execution_time_ms&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;5.163&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;execution_time_summary_ms.min&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5.154&lt;/td&gt;
&lt;td&gt;ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;execution_time_summary_ms.max&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5.168&lt;/td&gt;
&lt;td&gt;ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;execution_time_summary_ms.mean&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5.163&lt;/td&gt;
&lt;td&gt;ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;execution_time_summary_ms.stdev&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.00781&lt;/td&gt;
&lt;td&gt;ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cpu_cycles&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5,061,049&lt;/td&gt;
&lt;td&gt;cycles&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ref_cycles&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;6,771,220&lt;/td&gt;
&lt;td&gt;cycles&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;frontend_stall_cycles&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2,717,948&lt;/td&gt;
&lt;td&gt;cycles&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;instruction_count&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;19,962,138&lt;/td&gt;
&lt;td&gt;instructions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;branch_instructions&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3,114,390&lt;/td&gt;
&lt;td&gt;instructions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;branch_mispredictions&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;42,113&lt;/td&gt;
&lt;td&gt;mispredictions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cache_references&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1,353,453&lt;/td&gt;
&lt;td&gt;references&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ipc&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3.944269&lt;/td&gt;
&lt;td&gt;instructions/cycle&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These numbers are taken verbatim from the JSON artifact; no fabrication or rounding was performed.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to read the artifacts
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;csperf/machine.txt&lt;/code&gt;&lt;/strong&gt; – a plain‑text dump of &lt;code&gt;lscpu&lt;/code&gt; and other system probes.  Look for &lt;code&gt;CPU(s)&lt;/code&gt;, &lt;code&gt;Model name&lt;/code&gt;, &lt;code&gt;Kernel&lt;/code&gt;, and &lt;code&gt;Compiler&lt;/code&gt; lines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;csperf/meta.json&lt;/code&gt;&lt;/strong&gt; – a structured artifact that includes:

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;pipeline&lt;/code&gt; – the exact compiler steps and flags.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;hardware&lt;/code&gt; – the detected platform, CPU vendor, and supported metrics.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;metrics&lt;/code&gt; – raw numbers and statistical summaries.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;csperf/meta.csv&lt;/code&gt;&lt;/strong&gt; – a tabular view that can be opened in Excel or a spreadsheet program.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When you share the artifact, include all three files.  The CSV is handy for quick visual checks; the JSON is the authoritative source.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes (teacher checklist)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Mistake&lt;/th&gt;
&lt;th&gt;Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Skipping &lt;code&gt;csperf deviceinfo&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Always run &lt;code&gt;csperf deviceinfo&lt;/code&gt; before any &lt;code&gt;csperf run&lt;/code&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Using a different compiler than the one recorded&lt;/td&gt;
&lt;td&gt;Ensure the &lt;code&gt;compiler&lt;/code&gt; field in the artifact matches the actual compiler binary.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Ignoring the &lt;code&gt;--warmup-runs&lt;/code&gt; flag&lt;/td&gt;
&lt;td&gt;Warm‑up is essential; omit it only if you have a proven reason.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Comparing raw timings without metadata&lt;/td&gt;
&lt;td&gt;Never compare numbers from different machines unless the metadata is identical.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Overlooking the &lt;code&gt;--output&lt;/code&gt; flag&lt;/td&gt;
&lt;td&gt;Without it, csperf will not write the artifact; you lose the evidence.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Try this next (homework)
&lt;/h2&gt;

&lt;p&gt;Run the same &lt;code&gt;matrix_traversal.cpp&lt;/code&gt; benchmark on &lt;strong&gt;two&lt;/strong&gt; different machines you have access to.  Capture the device metadata on each and compare the &lt;code&gt;execution_time_ms&lt;/code&gt; and &lt;code&gt;ipc&lt;/code&gt; values.  Make sure to use the same compiler version and flags.  Document any differences you observe.&lt;/p&gt;

&lt;p&gt;Do this tonight — Episode 5 starts by assuming you did.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;We have seen how a single number can mislead.  In the next episode we’ll explore how csperf can be used as a lightweight observatory that records performance over time, without turning your CI into a performance gatekeeper.&lt;/p&gt;

&lt;h2&gt;
  
  
  The series so far
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Episode 1&lt;/strong&gt; – Why a Single &lt;code&gt;./a.out&lt;/code&gt; Time Misleads Your Performance Claims &lt;em&gt;(this article)&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Episode 2&lt;/strong&gt; – Warm vs Cold: Why a Single Trial Misleads Performance Claims&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Episode 3&lt;/strong&gt; – Screenshots Aren’t Evidence: Use csperf for Real Performance Artifacts&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Episode 4&lt;/strong&gt; – Comparing Across Machines: Why Metadata Matters in csperf&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>performance</category>
      <category>hardware</category>
      <category>compilers</category>
      <category>csperf</category>
    </item>
    <item>
      <title>Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Sun, 27 Sep 2026 22:41:44 +0000</pubDate>
      <link>https://dev.to/aabhinavg/screenshots-arent-evidence-use-csperf-for-real-performance-artifacts-4db6</link>
      <guid>https://dev.to/aabhinavg/screenshots-arent-evidence-use-csperf-for-real-performance-artifacts-4db6</guid>
      <description>&lt;h1&gt;
  
  
  Screenshots Aren't Evidence: Use csperf for Real Performance Artifacts
&lt;/h1&gt;

&lt;p&gt;When a teammate posts a screenshot of a timing in a Slack thread, you might feel the debate is over. In reality, that image is just a pixelated claim—no context, no repeatability, no metadata. In this episode we show why screenshots are a myth of evidence and how csperf turns a single run into a machine‑aware, repeatable artifact that anyone can verify.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this lesson exists
&lt;/h2&gt;

&lt;p&gt;Discord and Slack are great for quick chats, but they also become arenas for &lt;em&gt;screenshot wars&lt;/em&gt;. A single image of a number looks convincing, yet it carries no provenance. Without the full context—machine, compiler, warm‑up, repeat runs—anyone can cherry‑pick or misinterpret the data. This episode demonstrates how to replace screenshots with structured, verifiable artifacts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recap — where we are in the series
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ep 1:&lt;/strong&gt; Reproducible, machine‑aware performance evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ep 2:&lt;/strong&gt; Use warm‑up and repeat for credible performance claims.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We built on those foundations to tackle the next hurdle: the lack of metadata in informal screenshots.&lt;/p&gt;

&lt;h2&gt;
  
  
  The misconception
&lt;/h2&gt;

&lt;p&gt;A screenshot of a timing looks like a definitive result, but it is merely a &lt;em&gt;snapshot&lt;/em&gt; of a single execution. It hides:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The machine’s CPU model, clock speed, and cache hierarchy.&lt;/li&gt;
&lt;li&gt;The compiler version and flags used.&lt;/li&gt;
&lt;li&gt;Whether the run was warmed up or cold.&lt;/li&gt;
&lt;li&gt;The number of repetitions and statistical spread.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Without this information, the number is just a rumor.&lt;/p&gt;

&lt;h2&gt;
  
  
  What problem csperf solves (this episode's slice)
&lt;/h2&gt;

&lt;p&gt;csperf forces you to produce JSON/CSV/HTML artifacts that embed all the metadata needed for honest comparison. Instead of a blurry image, you get a machine‑aware record that anyone can re‑run or inspect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mental model
&lt;/h2&gt;

&lt;p&gt;Think of a performance claim as a &lt;em&gt;scientific experiment&lt;/em&gt;. The screenshot is the &lt;em&gt;observation&lt;/em&gt;; the artifact is the &lt;em&gt;full experimental protocol&lt;/em&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Observation&lt;/strong&gt;: 5.17 ms (from a screenshot)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Protocol&lt;/strong&gt;: 1 warm‑up run, 5 repeat runs, CPU‑specific counters, compiler details, etc.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The protocol is what makes the observation reproducible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lab: install and first commands
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install csperf from PyPI&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;csperf

&lt;span class="c"&gt;# Run a simple benchmark&lt;/span&gt;
csperf run &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--input&lt;/span&gt; examples/cpp/row_major_row_access.cpp &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--backend&lt;/span&gt; cpu &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--warmup-runs&lt;/span&gt; 1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--repeat-runs&lt;/span&gt; 5 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt; results/row.json

&lt;span class="c"&gt;# Visualise the JSON into an HTML report&lt;/span&gt;
csperf visualize results/row.json &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--format&lt;/span&gt; html &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt; results/row-report
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Lab: what we ran on this machine
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hostname&lt;/td&gt;
&lt;td&gt;f4c59d864117&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OS&lt;/td&gt;
&lt;td&gt;Linux 7.0.0-31-generic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU&lt;/td&gt;
&lt;td&gt;AMD Ryzen 7 9700X 8‑core&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Threads&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clock&lt;/td&gt;
&lt;td&gt;5.582 GHz (max)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compiler&lt;/td&gt;
&lt;td&gt;clang++ 18.1.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backend&lt;/td&gt;
&lt;td&gt;cpu&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Warm‑up runs&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repeat runs&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The command above produced &lt;code&gt;results/row.json&lt;/code&gt; and an HTML report in &lt;code&gt;results/row-report&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results (real numbers only)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;execution_time_ms&lt;/td&gt;
&lt;td&gt;5.1714&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;min&lt;/td&gt;
&lt;td&gt;5.161&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;max&lt;/td&gt;
&lt;td&gt;5.191&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mean&lt;/td&gt;
&lt;td&gt;5.1714&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;median&lt;/td&gt;
&lt;td&gt;5.169&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;stdev&lt;/td&gt;
&lt;td&gt;0.011502&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cpu_cycles&lt;/td&gt;
&lt;td&gt;6 780 514&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ref_cycles&lt;/td&gt;
&lt;td&gt;7 756 499&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;frontend_stall_cycles&lt;/td&gt;
&lt;td&gt;3 255 266&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;instruction_count&lt;/td&gt;
&lt;td&gt;31 230 947&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;branch_instructions&lt;/td&gt;
&lt;td&gt;4 609 954&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;branch_mispredictions&lt;/td&gt;
&lt;td&gt;35 186&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cache_references&lt;/td&gt;
&lt;td&gt;1 355 782&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cache_misses&lt;/td&gt;
&lt;td&gt;76 090&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;l1_cache_misses&lt;/td&gt;
&lt;td&gt;395 012&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ipc&lt;/td&gt;
&lt;td&gt;4.605985&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All numbers are extracted from &lt;code&gt;csperf/row.json&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to read the artifacts
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;JSON&lt;/strong&gt; (&lt;code&gt;row.json&lt;/code&gt;): Full machine, compiler, and metric snapshot. Open it in any editor or load it into a notebook.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CSV&lt;/strong&gt; (&lt;code&gt;row.csv&lt;/code&gt;): Tabular view suitable for spreadsheet analysis.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;HTML&lt;/strong&gt; (&lt;code&gt;row-report&lt;/code&gt;): Human‑friendly report with tables, charts, and embedded metadata.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each artifact includes a &lt;code&gt;metadata&lt;/code&gt; section that records the exact command line, compiler version, and system details.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes (teacher checklist)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Skipping warm‑up&lt;/strong&gt; – always set &lt;code&gt;--warmup-runs&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Using a single repeat&lt;/strong&gt; – set &lt;code&gt;--repeat-runs&lt;/code&gt; to ≥ 5 for statistical stability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ignoring the output path&lt;/strong&gt; – artifacts must be stored in a version‑controlled directory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relying on screenshots&lt;/strong&gt; – replace every image with a JSON/CSV artifact.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not sharing the artifact&lt;/strong&gt; – publish the artifact alongside the claim.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Try this next (homework)
&lt;/h2&gt;

&lt;p&gt;Run the same benchmark on a different compiler (e.g., GCC 13) and compare the JSON artifacts. Observe how the metrics change.&lt;/p&gt;

&lt;p&gt;Do this tonight — Episode 4 starts by assuming you did.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;We’ve seen how screenshots can mislead and how csperf turns a single run into a verifiable record. In the next episode we’ll tackle the &lt;em&gt;three‑machines problem&lt;/em&gt;: how to make runs comparable across different hardware by embedding full metadata.&lt;/p&gt;

&lt;h2&gt;
  
  
  The series so far
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ep 1&lt;/strong&gt; – Reproducible, machine‑aware performance evidence &lt;em&gt;(this article)&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ep 2&lt;/strong&gt; – Use warm‑up and repeat for credible performance claims &lt;em&gt;(this article)&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ep 3&lt;/strong&gt; – Screenshots aren’t evidence; use csperf for real artifacts &lt;em&gt;(this article)&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>performance</category>
      <category>compilers</category>
      <category>json</category>
      <category>csperf</category>
    </item>
    <item>
      <title>Warm vs Cold: Why a Single Trial Misleads Performance Claims</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Sat, 26 Sep 2026 13:51:28 +0000</pubDate>
      <link>https://dev.to/aabhinavg/warm-vs-cold-why-a-single-trial-misleads-performance-claims-5459</link>
      <guid>https://dev.to/aabhinavg/warm-vs-cold-why-a-single-trial-misleads-performance-claims-5459</guid>
      <description>&lt;h2&gt;
  
  
  Why this lesson exists
&lt;/h2&gt;

&lt;p&gt;Performance claims that rely on a single run of a binary are like a snapshot of a moving train – you see a moment, but you miss the whole journey. In the first episode we warned that a one‑shot timing can be wildly off because of warm‑up effects, OS noise, or just a lucky cache state.  &lt;/p&gt;

&lt;p&gt;This episode digs into the &lt;em&gt;warm‑up trap&lt;/em&gt; and shows how &lt;code&gt;csperf&lt;/code&gt; turns a single measurement into a statistically sound artifact.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recap — where we are in the series
&lt;/h2&gt;

&lt;p&gt;Ep 1: &lt;em&gt;Why a Single ./a.out Time Misleads Your Performance Claims&lt;/em&gt; – we saw that a single run can be 10‑30 % off the true steady‑state performance and that reproducibility is impossible without a repeatable protocol.&lt;/p&gt;

&lt;p&gt;Ep 2: &lt;em&gt;Warm vs Cold: Why a Single Trial Misleads Performance Claims&lt;/em&gt; – we are now looking at how to structure a run so that the first few executions warm the cache, branch predictor, and other micro‑architectural state, and how to capture the resulting data.&lt;/p&gt;

&lt;h2&gt;
  
  
  The misconception
&lt;/h2&gt;

&lt;p&gt;Many developers think “run once, read the number, publish.”  The problem is that the first execution of a program after compilation is usually &lt;em&gt;cold&lt;/em&gt;: the code is fetched from disk, the instruction cache is empty, the data cache is cold, and the CPU’s micro‑architectural state (branch predictor, prefetcher, etc.) is uninitialized.  Subsequent runs see a different performance profile.&lt;/p&gt;

&lt;p&gt;A single measurement cannot distinguish between a &lt;em&gt;warm&lt;/em&gt; run that reflects the steady‑state performance and a &lt;em&gt;cold&lt;/em&gt; run that is an outlier.&lt;/p&gt;

&lt;h2&gt;
  
  
  What problem csperf solves (this episode's slice)
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;csperf&lt;/code&gt; automates the &lt;em&gt;warm‑up / repeat&lt;/em&gt; pattern and stores the raw data in a JSON artifact.  The artifact contains:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the machine metadata (CPU, OS, compiler, etc.)&lt;/li&gt;
&lt;li&gt;the command line used for compilation and execution&lt;/li&gt;
&lt;li&gt;a list of execution times for each repeat&lt;/li&gt;
&lt;li&gt;summary statistics (min, max, mean, stdev)&lt;/li&gt;
&lt;li&gt;a full set of hardware counters collected by &lt;code&gt;perf&lt;/code&gt; or &lt;code&gt;papi&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With this data you can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;prove that your performance claim is reproducible&lt;/li&gt;
&lt;li&gt;compare different compiler flags or backends&lt;/li&gt;
&lt;li&gt;detect regressions in a CI pipeline&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Mental model
&lt;/h2&gt;

&lt;p&gt;Think of a performance run as a &lt;em&gt;warm‑up&lt;/em&gt; phase followed by a &lt;em&gt;steady‑state&lt;/em&gt; phase.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌───────────────────────┐
│  Warm‑up runs (2)      │
│  (discarded)           │
└─────────────┬─────────┘
              │
              ▼
┌───────────────────────┐
│  Repeat runs (5)       │
│  (kept for analysis)   │
└───────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The warm‑up runs are executed to bring the program into a stable state.  They are &lt;em&gt;not&lt;/em&gt; part of the final statistics.  The repeat runs are what you publish.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lab: install and first commands
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Install &lt;code&gt;csperf&lt;/code&gt;&lt;/strong&gt; (if you haven’t already):
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   pip &lt;span class="nb"&gt;install &lt;/span&gt;csperf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Compile and run a simple matrix traversal&lt;/strong&gt; (the same example used in Ep 1):
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   csperf run &lt;span class="se"&gt;\&lt;/span&gt;
     &lt;span class="nt"&gt;--input&lt;/span&gt; examples/cpp/matrix_traversal.cpp &lt;span class="se"&gt;\&lt;/span&gt;
     &lt;span class="nt"&gt;--backend&lt;/span&gt; cpu &lt;span class="se"&gt;\&lt;/span&gt;
     &lt;span class="nt"&gt;--warmup-runs&lt;/span&gt; 2 &lt;span class="se"&gt;\&lt;/span&gt;
     &lt;span class="nt"&gt;--repeat-runs&lt;/span&gt; 5 &lt;span class="se"&gt;\&lt;/span&gt;
     &lt;span class="nt"&gt;--output&lt;/span&gt; results/warm.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This command:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;compiles the C++ file with &lt;code&gt;clang++ -O3&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;runs the binary twice as warm‑up (discarded)&lt;/li&gt;
&lt;li&gt;runs it five times, recording each execution time&lt;/li&gt;
&lt;li&gt;writes a JSON artifact to &lt;code&gt;results/warm.json&lt;/code&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Profile the artifact&lt;/strong&gt; to see the raw numbers:
&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   csperf profile results/warm.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The profile command prints a concise table of the summary statistics and the raw per‑run times.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lab: what we ran on this machine
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hostname&lt;/td&gt;
&lt;td&gt;f4c59d864117&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Date&lt;/td&gt;
&lt;td&gt;2026‑09‑26T19:21:22+05:30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU&lt;/td&gt;
&lt;td&gt;AMD Ryzen 7 9700X (8‑core, 16‑thread)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OS&lt;/td&gt;
&lt;td&gt;Ubuntu 24.04 (Linux‑7.0.0‑31‑generic)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compiler&lt;/td&gt;
&lt;td&gt;clang++ (LLVM 18)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backend&lt;/td&gt;
&lt;td&gt;CPU (native)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Warm‑up runs&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repeat runs&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Artifact&lt;/td&gt;
&lt;td&gt;&lt;code&gt;results/warm.json&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The machine metadata is embedded in the artifact; you can inspect it with &lt;code&gt;csperf profile&lt;/code&gt; or by opening the JSON file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results (real numbers only)
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;csperf&lt;/code&gt; run produced the following execution times (in milliseconds) for the five repeat runs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Time (ms)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;5.151&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;5.165&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;5.163&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;5.182&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;5.165&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Summary statistics (from the artifact):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Count&lt;/strong&gt;: 5&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Min&lt;/strong&gt;: 5.151 ms&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Max&lt;/strong&gt;: 5.182 ms&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mean&lt;/strong&gt;: 5.1634 ms&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Median&lt;/strong&gt;: 5.165 ms&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Standard Deviation&lt;/strong&gt;: 0.0127 ms&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Hardware counters (subset):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Counter&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CPU cycles&lt;/td&gt;
&lt;td&gt;5 407 786&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reference cycles&lt;/td&gt;
&lt;td&gt;6 685 659&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction count&lt;/td&gt;
&lt;td&gt;20 043 266&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Branch instructions&lt;/td&gt;
&lt;td&gt;3 125 327&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Branch mispredictions&lt;/td&gt;
&lt;td&gt;35 076&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache references&lt;/td&gt;
&lt;td&gt;1 243 626&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache misses&lt;/td&gt;
&lt;td&gt;85 371&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IPC&lt;/td&gt;
&lt;td&gt;3.706&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These numbers come straight from &lt;code&gt;csperf/warm.json&lt;/code&gt; and are &lt;em&gt;verifiable&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to read the artifacts
&lt;/h2&gt;

&lt;p&gt;The JSON artifact is a self‑contained record of the experiment.  Key sections:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;pipeline&lt;/code&gt;&lt;/strong&gt; – shows the compiler steps and the profiler used.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;commands&lt;/code&gt;&lt;/strong&gt; – the exact shell commands that were executed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;metrics&lt;/code&gt;&lt;/strong&gt; – raw measurements and summary statistics.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;artifacts&lt;/code&gt;&lt;/strong&gt; – paths to the binary and any exported CSV/Excel files.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can use &lt;code&gt;csperf profile&lt;/code&gt; to pretty‑print the summary, or &lt;code&gt;csperf export --format csv results/warm.json&lt;/code&gt; to get a CSV for spreadsheet analysis.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes (teacher checklist)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Skipping warm‑up runs&lt;/strong&gt; – always set &lt;code&gt;--warmup-runs&lt;/code&gt; to at least 1.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Using a single repeat run&lt;/strong&gt; – set &lt;code&gt;--repeat-runs&lt;/code&gt; to ≥ 5 for a stable mean.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ignoring the artifact&lt;/strong&gt; – publish the JSON, not just the printed numbers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Running on a shared machine&lt;/strong&gt; – background processes can skew the results; use a dedicated test rig.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not pinning the binary to a CPU&lt;/strong&gt; – &lt;code&gt;csperf&lt;/code&gt; can set CPU affinity; otherwise the OS may migrate the process.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Try this next (homework)
&lt;/h2&gt;

&lt;p&gt;Run the same experiment on a different compiler backend (e.g., &lt;code&gt;gcc -O3&lt;/code&gt;) and compare the mean execution time and IPC.  Document the differences in a short markdown file.&lt;/p&gt;

&lt;p&gt;Do this tonight — Episode 3 starts by assuming you did.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;The warm‑up / repeat pattern is the foundation of any credible performance measurement.  &lt;code&gt;csperf&lt;/code&gt; removes the guesswork and gives you a machine‑aware, repeatable artifact that anyone can audit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The series so far
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Ep 1: Why a Single ./a.out Time Misleads Your Performance Claims &lt;em&gt;(this article)&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Ep 2: Warm vs Cold: Why a Single Trial Misleads Performance Claims&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Teaser for next episode
&lt;/h2&gt;

&lt;p&gt;This episode builds on Ep 1’s warning about single‑shot timings and sets the stage for next time when we show why screenshots aren’t evidence and how to keep your artifacts in a reproducible repository.&lt;/p&gt;

</description>
      <category>performance</category>
      <category>benchmarking</category>
      <category>compilers</category>
      <category>csperf</category>
    </item>
    <item>
      <title>Why a Single ./a.out Time Misleads Your Performance Claims</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Fri, 25 Sep 2026 16:40:48 +0000</pubDate>
      <link>https://dev.to/aabhinavg/why-a-single-aout-time-misleads-your-performance-claims-33l8</link>
      <guid>https://dev.to/aabhinavg/why-a-single-aout-time-misleads-your-performance-claims-33l8</guid>
      <description>&lt;h2&gt;
  
  
  Why this lesson exists
&lt;/h2&gt;

&lt;p&gt;Single‑shot timings from &lt;code&gt;time ./a.out&lt;/code&gt; or screenshots are seductive but fragile. They ignore machine state, compiler flags, and the stochastic nature of modern CPUs. A single number cannot be shared, diffed, or reproduced reliably.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we're building toward
&lt;/h2&gt;

&lt;p&gt;We want a reproducible, machine‑aware evidence trail that anyone can inspect, compare, and trust. &lt;code&gt;csperf&lt;/code&gt; is that observatory.&lt;/p&gt;

&lt;h2&gt;
  
  
  The misconception
&lt;/h2&gt;

&lt;p&gt;People think “my program runs in 5 ms” is enough proof. In reality, that 5 ms hides cache warm‑ups, branch predictor state, and even the current frequency governor.&lt;/p&gt;

&lt;h2&gt;
  
  
  What problem csperf solves (this episode's slice)
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;csperf&lt;/code&gt; turns a raw run into a structured artifact: a CSV, JSON, and optional Excel sheet that record the machine, compiler, and timing statistics. It also runs a sanity check (&lt;code&gt;csperf doctor&lt;/code&gt;) to make sure the environment is ready.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mental model
&lt;/h2&gt;

&lt;p&gt;Think of a single &lt;code&gt;time&lt;/code&gt; call as a photograph. &lt;code&gt;csperf&lt;/code&gt; is a full‑body scan that records every relevant metric and the context that produced it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lab: install and first commands
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;compilersutra-perf
csperf doctor
csperf quickstart &lt;span class="nt"&gt;--output-dir&lt;/span&gt; results/quickstart
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The quickstart builds &lt;code&gt;matrix_traversal.cpp&lt;/code&gt; with &lt;code&gt;clang++ -O3&lt;/code&gt;, runs it three times, and writes the artifacts under &lt;code&gt;results/quickstart&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lab: what we ran on this machine
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Host: &lt;code&gt;f4c59d864117&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;OS: Ubuntu 24.04 (Linux 7.0.0‑31‑generic)&lt;/li&gt;
&lt;li&gt;CPU: AMD Ryzen 7 9700X, 16 cores, 5582 MHz max&lt;/li&gt;
&lt;li&gt;Compiler: clang++ 18.1.3&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;csperf&lt;/code&gt; version: 0.2.0&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Results (real numbers only)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Trial&lt;/th&gt;
&lt;th&gt;Time (ms)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;5.162&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;5.144&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;5.148&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Mean&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;5.151333&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Std dev&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.009452&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All numbers come from &lt;code&gt;results/quickstart/quickstart.csv&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to read the artifacts
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CSV&lt;/strong&gt; – human‑readable table of trials and summary stats.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;JSON&lt;/strong&gt; – machine‑friendly representation, useful for CI pipelines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Excel&lt;/strong&gt; – for quick visual inspection.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each artifact includes a &lt;code&gt;machine.txt&lt;/code&gt; snapshot that records the exact CPU model, clock, and available counters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes (teacher checklist)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Relying on a single &lt;code&gt;time&lt;/code&gt; invocation.&lt;/li&gt;
&lt;li&gt;Forgetting to run &lt;code&gt;csperf doctor&lt;/code&gt; before measuring.&lt;/li&gt;
&lt;li&gt;Ignoring the &lt;code&gt;runs&lt;/code&gt; field; 1‑shot runs are not statistically sound.&lt;/li&gt;
&lt;li&gt;Not checking the &lt;code&gt;cpu_model&lt;/code&gt; in the artifact to ensure consistency.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Try this next (homework)
&lt;/h2&gt;

&lt;p&gt;Do this tonight — Episode 2 starts by assuming you did.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;A single &lt;code&gt;./a.out&lt;/code&gt; time is a myth. &lt;code&gt;csperf&lt;/code&gt; gives you a verifiable, repeatable evidence trail. Next time, we’ll explore why warm‑up and repeat runs matter.&lt;/p&gt;

</description>
      <category>compilers</category>
      <category>performance</category>
      <category>benchmarking</category>
      <category>cpp</category>
    </item>
    <item>
      <title>Create Your Own LLVM Pass as a Plugin (Step-by-Step Guide)</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Tue, 21 Apr 2026 16:48:32 +0000</pubDate>
      <link>https://dev.to/aabhinavg/create-your-own-llvm-pass-as-a-plugin-step-by-step-guide-4cnh</link>
      <guid>https://dev.to/aabhinavg/create-your-own-llvm-pass-as-a-plugin-step-by-step-guide-4cnh</guid>
      <description>&lt;p&gt;f you're getting into compilers or exploring LLVM internals, one of the most powerful things you can learn is how to build your own LLVM Pass.&lt;/p&gt;

&lt;p&gt;I’ve put together a hands-on guide that walks you through creating an LLVM pass as a plugin — something that’s actually used in real-world compiler workflows.&lt;/p&gt;

&lt;p&gt;💡 What You’ll Learn&lt;br&gt;
How LLVM passes work internally&lt;br&gt;
How to create a custom pass from scratch&lt;br&gt;
How to structure your plugin&lt;br&gt;
How to build it using CMake&lt;br&gt;
How to dynamically load and run it in LLVM&lt;br&gt;
🔧 Why This Matters&lt;/p&gt;

&lt;p&gt;Most developers use compilers as a black box. But if you're serious about:&lt;/p&gt;

&lt;p&gt;Compiler development&lt;br&gt;
Performance optimization&lt;br&gt;
Static analysis&lt;br&gt;
Systems programming&lt;/p&gt;

&lt;p&gt;…then understanding LLVM passes is a game-changer.&lt;/p&gt;

&lt;p&gt;📚 Full Guide&lt;/p&gt;

&lt;p&gt;👉 &lt;a href="https://www.compilersutra.com/docs/llvm/llvm_basic/pass/Create_LLVM_Pass_As_A_Plugin" rel="noopener noreferrer"&gt;https://www.compilersutra.com/docs/llvm/llvm_basic/pass/Create_LLVM_Pass_As_A_Plugin&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;🧠 Who Is This For?&lt;br&gt;
Students learning compilers&lt;br&gt;
Engineers exploring LLVM&lt;br&gt;
Anyone building tooling around code analysis or optimisation&lt;/p&gt;

&lt;p&gt;If you're building something cool with LLVM, I’d love to hear about it 👇&lt;br&gt;
Let’s push the boundaries of compilers together.&lt;/p&gt;

&lt;h1&gt;
  
  
  llvm #compilers #cpp #systems #opensource #programming
&lt;/h1&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>From 0 to 1M Impressions: Building a Niche Compiler Blog</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Tue, 14 Apr 2026 14:30:05 +0000</pubDate>
      <link>https://dev.to/aabhinavg/from-0-to-1m-impressions-building-a-niche-compiler-blog-2hk1</link>
      <guid>https://dev.to/aabhinavg/from-0-to-1m-impressions-building-a-niche-compiler-blog-2hk1</guid>
      <description>&lt;p&gt;I’ve been working on a niche site focused on compilers and systems:&lt;br&gt;
👉 &lt;a href="https://www.compilersutra.com" rel="noopener noreferrer"&gt;https://www.compilersutra.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Recently, I checked my performance over the last 12 months — and the results surprised me:&lt;/p&gt;

&lt;p&gt;1.09M impressions&lt;br&gt;
11K clicks&lt;br&gt;
Ranking for topics like LLVM, OpenCL, TVM&lt;/p&gt;

&lt;p&gt;All of this came purely from organic search.&lt;/p&gt;

&lt;p&gt;💡 What worked for me&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Picking a niche most people ignore&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Compilers, low-level systems, and ML compilers aren’t “mainstream” topics.&lt;/p&gt;

&lt;p&gt;But that’s exactly why they work.&lt;/p&gt;

&lt;p&gt;Less noise → more authority over time.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Writing for depth, not just keywords&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Instead of chasing trends, I focused on:&lt;/p&gt;

&lt;p&gt;Explaining concepts deeply&lt;br&gt;
Covering real-world use cases&lt;br&gt;
Connecting theory → practical systems&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Consistency over intensity&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No crazy posting schedule.&lt;/p&gt;

&lt;p&gt;Just consistent effort over time.&lt;/p&gt;

&lt;p&gt;That’s what compounds.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Structured content &amp;gt; random blogs&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I started thinking in terms of:&lt;/p&gt;

&lt;p&gt;Learning paths&lt;br&gt;
Roadmaps&lt;br&gt;
Connected topics&lt;/p&gt;

&lt;p&gt;Instead of isolated articles.&lt;/p&gt;

&lt;p&gt;📈 What I’m focusing on next&lt;br&gt;
Expanding compiler-related topics (LLVM, MLIR, TVM)&lt;br&gt;
Building structured learning tracks&lt;br&gt;
Improving user experience and engagement&lt;br&gt;
[🤝 Let’s connect]&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmddphelr8xmbhew9iye1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmddphelr8xmbhew9iye1.png" alt="compilersutra seo" width="800" height="299"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you’re interested in:&lt;/p&gt;

&lt;p&gt;Compilers&lt;br&gt;
Systems programming&lt;br&gt;
Low-level optimization&lt;/p&gt;

&lt;p&gt;Check it out 👉 &lt;a href="https://www.compilersutra.com" rel="noopener noreferrer"&gt;https://www.compilersutra.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Would love your feedback and suggestions!&lt;/p&gt;

</description>
      <category>buildinpublic</category>
      <category>computerscience</category>
      <category>marketing</category>
      <category>programming</category>
    </item>
    <item>
      <title>Memory Hierarch</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Mon, 13 Apr 2026 11:49:45 +0000</pubDate>
      <link>https://dev.to/aabhinavg/memory-hierarch-ej1</link>
      <guid>https://dev.to/aabhinavg/memory-hierarch-ej1</guid>
      <description>&lt;p&gt;If you're working on compilers, runtimes, or low-level systems…&lt;br&gt;
Stop asking “what is cache?”&lt;/p&gt;

&lt;p&gt;Start asking 👉 “what kind of miss did my code create?”&lt;/p&gt;

&lt;p&gt;💡 One bad memory access = hundreds of cycles lost&lt;br&gt;
💡 L1 → L3 → DRAM = massive slowdown&lt;br&gt;
💡 Performance = access pattern, not just instructions&lt;br&gt;
I broke it all down with real benchmarks (Ryzen 9700X) 👇&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.compilersutra.com/docs/coa/memory-hierarchy/" rel="noopener noreferrer"&gt;https://www.compilersutra.com/docs/coa/memory-hierarchy/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;⚡ Learn:&lt;br&gt;
• Cache misses &amp;amp; set conflicts&lt;br&gt;
• False sharing &amp;amp; multithreading pitfalls&lt;br&gt;
• TLB &amp;amp; page-walk cost&lt;br&gt;
• Why loop tiling gives 30x speedups&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>beginners</category>
    </item>
    <item>
      <title>AMD ML Complete Stack</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Sun, 12 Apr 2026 07:02:28 +0000</pubDate>
      <link>https://dev.to/aabhinavg/amd-ml-complete-stack-3hnm</link>
      <guid>https://dev.to/aabhinavg/amd-ml-complete-stack-3hnm</guid>
      <description>&lt;p&gt;I wrote 6 lines of Triton…&lt;/p&gt;

&lt;p&gt;and it turned into thousands of GPU instructions.&lt;/p&gt;

&lt;p&gt;Python → TTIR → TTGIR → LLVM → AMDGCN → HSACO&lt;/p&gt;

&lt;p&gt;👉 a + b → buffer_load_b128&lt;/p&gt;

&lt;p&gt;👉 mask → v_cmp + conditional execution&lt;/p&gt;

&lt;p&gt;Here’s the truth:&lt;/p&gt;

&lt;p&gt;Your code is NOT what runs on the GPU.&lt;/p&gt;

&lt;p&gt;The compiler builds an entire execution pipeline in between.&lt;/p&gt;

&lt;p&gt;I dumped every stage and traced one kernel end-to-end 👇&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.compilersutra.com/docs/ml-compilers/mlcompilerstack/" rel="noopener noreferrer"&gt;https://www.compilersutra.com/docs/ml-compilers/mlcompilerstack/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After this, ML compilers don’t feel like “magic” anymore.&lt;/p&gt;

</description>
      <category>gpu</category>
      <category>cpu</category>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>Introduction to ML Compilers + Roadmap (MLIR, TVM, GPU Kernels)</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Sat, 11 Apr 2026 11:33:31 +0000</pubDate>
      <link>https://dev.to/aabhinavg/introduction-to-ml-compilers-roadmap-mlir-tvm-gpu-kernels-24hb</link>
      <guid>https://dev.to/aabhinavg/introduction-to-ml-compilers-roadmap-mlir-tvm-gpu-kernels-24hb</guid>
      <description>&lt;p&gt;Most people think they are running Python when they train ML models.&lt;/p&gt;

&lt;p&gt;They are not.&lt;/p&gt;

&lt;p&gt;Python is only the interface.&lt;/p&gt;

&lt;p&gt;The real execution happens somewhere completely different — inside an ML compiler stack.&lt;/p&gt;




&lt;h2&gt;
  
  
  🧠 What actually happens?
&lt;/h2&gt;

&lt;p&gt;When you write something like:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;matmul → add → relu&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;It looks simple.&lt;/p&gt;

&lt;p&gt;But internally, the system transforms it into multiple layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Python (model definition)&lt;/li&gt;
&lt;li&gt;Graph (tensor operations)&lt;/li&gt;
&lt;li&gt;Execution plan (optimized structure)&lt;/li&gt;
&lt;li&gt;Kernels (GPU/CPU instructions)&lt;/li&gt;
&lt;li&gt;Hardware execution&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At no point does the GPU “run Python”.&lt;/p&gt;

&lt;p&gt;It runs &lt;strong&gt;compiled kernels&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚙️ Why ML Compilers exist
&lt;/h2&gt;

&lt;p&gt;Because raw model code is inefficient for hardware.&lt;/p&gt;

&lt;p&gt;Without a compiler:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Too many kernel launches&lt;/li&gt;
&lt;li&gt;Unnecessary memory transfers&lt;/li&gt;
&lt;li&gt;No operator fusion&lt;/li&gt;
&lt;li&gt;Poor GPU utilization&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With a compiler:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Operations are fused&lt;/li&gt;
&lt;li&gt;Memory movement is reduced&lt;/li&gt;
&lt;li&gt;Execution is optimized for hardware&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  🔥 Key concepts covered
&lt;/h2&gt;

&lt;p&gt;This article builds the foundation for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;MLIR (multi-level IR systems)&lt;/li&gt;
&lt;li&gt;TVM (end-to-end ML compiler stack)&lt;/li&gt;
&lt;li&gt;GPU kernel execution model&lt;/li&gt;
&lt;li&gt;Operator fusion &amp;amp; memory planning&lt;/li&gt;
&lt;li&gt;Compilation pipeline design&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  🧭 Roadmap (what you’ll learn)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Tensors, shapes, memory layout&lt;/li&gt;
&lt;li&gt;CPU vs GPU execution model&lt;/li&gt;
&lt;li&gt;Compiler basics (IR, lowering, passes)&lt;/li&gt;
&lt;li&gt;ML compiler optimizations&lt;/li&gt;
&lt;li&gt;Real systems (TVM, MLIR, XLA)&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  📘 Full Article
&lt;/h2&gt;

&lt;p&gt;👉 [&lt;a href="https://www.compilersutra.com/docs/ml-compile" rel="noopener noreferrer"&gt;https://www.compilersutra.com/docs/ml-compile&lt;/a&gt;&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>deeplearning</category>
      <category>machinelearning</category>
      <category>performance</category>
    </item>
    <item>
      <title>Building CompilerSutra</title>
      <dc:creator>compilersutra</dc:creator>
      <pubDate>Thu, 02 Apr 2026 04:09:59 +0000</pubDate>
      <link>https://dev.to/aabhinavg/building-compilersutra-26a9</link>
      <guid>https://dev.to/aabhinavg/building-compilersutra-26a9</guid>
      <description>&lt;p&gt;🚀 Building practical content on compilers, LLVM, MLIR, and performance.&lt;/p&gt;

&lt;p&gt;If this sounds interesting, you can join here:&lt;br&gt;
&lt;a href="https://docs.google.com/forms/d/e/1FAIpQLSebP1JfLFDp0ckTxOhODKPNVeI1e21rUqMJ0fbBwJoaa-i4Yw/viewform" rel="noopener noreferrer"&gt;link&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Would love to know what topics you’d like covered.&lt;/p&gt;

</description>
      <category>llvm</category>
      <category>computerscience</category>
      <category>college</category>
    </item>
  </channel>
</rss>
