<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: 4thwithme</title>
    <description>The latest articles on DEV Community by 4thwithme (@4thwithme).</description>
    <link>https://dev.to/4thwithme</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F132255%2F41d4b73d-4ede-4b46-b5ff-4b048afbc3d8.webp</url>
      <title>DEV Community: 4thwithme</title>
      <link>https://dev.to/4thwithme</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/4thwithme"/>
    <language>en</language>
    <item>
      <title>Cron jobs: the tiny line that runs half your backend</title>
      <dc:creator>4thwithme</dc:creator>
      <pubDate>Sun, 27 Sep 2026 19:36:51 +0000</pubDate>
      <link>https://dev.to/4thwithme/cron-jobs-the-tiny-line-that-runs-half-your-backend-3o8a</link>
      <guid>https://dev.to/4thwithme/cron-jobs-the-tiny-line-that-runs-half-your-backend-3o8a</guid>
      <description>&lt;p&gt;Every backend has a few jobs nobody looks at. The nightly backup. The 9 AM report. The script that deletes old files. Each one is a single line of cron. They work fine for months. Then one day you find out the backup hasn't run since March. Nobody noticed. This post is about that one line: what it means, where to run it, and how to hear about it when it stops.&lt;/p&gt;

&lt;h2&gt;
  
  
  What cron actually is
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;cron&lt;/strong&gt; - a program that runs in the background on Unix machines. It starts commands on a schedule. Each scheduled command is a &lt;strong&gt;cron job&lt;/strong&gt;. All the jobs live in one file, the &lt;strong&gt;crontab&lt;/strong&gt; (cron table). You edit it with &lt;code&gt;crontab -e&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;the name comes from Chronos, the Greek word for time&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Think of it as an alarm clock. But instead of ringing, it runs a program. Cron is old. It first shipped in 1979. Most Linux machines today still run a version based on Paul Vixie's rewrite from 1987. In 2025 the syntax finally got a real written standard, OCPS 1.0. GitHub Actions, Kubernetes, and AWS all use the same kind of line, with small changes.&lt;/p&gt;

&lt;p&gt;One thing to know first: cron has no memory. Every minute it wakes up. It checks each line: "should this run right now?" It runs the ones that match. Then it goes back to sleep. It doesn't know what it ran yesterday. If the machine was off at 02:00, the 02:00 job just doesn't happen. Nothing runs it later. (For laptops that sleep a lot, use &lt;code&gt;anacron&lt;/code&gt; instead.) Remember this. Most of the tips below exist because of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to read a cron line
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmjbkm8xcsv9wcxg5b8n3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmjbkm8xcsv9wcxg5b8n3.png" alt="A cron line split into five boxes: minute every-15, hour 9-17, day of month any, month any, day of week 1-5, followed by the command ./sync.sh, each box labeled with its allowed range" width="800" height="391"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;read left to right: minute, hour, day of month, month, day of week, then the command&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Each field is a filter. Cron checks all five every minute. If all five say yes, the job runs. You only need four symbols for almost everything:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symbol&lt;/th&gt;
&lt;th&gt;Means&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;*&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;every value&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;*&lt;/code&gt; in hour = every hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;,&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;a list&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;1,15&lt;/code&gt; = the 1st and the 15th&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;-&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;a range&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;1-5&lt;/code&gt; = Monday to Friday&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;a step&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;*/15&lt;/code&gt; = 0, 15, 30, 45&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;There are also shortcuts, like &lt;code&gt;@daily&lt;/code&gt; (same as &lt;code&gt;0 0 * * *&lt;/code&gt;) and &lt;code&gt;@hourly&lt;/code&gt;. But not every tool supports them. GitHub Actions doesn't. Write all five fields and it works everywhere.&lt;/p&gt;
&lt;h2&gt;
  
  
  Two examples, read slowly
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;*/15 9-17 * * 1-5&lt;/code&gt; - every 15 minutes, starting at 9 AM, Monday to Friday.&lt;/p&gt;

&lt;p&gt;The trap: &lt;code&gt;9-17&lt;/code&gt; means "any minute where the hour is 9 to 17". So the whole 17th hour counts. The last run is at &lt;strong&gt;17:45&lt;/strong&gt;, not 17:00. That's &lt;strong&gt;36 runs a day&lt;/strong&gt;. Want to stop at exactly 17:00? Use &lt;code&gt;9-16&lt;/code&gt;, then add one more line: &lt;code&gt;0 17 * * 1-5&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;next runs from Mon 2026-10-05: 09:00, 09:15, 09:30 ... 17:45 (checked with croniter)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;0 3 1,15 * 5&lt;/code&gt; - you'd read it as "03:00 on the 1st and 15th, but only if it's Friday". Wrong.&lt;/p&gt;

&lt;p&gt;If &lt;strong&gt;both&lt;/strong&gt; day fields are set (neither one is &lt;code&gt;*&lt;/code&gt;), cron uses &lt;strong&gt;OR&lt;/strong&gt;, not AND. So it runs on the 1st, on the 15th, &lt;strong&gt;and every Friday&lt;/strong&gt;. In October 2026 that's the 1st, 2nd, 9th, 15th, 16th, 23rd, and 30th. Seven runs in one month. You probably expected zero.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;the oldest cron surprise, and it's in the rules on purpose&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Not sure what a line does? Paste it into &lt;a href="https://crontab.guru" rel="noopener noreferrer"&gt;crontab.guru&lt;/a&gt;. Or print the next few run times with a library: &lt;code&gt;croniter&lt;/code&gt; in Python, &lt;code&gt;cron-parser&lt;/code&gt; in Node. Do this before you deploy.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where to run it: GitHub Actions
&lt;/h2&gt;

&lt;p&gt;You don't need a server for this. Add a workflow file to &lt;code&gt;.github/workflows/&lt;/code&gt; with a &lt;code&gt;schedule&lt;/code&gt; trigger:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;schedule&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;cron&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;30&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;5&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;1-5'&lt;/span&gt;        &lt;span class="c1"&gt;# always quote: a value starting with * breaks YAML&lt;/span&gt;
      &lt;span class="na"&gt;timezone&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;America/New_York"&lt;/span&gt; &lt;span class="c1"&gt;# optional, default is UTC&lt;/span&gt;
  &lt;span class="na"&gt;workflow_dispatch&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;               &lt;span class="c1"&gt;# lets you run it by hand too&lt;/span&gt;
&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;report&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;./scripts/daily-report.sh&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What people learn the hard way:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It runs in UTC by default.&lt;/strong&gt; The &lt;code&gt;timezone&lt;/code&gt; key is optional and fairly new.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every 5 minutes&lt;/strong&gt; is the fastest you can go.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runs can be late, or dropped.&lt;/strong&gt; This happens when GitHub is busy. The top of every hour is the busiest time. So use minute 17, not minute 0.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Only the default branch counts.&lt;/strong&gt; A schedule on your feature branch never runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Public repos with no activity for 60 days&lt;/strong&gt; get their schedules turned off.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Good for anything that's fine a few minutes late. Bad for anything that must run on time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to run it: Docker
&lt;/h2&gt;

&lt;p&gt;You can put normal &lt;code&gt;cron&lt;/code&gt; in a container. It will start. But it fails in three ways, and none of them show an error:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your job can't see env vars.&lt;/strong&gt; Cron starts each job with an almost empty env. You set &lt;code&gt;DATABASE_URL&lt;/code&gt; on the container, but the job never gets it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your logs are empty.&lt;/strong&gt; Cron sends job output to email or syslog, not to stdout. So &lt;code&gt;docker logs&lt;/code&gt; shows nothing, even when the job fails.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stopping the container can break data.&lt;/strong&gt; In a container, cron is often the main process (PID 1). When Docker stops the container, cron doesn't wait for jobs to finish. A job can get killed halfway through writing a file.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most people fix this with &lt;a href="https://github.com/aptible/supercronic" rel="noopener noreferrer"&gt;supercronic&lt;/a&gt;. It's a cron made for containers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; debian:bookworm-slim&lt;/span&gt;
&lt;span class="c"&gt;# install supercronic binary (see its README for the pinned URL + checksum)&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; crontab /app/crontab&lt;/span&gt;
&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["supercronic", "/app/crontab"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It gives every job the container's env vars. It writes job output to the container log. It shuts down cleanly on &lt;code&gt;SIGTERM&lt;/code&gt;. And it won't start a job again while the last run is still going. Add &lt;code&gt;supercronic -test crontab&lt;/code&gt; to your CI. It catches a broken line before you deploy.&lt;/p&gt;

&lt;p&gt;Don't want to change your images? Try &lt;a href="https://github.com/mcuadros/ofelia" rel="noopener noreferrer"&gt;ofelia&lt;/a&gt;. It reads schedules from Docker labels. It can run a command inside a container that's already running (&lt;code&gt;job-exec&lt;/code&gt;). Or it can start a new container for each run (&lt;code&gt;job-run&lt;/code&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to run it: Kubernetes
&lt;/h2&gt;

&lt;p&gt;Kubernetes has cron built in. It's called a &lt;code&gt;CronJob&lt;/code&gt;. It doesn't run your code by itself. It's a template. Each time the schedule fires, it creates a new &lt;code&gt;Job&lt;/code&gt;. Each Job starts one or more Pods.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvnny9vgmh9c6u8u8g96r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvnny9vgmh9c6u8u8g96r.png" alt="Diagram: one CronJob creates a Job for each scheduled time; each Job runs a Pod, and a crashed Pod is retried with a new Pod" width="800" height="320"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;CronJob → one Job per run → Pods, with retries inside each Job&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;batch/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;CronJob&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nightly-backup&lt;/span&gt;            &lt;span class="c1"&gt;# max 52 characters&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;schedule&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;17&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;2&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;
  &lt;span class="na"&gt;timeZone&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Europe/Kyiv"&lt;/span&gt;         &lt;span class="c1"&gt;# GA since v1.27&lt;/span&gt;
  &lt;span class="na"&gt;concurrencyPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Forbid&lt;/span&gt;       &lt;span class="c1"&gt;# never two at once&lt;/span&gt;
  &lt;span class="na"&gt;startingDeadlineSeconds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;600&lt;/span&gt;    &lt;span class="c1"&gt;# too late by 10 min? skip it&lt;/span&gt;
  &lt;span class="na"&gt;successfulJobsHistoryLimit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3&lt;/span&gt;
  &lt;span class="na"&gt;failedJobsHistoryLimit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3&lt;/span&gt;
  &lt;span class="na"&gt;jobTemplate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;backoffLimit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;             &lt;span class="c1"&gt;# default is 6 retries&lt;/span&gt;
      &lt;span class="na"&gt;activeDeadlineSeconds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3600&lt;/span&gt; &lt;span class="c1"&gt;# kill it after an hour&lt;/span&gt;
      &lt;span class="na"&gt;ttlSecondsAfterFinished&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;86400&lt;/span&gt;
      &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;restartPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Never&lt;/span&gt;
          &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;backup&lt;/span&gt;
              &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;myorg/backup:1.4.2&lt;/span&gt;
              &lt;span class="na"&gt;args&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--target"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s3://backups"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The most important field is &lt;code&gt;concurrencyPolicy&lt;/code&gt;. It answers one question: the next run is due, but the last one isn't done yet. What now?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fizg9qe8ecvi2s6drlqce.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fizg9qe8ecvi2s6drlqce.png" alt="Timeline of a job scheduled every 10 minutes that takes 15 minutes: Allow stacks overlapping runs, Forbid skips the 10-minute tick, Replace kills each run when the next starts" width="800" height="338"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Allow piles up, Forbid skips, Replace kills. Choose on purpose. The default is Allow.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two more things from the docs. First, without &lt;code&gt;timeZone&lt;/code&gt;, the schedule uses the time zone of the cluster's controller. Not yours. Second, Kubernetes says a CronJob can create two Jobs for one run. Or none at all. So, in their words, your jobs "should be idempotent."&lt;/p&gt;

&lt;p&gt;To test it now, start one run by hand: &lt;code&gt;kubectl create job --from=cronjob/nightly-backup test-run&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;On AWS, use EventBridge Scheduler. Its cron has six fields (a year at the end), and one day field must be &lt;code&gt;?&lt;/code&gt;. A normal crontab line gets rejected.&lt;/p&gt;
&lt;h2&gt;
  
  
  Best practices
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Idempotent&lt;/strong&gt; - running it twice gives the same result as running it once. "Set today's total to X" is idempotent. "Add today's sales to the total" is not. Run that twice and your revenue doubles.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;the one rule that makes every other failure harmless&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Make every job idempotent.&lt;/strong&gt; Then retries, double runs, and manual reruns can't hurt you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stop overlaps.&lt;/strong&gt; On a normal server, use &lt;code&gt;flock&lt;/code&gt;: &lt;code&gt;*/5 * * * * flock -n /run/lock/sync.lock /usr/local/bin/sync.sh&lt;/code&gt;. If the lock is taken, the new run just exits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set a time limit.&lt;/strong&gt; Use &lt;code&gt;timeout 2h ./job.sh&lt;/code&gt;, or &lt;code&gt;activeDeadlineSeconds&lt;/code&gt; in Kubernetes. Without it, one stuck job blocks every run after it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use UTC.&lt;/strong&gt; Local time has daylight saving. In spring, a 02:30 job can get skipped. In autumn, it can run twice. It depends on your cron version. If you must use local time, don't schedule between 01:00 and 03:00.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't use minute 0.&lt;/strong&gt; Everyone does. Spread your jobs out. Or add a small random delay, so a hundred servers don't hit one database at the same second.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use full paths and set your env.&lt;/strong&gt; Cron's &lt;code&gt;PATH&lt;/code&gt; is tiny (&lt;code&gt;/usr/bin:/bin&lt;/code&gt;). Its shell is &lt;code&gt;/bin/sh&lt;/code&gt;, not bash. "It works in my terminal" means nothing here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escape &lt;code&gt;%&lt;/code&gt;.&lt;/strong&gt; In a crontab, &lt;code&gt;%&lt;/code&gt; means a new line. So &lt;code&gt;date +%F&lt;/code&gt; breaks without an error. Write &lt;code&gt;date +\%F&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log to somewhere you'll read.&lt;/strong&gt; Use &lt;code&gt;&amp;gt;&amp;gt; /var/log/job.log 2&amp;gt;&amp;amp;1&lt;/code&gt;, or stdout in a container. But a log is not an alert. Nobody reads the log of a job that never started.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep crontabs in git.&lt;/strong&gt; Someone typed a line into &lt;code&gt;crontab -e&lt;/code&gt; on a server in 2021? Nobody will ever find it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat jobs like prod services.&lt;/strong&gt; Wire up the same APM and error tracking as your app: Sentry, New Relic, Rollbar. Ship logs, traces, and metrics like duration and memory. Then a slow or failing job shows up on the same dashboards as your API.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;
  
  
  Know when it didn't run
&lt;/h2&gt;

&lt;p&gt;That last problem is the big one. When a job crashes, you see an error. When it never runs, you see nothing. Maybe the server was down. Maybe cron was off. Maybe someone deleted the line. There's no error to catch. So flip it around: alert on silence.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fax7qr7kt8qdpdz4f7qh3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fax7qr7kt8qdpdz4f7qh3.png" alt="Flow: the cron job sends start, success, or fail pings to a monitor; if no ping arrives by the expected time plus a grace period, the monitor alerts you" width="800" height="320"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;a dead man's switch: silence is the alarm&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Here's how it works. The job sends a ping to a monitor when it's done. The monitor knows the schedule. If no ping comes by the expected time (plus a little extra time), you get an alert. It's one &lt;code&gt;curl&lt;/code&gt; at the end of the job:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;0 2 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; /opt/backup.sh &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; curl &lt;span class="nt"&gt;-fsS&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; 10 &lt;span class="nt"&gt;--retry&lt;/span&gt; 5 &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null https://hc-ping.com/&amp;lt;your-uuid&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  3 services that do this for you
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Service&lt;/th&gt;
&lt;th&gt;Free plan&lt;/th&gt;
&lt;th&gt;Why pick it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://healthchecks.io" rel="noopener noreferrer"&gt;Healthchecks.io&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;20 checks&lt;/td&gt;
&lt;td&gt;Open source (BSD). You can self-host it. Knows cron syntax and time zones. Has &lt;code&gt;/start&lt;/code&gt; and &lt;code&gt;/fail&lt;/code&gt; pings. Can send the exit code with &lt;code&gt;/$?&lt;/code&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://cronitor.io" rel="noopener noreferrer"&gt;Cronitor&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;5 monitors&lt;/td&gt;
&lt;td&gt;Its CLI reads your crontab and wraps each line for you (&lt;code&gt;cronitor exec&lt;/code&gt;). Also tracks how long each run takes.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://docs.sentry.io/product/crons/" rel="noopener noreferrer"&gt;Sentry Crons&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;1 monitor&lt;/td&gt;
&lt;td&gt;Already on Sentry? Missed or slow runs show up as issues, right next to that job's errors.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Prices change, so check first. For a side project, free Healthchecks.io is enough. Team already on Sentry? Keep alerts there.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to actually do
&lt;/h2&gt;

&lt;p&gt;Read each line out loud before you deploy it. Print the next five runs. Use the platform you already have. GitHub Actions if late is okay. supercronic in Docker. A CronJob with &lt;code&gt;Forbid&lt;/code&gt; and a deadline in Kubernetes. Make each job safe to run twice. Then add that one &lt;code&gt;curl&lt;/code&gt;. It tells you when a job didn't run at all, and that's the failure you'd never see on your own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Useful links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://man7.org/linux/man-pages/man5/crontab.5.html" rel="noopener noreferrer"&gt;crontab(5) man page&lt;/a&gt; - the official syntax reference&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://crontab.guru" rel="noopener noreferrer"&gt;crontab.guru&lt;/a&gt; - paste a line, read it in plain English&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.github.com/en/actions/writing-workflows/choosing-when-your-workflow-runs/events-that-trigger-workflows#schedule" rel="noopener noreferrer"&gt;GitHub Docs: schedule event&lt;/a&gt; - limits and delays&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://kubernetes.io/docs/concepts/workloads/controllers/cron-jobs/" rel="noopener noreferrer"&gt;Kubernetes: CronJob&lt;/a&gt; - every field explained&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/aptible/supercronic" rel="noopener noreferrer"&gt;supercronic&lt;/a&gt; - cron for containers&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://blog.healthchecks.io/2021/10/how-debian-cron-handles-dst-transitions/" rel="noopener noreferrer"&gt;Healthchecks.io: how Debian cron handles DST&lt;/a&gt; - the full daylight saving story&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Elsewhere
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/4thwithme" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; · &lt;a href="https://www.linkedin.com/in/andrii-popenko-3331b3193/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://dev.to/4thwithme"&gt;Dev.to&lt;/a&gt; · &lt;a href="https://substack.com/@andriipopenko124891" rel="noopener noreferrer"&gt;Substack&lt;/a&gt; · &lt;a href="https://4thwithme.dev/blog" rel="noopener noreferrer"&gt;4thwithme.dev/blog&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;May the &lt;code&gt;--force&lt;/code&gt; be with you. See you next week.&lt;/p&gt;

</description>
      <category>cron</category>
      <category>devops</category>
      <category>kubernetes</category>
      <category>docker</category>
    </item>
    <item>
      <title>How to measure color the way your eyes actually see it</title>
      <dc:creator>4thwithme</dc:creator>
      <pubDate>Mon, 21 Sep 2026 08:37:25 +0000</pubDate>
      <link>https://dev.to/4thwithme/how-to-measure-color-the-way-your-eyes-actually-see-it-2dcn</link>
      <guid>https://dev.to/4thwithme/how-to-measure-color-the-way-your-eyes-actually-see-it-2dcn</guid>
      <description>&lt;p&gt;Source post: &lt;a href="https://4thwithme.dev/blog/color-distance-why-rgb-math-lies/" rel="noopener noreferrer"&gt;https://4thwithme.dev/blog/color-distance-why-rgb-math-lies/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Take two colors, (0, 64, 0) and (255, 64, 0). The plain math says they're 255 apart. Now compare (255, 64, 0) to (255, 64, 128): only 128 apart, half as much. Trust just the numbers and the second pair looks closer together. Look at them with your own eyes and that can flip completely, because plain math has no idea what your eyes actually see. Good news: people already fixed this. It just took fifty years and four tries. Two real examples with real numbers, later in this post, show exactly how big the gap can be.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ WHY YOU CAN'T JUST SUBTRACT THE NUMBERS ]
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;distance&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sqrt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;R1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;R2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;^&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;G1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;G2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;^&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;B1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;B2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;^&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the first thing everyone tries, and it's wrong. Your eyes aren't equally good at spotting changes in red, green, and blue. A change of 10 in green is much easier to see than the same change of 10 in blue. Plain math doesn't know that, it treats every kind of "different" the same. It isn't the same.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;ΔE (Delta E)&lt;/strong&gt; - the standard way to measure how far apart two colors are. A score of 1.0 is supposed to be the smallest gap a human eye can notice, the Just Noticeable Difference. Every formula below is an attempt to make that number mean the same thing everywhere in the color space.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  [ WHERE THESE NUMBERS ACTUALLY COME FROM ]
&lt;/h2&gt;

&lt;p&gt;Every color space in this post, RGB, CIELAB, Oklab, comes from the same place: tests on real people in the 1920s. Someone looked at a light and mixed red, green, and blue until the mix matched it, across every color the eye can see.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0nwk4x4bjlxdjjy3vja8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0nwk4x4bjlxdjjy3vja8.png" alt="CIE 1931 xy chromaticity diagram, the horseshoe-shaped map of all colors visible to the human eye" width="800" height="800"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;the 1931 chromaticity diagram: every color the human eye can see, mapped from those matching experiments (Sakurambo, Wikimedia Commons, CC BY-SA)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Some matches needed a &lt;em&gt;negative&lt;/em&gt; amount of a light, which isn't possible in real life. So the results got turned into a new system, XYZ, where every visible color gets a normal, positive number. Every color space since is built on top of XYZ. But none of it matches what "similar" looks like to a person: XYZ is complete, not built to match perception. That gap is why the next fifty years of formulas exist.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foyuyzph275qrqy1ty8gd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foyuyzph275qrqy1ty8gd.png" alt="MacAdam ellipses plotted on the CIE 1931 chromaticity diagram, showing regions of colors indistinguishable to a human observer, magnified 10x" width="800" height="884"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;MacAdam ellipses (10x scale): equal steps in this raw space aren't equal steps in perception, some directions need way more distance before you'd notice&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  [ THE RACE TO FIX IT ]
&lt;/h2&gt;

&lt;p&gt;CIELAB (1976) was the first real attempt: a space where an equal number should mean an equal amount of visible change. It still gets strong blues wrong, so every formula after it patches the last one's mistake:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Formula&lt;/th&gt;
&lt;th&gt;Year&lt;/th&gt;
&lt;th&gt;What changed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ΔE76&lt;/td&gt;
&lt;td&gt;1976&lt;/td&gt;
&lt;td&gt;Plain distance in CIELAB. The smallest visible gap lands near 2.3, not the intended 1.0.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CMC l:c&lt;/td&gt;
&lt;td&gt;1984&lt;/td&gt;
&lt;td&gt;Weighted specifically for the textile industry.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ΔE94&lt;/td&gt;
&lt;td&gt;1995&lt;/td&gt;
&lt;td&gt;Weights that change by region, from car-paint testing. Still weak on blue-violet.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CIEDE2000&lt;/td&gt;
&lt;td&gt;2001&lt;/td&gt;
&lt;td&gt;Weights that adapt, plus a fix aimed straight at blue.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Some industries take this dead seriously: car paint has to stay under 0.5 ΔE, printing under 2.0 ΔE. A narrower formula, ΔE_NS (2019), exists just for that near-zero range: the pass/fail check for "are these basically the same color."&lt;/p&gt;
&lt;h2&gt;
  
  
  [ WHAT'S ACTUALLY INSIDE CIEDE2000 ]
&lt;/h2&gt;

&lt;p&gt;Here's the kid version. Say you have two crayons and want to know how different they look. You could measure three things: which one is lighter or darker, which one is more vivid or more washed-out, and which one leans more red vs. more blue. Add those three differences together (with a little extra twisting for blue, because blue kept fooling the older formulas) and that's basically the whole idea.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;ΔE00&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sqrt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ΔL&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; / (kL·SL))^2 +
  (ΔC&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;kC&lt;/span&gt;&lt;span class="err"&gt;·&lt;/span&gt;&lt;span class="n"&gt;SC&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="o"&gt;^&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ΔH&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; / (kH·SH))^2 +
  RT · (ΔC&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;kC&lt;/span&gt;&lt;span class="err"&gt;·&lt;/span&gt;&lt;span class="n"&gt;SC&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="err"&gt;·&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ΔH&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; / (kH·SH))
)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Scary-looking, but every piece is simple once you name it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ΔL'&lt;/strong&gt; - the lightness difference. How much lighter or darker is one color than the other.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ΔC'&lt;/strong&gt; - the chroma difference. How much more vivid or washed-out one is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ΔH'&lt;/strong&gt; - the hue difference. How much the actual color type shifted, more red vs. more blue.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;S_L, S_C, S_H&lt;/strong&gt; - "stretchy rulers." The same real gap counts as bigger or smaller depending on where you are in the color space, because our eyes aren't equally picky everywhere. Near black, a tiny lightness change stands out, so that ruler stretches. In a very vivid color, the same numeric change is harder to spot, so it shrinks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;R_T&lt;/strong&gt; - the twist. A small correction added just because blues specifically kept breaking the older formulas, it nudges the chroma and hue pieces to work together in that one problem region.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;k_L, k_C, k_H&lt;/strong&gt; - dials you can turn for your specific use case (lighting, industry). Left at 1 almost always.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's the whole trick: three simple differences, three rulers that stretch and shrink depending on where you're measuring, and one small twist for blue. What makes it hard isn't the idea, it's getting all those moving pieces to line up in code without a sign error. A 2005 paper by Sharma, Wu, and Dalal exists because early implementations, including ones people pointed to as "the right way," had bugs in exactly that plumbing.&lt;/p&gt;

&lt;p&gt;Two real pairs, both exactly 30 apart in plain RGB math. Run them through CIEDE2000 and the story flips completely:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🟦 &lt;code&gt;rgb(0,0,255)&lt;/code&gt; vs &lt;code&gt;rgb(30,0,255)&lt;/code&gt; → &lt;strong&gt;ΔE2000 = 0.62&lt;/strong&gt;. Below the notice-it threshold. Nobody sees a difference.&lt;/li&gt;
&lt;li&gt;⬛ &lt;code&gt;rgb(10,10,10)&lt;/code&gt; vs &lt;code&gt;rgb(10,10,40)&lt;/code&gt; → &lt;strong&gt;ΔE2000 = 14.99&lt;/strong&gt;. Same RGB gap, an obviously different color.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same 30-point RGB step, roughly 24x difference in how much it actually matters (computed with the &lt;code&gt;colour-science&lt;/code&gt; Python library).&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where it breaks&lt;/strong&gt; - the hue-angle math is easy to flip a sign on, the wraparound at 0°/360° needs its own special handling, not a plain remainder, and the weights genuinely change by region. Copy a CIE94-shaped formula (one formula, no branches) and you get a wrong number back, not an error. The good news: use a library people already tested, and check it against the published test numbers.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  [ THE CHEAP TRICK: STAY IN RGB ]
&lt;/h2&gt;

&lt;p&gt;All that trig and region-hopping is worth the trouble when accuracy is the whole point. It stops being worth it the moment you're running the comparison millions of times a second, palette matching, image quantization, because converting every pixel to Lab first is exactly the kind of work that piles up faster at scale. Redmean skips the trip entirely. It never leaves RGB, it just leans harder on red or blue depending on how much red is already in the pair:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;r_avg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;R1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;R2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;

&lt;span class="n"&gt;distance&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sqrt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;r_avg&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;         &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;R1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;R2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;^&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="mi"&gt;4&lt;/span&gt;                       &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;G1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;G2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;^&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;r_avg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;B1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;B2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;^&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It's not the most accurate choice, plain CIELUV usually beats it on average. Its real strength is that it never has a bad day: Luv can go badly wrong on some colors, skin tones especially, and redmean doesn't, for a lot less computer work.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ OKLAB: A FRESH START, NOT ANOTHER PATCH ]
&lt;/h2&gt;

&lt;p&gt;Redmean and CIEDE2000 are still playing the same game: patch RGB, patch CIELAB, patch the patch. Oklab walks away from the table instead. Björn Ottosson published it in 2020 after running a simple test on CIELAB: hold lightness and saturation fixed, slide through every hue, and watch what happens to the color wheel. Yellow, magenta, and cyan come out visibly brighter than red and blue, a strip that flickers as it spins, not because the colors changed but because the space measuring them is uneven. Oklab was built from real perception data specifically to make that flicker disappear.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuwubomai9ies4kw4m2jb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuwubomai9ies4kw4m2jb.png" alt="HSV hue gradient at constant value and saturation, showing yellow and cyan appearing noticeably lighter than red and blue" width="800" height="120"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsqv2wjlwt5dbd0je4dp9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsqv2wjlwt5dbd0je4dp9.png" alt="Oklab hue gradient at constant lightness and chroma, showing even perceived brightness across all hues" width="800" height="120"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;same hue sweep, top (HSV/sRGB) visibly flickers in brightness, bottom (Oklab) stays flat (Björn Ottosson)&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Oklab / Oklch&lt;/strong&gt; - XYZ turns into Oklab through two simple matrix multiplications with a cube root in between: cheap, stable, no special cases by region. Make a color brighter or darker and the numbers just scale, nothing else needs fixing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's been part of the CSS color spec since 2021, works in every major browser, and is Photoshop's default way to blend colors: the better starting point for gradients and color pickers on the web today, not CIELAB. (Two colors can share the exact same XYZ numbers and still look different under different lighting, that's what a color appearance model like CIECAM02 is built to handle.)&lt;/p&gt;

&lt;h2&gt;
  
  
  [ WHERE THIS SHOWS UP IN A REAL FEATURE ]
&lt;/h2&gt;

&lt;p&gt;A tool that searches photos by color first has to decide what "the main color" even means. Just counting pixels doesn't work: point it at a beach photo and "blue" always finds the sky, never someone's blue shirt. One real version of this uses plain Lab distance, but turns the lightness weight down to 0.7, so a light and dark version of the same hue don't get scored as too far apart. The color math here was the easy part. The hard part was speeding up the database and showing results as they arrive, instead of making people wait. Swapping in CIEDE2000 wouldn't have made this feature faster. It usually doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ WHAT TO ACTUALLY USE ]
&lt;/h2&gt;

&lt;p&gt;Use plain Lab distance by default, it's good enough for almost everything. Reach for CIEDE2000 only when checking against a limit already written in ΔE2000: printing, paint, fabric. Reach for redmean when converting to Lab is genuinely too slow. And for anything visual on the web today, start with Oklab/Oklch: already built into your browser, designed to skip the problems everything above it has been patching since 1976.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ LIBRARIES THAT DO THIS FOR YOU ]
&lt;/h2&gt;

&lt;p&gt;Don't write CIEDE2000 by hand. Here's what already does it, checked against current docs, not memory:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Python&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Library&lt;/th&gt;
&lt;th&gt;What it covers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;colour-science&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Broadest: ΔE76/94/2000, CMC, DIN99, HyAB, plus Oklab and CIECAM02/CIECAM16.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;scikit-image&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;skimage.color.deltaE_cie76&lt;/code&gt;, &lt;code&gt;deltaE_ciede94&lt;/code&gt;, &lt;code&gt;deltaE_ciede2000&lt;/code&gt;, &lt;code&gt;deltaE_cmc&lt;/code&gt;. No Oklab.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;coloraide&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Native Oklab/Oklch, &lt;code&gt;delta_e()&lt;/code&gt; with CIE76, DE2000, HyAB, and an Oklab-based "OK" method.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;colormath&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The old classic, archived since Dec 2023. Use the &lt;code&gt;colormath2&lt;/code&gt; fork instead.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Node.js&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Library&lt;/th&gt;
&lt;th&gt;What it covers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;culori&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;CIE76/94/2000 difference functions, native Oklab/Oklch parsing and conversion.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;colorjs.io&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;By the CSS Color 4 spec editors. &lt;code&gt;deltaE()&lt;/code&gt; with &lt;code&gt;76&lt;/code&gt;, &lt;code&gt;CMC&lt;/code&gt;, &lt;code&gt;2000&lt;/code&gt;, &lt;code&gt;Jz&lt;/code&gt;, &lt;code&gt;deltaEOK&lt;/code&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;color-diff&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;CIEDE2000 specifically, plus palette-mapping (&lt;code&gt;closest&lt;/code&gt;, &lt;code&gt;mapPalette&lt;/code&gt;).&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;nearest-color&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Plain Euclidean RGB nearest-neighbor, no perceptual conversion, the cheap option from earlier.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One thing to watch for: &lt;code&gt;chroma-js&lt;/code&gt; ships a &lt;code&gt;deltaE()&lt;/code&gt; too, but it's CMC l:c under the hood, not CIEDE2000, despite the name. Don't use it for perceptual ΔE2000 matching.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ USEFUL LINKS ]
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://bottosson.github.io/posts/oklab/" rel="noopener noreferrer"&gt;Björn Ottosson: "A perceptual color space for image processing"&lt;/a&gt; - the original Oklab post&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://hajim.rochester.edu/ece/sites/gsharma/papers/CIEDE2000CRNAFeb05.pdf" rel="noopener noreferrer"&gt;Sharma, Wu, Dalal: CIEDE2000 implementation notes&lt;/a&gt; - the fix-your-bugs paper&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.compuphase.com/cmetric.htm" rel="noopener noreferrer"&gt;CompuPhase: "Colour metric"&lt;/a&gt; - the redmean derivation&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://en.wikipedia.org/wiki/CIE_1931_color_space" rel="noopener noreferrer"&gt;Wikipedia: CIE 1931 color space&lt;/a&gt; - where this all starts&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Originally posted on &lt;a href="https://4thwithme.dev/blog/color-distance-why-rgb-math-lies/" rel="noopener noreferrer"&gt;4thwithme.dev/blog&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/4thwithme" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; · &lt;a href="https://www.linkedin.com/in/andrii-popenko-3331b3193/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://substack.com/@andriipopenko124891" rel="noopener noreferrer"&gt;Substack&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;May the &lt;code&gt;--force&lt;/code&gt; be with you. See you next week.&lt;/p&gt;

</description>
      <category>color</category>
      <category>colormatching</category>
      <category>colordistance</category>
      <category>webdev</category>
    </item>
    <item>
      <title>A/B testing: how to test your test</title>
      <dc:creator>4thwithme</dc:creator>
      <pubDate>Mon, 14 Sep 2026 08:47:13 +0000</pubDate>
      <link>https://dev.to/4thwithme/ab-testing-how-to-test-your-test-enn</link>
      <guid>https://dev.to/4thwithme/ab-testing-how-to-test-your-test-enn</guid>
      <description>&lt;p&gt;&lt;em&gt;Source post: &lt;a href="https://4thwithme.dev/blog/ab-testing-how-to-test-your-test/" rel="noopener noreferrer"&gt;https://4thwithme.dev/blog/ab-testing-how-to-test-your-test/&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Posts 3 and 4 were about the math: what "significant" actually means, and which test design fits which situation. This one is about a question that comes before either: how do you know the test itself is working? A p-value can be perfectly correct and still be answering a question corrupted by a tracking bug, a broken randomizer, or a confound nobody looked for. Here's how to check.&lt;/p&gt;

&lt;h2&gt;
  
  
  A/A tests
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A/A test&lt;/strong&gt; — run two completely identical experiences against each other and check whether your pipeline reports a difference. There shouldn't be one. If it finds one anyway, the problem isn't your product, it's your testing setup.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's the cheapest validity check available and it's underused precisely because it feels like it can't fail - you're testing nothing against nothing, what's there to break? Plenty. Randomization bugs, tracking that double-fires for one arm, IDs that leak between groups: none of that shows up until you look. Microsoft's experimentation team &lt;a href="https://medium.com/data-science-at-microsoft/how-to-better-control-for-false-positives-while-monitoring-your-experiment-62df02e8a451" rel="noopener noreferrer"&gt;documented a platform where repeated A/A tests&lt;/a&gt; turned up a false-positive rate as high as 30%, against a nominal 5% - the stats engine itself was miscalibrated, and no amount of correct math downstream would have caught that. The fix isn't a one-time check before launch, it's continuous: route any traffic that isn't allocated to a real test into a standing A/A test, so a regression in the pipeline shows up before it corrupts something you actually care about.&lt;/p&gt;

&lt;p&gt;In plain words: two vending machines, stocked with the exact same snacks at the exact same prices. If your sales dashboard tells you machine B is suddenly selling 20% more candy than machine A, the candy isn't the story - one of the machines is broken, or someone's miscounting the coins.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sample ratio mismatch
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Sample ratio mismatch (SRM)&lt;/strong&gt; — you configured a 50/50 split, but the traffic that actually landed in each group isn't 50/50. A chi-squared goodness-of-fit test on the observed counts tells you whether that gap is noise or a real problem.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The industry convention here is a much stricter bar than the usual 5%: teams flag SRM at &lt;strong&gt;p &amp;lt; 0.001&lt;/strong&gt;, deliberately tighter, because at real traffic volumes even a trivial imbalance clears 0.05 easily, and you'd rather tolerate a few false alarms than miss a broken splitter. Here's what that looks like with actual numbers: 100,000 visitors, intended 50/50, and you actually get 52,000 in control against 48,000 in test. Run the chi-squared test on that and you get χ² = 160, which lands around p ≈ 10⁻³⁶ - not a borderline call, a five-sigma-and-then-some alarm bell. That kind of gap is almost never real randomness; it's bot traffic skewed into one arm, an ID collision, or a redirect that leaks users out of their assigned bucket before they're logged.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fri1pmh1qtngghz3ke305.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fri1pmh1qtngghz3ke305.png" alt="Bar chart showing 50,000 expected visitors per arm against an actual observed split of 52,000 control vs. 48,000 test, flagged by a chi-squared value of 160" width="800" height="462"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;a 4% imbalance looks small until the chi-squared test does the math on it&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Run this check before you read any other result from the test. A significant SRM doesn't tell you your product change failed, it tells you the comparison itself isn't trustworthy, full stop - fix the split, don't interpret the metric.&lt;/p&gt;

&lt;p&gt;In plain words: a teacher splits 100 kids into two identical classrooms by coin flip, expecting roughly 50 and 50. If 62 end up in room A and 38 in room B, you don't go looking for what's different about the kids in room A - you go check the coin.&lt;/p&gt;

&lt;h2&gt;
  
  
  Simpson's Paradox
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simpson's Paradox&lt;/strong&gt; — the aggregate result shows B beating A, but every individual segment shows A beating B. The overall number and the segment-level numbers are both correct; they just tell opposite stories, because something is unevenly distributed between the two.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The textbook real-world case is a 1986 study of 700 kidney stone patients comparing two treatments. Overall, Treatment B (a less invasive procedure) had a better success rate than Treatment A, 83% against 78%. But split the same patients by stone size, and Treatment A wins in both the small-stone group and the large-stone group. What happened: doctors preferentially gave the gentler Treatment B to patients with smaller stones, who were easier cases to begin with regardless of treatment. Size was doing the real work; treatment assignment was riding along on top of it.&lt;/p&gt;

&lt;p&gt;That case is a selection-bias story, not a randomization bug - nobody split anything wrong, doctors were just choosing based on severity. In an actual A/B test, where assignment is supposed to be random, the same pattern showing up is a stronger signal: something correlated with the outcome ended up unevenly split between your two groups. Check segment-level performance after any test that looks clean in aggregate - by platform, region, new vs. returning, whatever you track - and if every segment disagrees with the topline number, suspect the split before you trust the topline.&lt;/p&gt;

&lt;p&gt;In plain words: two delivery drivers. Driver A is faster than Driver B on short routes, and faster than Driver B on long routes too. But overall, across the month, Driver B has the better average time - because Driver B mostly got assigned short routes, and Driver A got stuck with the long ones. Add it all up and the assignment is doing the work, not the driving.&lt;/p&gt;

&lt;h2&gt;
  
  
  Other checks worth running
&lt;/h2&gt;

&lt;p&gt;A few more from the same family, lower ceremony than the three above but worth having in the rotation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Guardrail metrics&lt;/strong&gt; - track a small set of core health metrics (page load time, crash rate, unsubscribe rate) on every test regardless of what it's meant to measure, so a change that "wins" on its target metric while quietly breaking something else gets caught immediately. Example: a checkout redesign lifts conversion 3%, but the page-load guardrail shows +400ms - that's not a false alarm, that's the guardrail doing exactly its job.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multiple-comparison correction&lt;/strong&gt; - the more metrics or segments you check on a single test, the more of those checks will clear a 5% bar by chance alone. If you're tracking a dashboard of twenty metrics per test, expect roughly one of them to look "significant" for no real reason - correct your threshold (Bonferroni or similar) or treat any single metric flagged out of a large batch skeptically until it replicates. Example: scroll depth alone spikes to p=0.03 out of twenty tracked metrics where nothing else moved - that's the base rate you'd expect from chance, not a discovery.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pre-launch bias check&lt;/strong&gt; - before a test goes live, run the intended split for a short window with no actual treatment applied yet (an A/A period baked into the same pipeline) and confirm it comes back clean. Catches the same class of bug as a standing A/A test, just gated in front of the specific experiment instead of running generally. Example: a 48-hour blank dry run on a pricing test's exact split logic, confirming both arms convert identically before the real price difference goes live.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The checklist
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check&lt;/th&gt;
&lt;th&gt;Catches&lt;/th&gt;
&lt;th&gt;Run it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A/A test&lt;/td&gt;
&lt;td&gt;Tracking, randomization, pipeline bugs&lt;/td&gt;
&lt;td&gt;Continuously, on idle traffic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SRM (chi-squared)&lt;/td&gt;
&lt;td&gt;Broken splitter, bot traffic, ID leaks&lt;/td&gt;
&lt;td&gt;Before reading any other result&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Segment breakdown&lt;/td&gt;
&lt;td&gt;Simpson's Paradox, hidden confounds&lt;/td&gt;
&lt;td&gt;After any "clean" result&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Guardrail metrics&lt;/td&gt;
&lt;td&gt;A win that breaks something else&lt;/td&gt;
&lt;td&gt;On every test, by default&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multiple-comparison correction&lt;/td&gt;
&lt;td&gt;False positives from dashboard-scanning&lt;/td&gt;
&lt;td&gt;Whenever you track many metrics per test&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A statistically significant result off a broken pipeline is just a precise measurement of the wrong thing. These checks are cheap relative to what they protect against - run them before you trust the number, not after someone's already acted on it.&lt;/p&gt;




&lt;p&gt;&lt;a href="https://github.com/4thwithme" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; · &lt;a href="https://www.linkedin.com/in/andrii-popenko-3331b3193/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://substack.com/@andriipopenko124891" rel="noopener noreferrer"&gt;Substack&lt;/a&gt; · &lt;a href="https://4thwithme.dev/blog" rel="noopener noreferrer"&gt;4thwithme.dev/blog&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;May the &lt;code&gt;--force&lt;/code&gt; be with you. See you next week.&lt;/p&gt;

</description>
      <category>abtesting</category>
      <category>statistics</category>
      <category>datascience</category>
      <category>management</category>
    </item>
    <item>
      <title>A/B testing: when the standard test is the wrong tool</title>
      <dc:creator>4thwithme</dc:creator>
      <pubDate>Mon, 07 Sep 2026 07:44:07 +0000</pubDate>
      <link>https://dev.to/4thwithme/ab-testing-when-the-standard-test-is-the-wrong-tool-40an</link>
      <guid>https://dev.to/4thwithme/ab-testing-when-the-standard-test-is-the-wrong-tool-40an</guid>
      <description>&lt;p&gt;Last week was the standard fixed-horizon test: split traffic 50/50, wait for your pre-calculated sample size, look once, done. That's still the right default most of the time. It's not the only tool, though, and treating it as the only one gets expensive once its assumptions stop matching your situation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the default isn't always right
&lt;/h2&gt;

&lt;p&gt;A fixed-horizon test answers one question: is there a detectable difference between two fixed groups, at a preset error tolerance, once enough data has come in. That's the whole job. It doesn't answer "how do I make the most money while this test runs," and it assumes your two groups don't affect each other. Both assumptions can quietly fail.&lt;/p&gt;

&lt;p&gt;Here's a real version of why that matters. A subscription business earning roughly $1M a month runs a pricing test: 10% baseline conversion, a 5-week horizon at the standard 5%/80% thresholds. One week in, the test group converts at 15% against the control's 10% - a trend that, if it held, would add roughly $500K to that month's revenue. The textbook says don't stop, the sample size isn't there yet.&lt;/p&gt;

&lt;p&gt;That's not a math problem, it's the wrong question. The test was built to answer "is there a detectable difference at a fixed error rate," not "how do I make the most money while this runs." Point it at the second question and it keeps giving you the wrong kind of answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi-armed bandits
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Multi-armed bandit&lt;/strong&gt; - instead of holding a fixed split for a fixed duration, you continuously shift more traffic toward whichever variant is currently winning, while still sending a smaller slice to the others so you don't lock in a lucky early read. The name comes from a gambler in front of a row of slot machines ("one-armed bandits"), trying to find the one that pays out best while losing as little as possible to the rest.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That tension has a name: the &lt;a href="https://en.wikipedia.org/wiki/Exploration%E2%80%93exploitation_dilemma" rel="noopener noreferrer"&gt;exploration-exploitation tradeoff&lt;/a&gt;. Exploit too early and you lock in the wrong winner off a noisy first week; explore forever and you never cash in on knowing which one wins. Studied by Allied scientists in World War II, and found so hard that mathematician Peter Whittle joked it should be dropped over Germany so enemy scientists could waste their time on it too, before Herbert Robbins gave it a serious mathematical treatment in 1952. Three practical strategies still run in production today: epsilon-greedy (mostly exploit, explore at random sometimes), UCB (favor whatever you're least certain about), and Thompson sampling (a Bayesian approach that updates its belief as data comes in).&lt;/p&gt;

&lt;p&gt;Run as a bandit, the pricing test above looks different: most traffic goes to whichever price is winning, a smaller slice keeps testing the alternative, and the split keeps moving as the picture updates, instead of freezing at 50/50 for five weeks.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ey030x943t0surrggaz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ey030x943t0surrggaz.png" alt="Line chart comparing a flat 50/50 traffic split against a bandit's allocation climbing from 50% toward 90% over 20 days" width="800" height="462"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;the losing variant never gets shut off entirely, it just gets starved of traffic&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The trade-off: a bandit chases money made during the test, not a clean significance result, and usually needs longer to reach the same confidence one calculated sample size gets you. Use it when running a worse variant a little longer costs more than a tidy p-value is worth.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;In simple words&lt;/strong&gt; - Three ice cream stands, and you want to know which one people like best. A regular test sends exactly one-third of people to each stand for a whole month, even once it's obvious stand #2 is winning. A bandit is smarter: it quietly sends more people to stand #2, while still sending a few to the others, just in case they get better later.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Switchback tests
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Switchback test&lt;/strong&gt; - instead of splitting your population into two groups, split time into blocks and alternate the whole population between A and B as blocks pass: this hour everyone gets A, next hour everyone gets B, on a randomized schedule.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The case for this shows up once your users stop being independent of each other. A standard split assumes what happens to user A doesn't leak into user B's experience. That breaks in a marketplace: if half of riders see one price and half see another, drivers respond to whichever price is in front of them, and both groups end up competing for the same pool of drivers. The "control" group is now polluted by the treatment group's effect on shared supply. Alternating the whole market by time block avoids that: "this city under A" versus "the same city under B," not two artificially split groups sharing the same supply.&lt;/p&gt;

&lt;p&gt;There's no single textbook citation for this the way "multi-armed bandit" has one; the closest formal relative is what statisticians call crossover design, adapted for markets instead of individual subjects. Uber and Lyft have both published on using it for dynamic-pricing and dispatch experiments, exactly because their two-sided marketplaces make a standard user-level split unreliable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjq5haej8y4j5dm49gt7w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjq5haej8y4j5dm49gt7w.png" alt="Diagram contrasting a standard user-level split against a switchback design that alternates the whole market between A and B by time block" width="799" height="373"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;every time block is its own control for the block right next to it&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;In simple words&lt;/strong&gt; - One playground, one set of kids. You can't fairly split them "half get the red slide, half get the blue slide," because they all play together and one group's fun affects the other's. So instead: everyone gets red on Monday, blue on Tuesday, red again on Wednesday - same kids, same playground, just switching what they try on different days.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Hold-out tests
&lt;/h2&gt;

&lt;p&gt;A hold-out test is the simplest design here: keep a slice of your population permanently untouched by the change, and compare how their behavior evolves against everyone else's over time. No alternating, no traffic-shifting, just a group frozen in the "before" state as a long-running reference point. It earns its place for effects a short test can't see - trust, habit, long-term retention - anything that shows up as a slow drift over months, not a clean signal within weeks. The cost: you're deliberately holding back a possibly-better experience from real people, for a long time, just to keep a clean comparison available later.&lt;/p&gt;

&lt;p&gt;A real one: Uber's rider-acquisition team saw wild week-to-week swings in Meta ad cost-per-acquisition that looked more like noise than signal, so they ran a 3-month hold-out - just turned Meta ads off for a slice of the market and watched. Nothing happened. Signups held steady, meaning the ads were just taking credit for people who'd have signed up anyway. That freed up &lt;a href="https://experimental.beehiiv.com/p/uber-saved-35m-ads" rel="noopener noreferrer"&gt;roughly $35M a year&lt;/a&gt; for channels that actually moved the number - a finding a two-week test would never have had the runway to catch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pre/post &amp;amp; synthetic control
&lt;/h2&gt;

&lt;p&gt;Sometimes there's no way to randomize at all, because there's only one of you: one product launching a redesign, one market entering a new pricing policy. You can't hold half a country in a control group. The blunt option is a pre/post comparison - metric before, metric after, call the difference the effect. It's also the weakest design here: everything else moving in that window (seasonality, a competitor's move, the economy) gets mixed into your "effect" with no way to pull it back out.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Synthetic control&lt;/strong&gt; - when you have exactly one treated unit and no real control group, build a fake one: a weighted blend of other similar untreated units, chosen so the blend tracks your unit's trend closely &lt;em&gt;before&lt;/em&gt; the change. After the change, the gap between what actually happened and what the blend predicts is your estimated effect.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This isn't an internet-testing invention, it comes from economics. Alberto Abadie and Javier Gardeazabal introduced it in 2003, building a "synthetic Basque Country" out of a blend of other Spanish regions to estimate what its economy would have done without a decades-long conflict. The &lt;a href="https://en.wikipedia.org/wiki/Synthetic_control_method" rel="noopener noreferrer"&gt;method's Wikipedia page&lt;/a&gt; is a good primer; the idea carries over cleanly to product work, where "other regions" become other markets, stores, or cohorts that never saw your change.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvb83zw9p5uwgyz4ucx4z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvb83zw9p5uwgyz4ucx4z.png" alt="Line chart showing an actual trajectory and a synthetic control trajectory tracking closely before a change, then diverging after it" width="800" height="462"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;before the change, the lines sit on top of each other - that's what makes the later gap meaningful&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;In simple words&lt;/strong&gt; - One plant, and you want to know if a new fertilizer helped it grow. No second identical plant to compare against. So you blend ten other, similar-but-not-identical plants into one imaginary "average plant" that grew just like yours did before the fertilizer. If your real plant suddenly grows taller than that imaginary one, the gap is probably the fertilizer.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Don't reach for A/B/C/D early
&lt;/h2&gt;

&lt;p&gt;Testing four or five variants at once is tempting when you have a backlog and want an answer fast. Resist it. A clean multi-variant test needs roughly the same sample size as a sequence of pairwise tests, but it's far harder to tell which comparison drove the result, and far easier to p-hack yourself by checking every pair until one clears the bar. A sequence costs calendar time, not data, and buys back your ability to actually interpret what happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to actually choose
&lt;/h2&gt;

&lt;p&gt;No universal method fits every situation - last week's flow solves one problem, nothing beyond that. You'll hit other limits too: overlapping tests, insufficient traffic, a goal about maximizing an outcome instead of detecting a difference. A rough guide:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Your situation&lt;/th&gt;
&lt;th&gt;Reach for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Can randomize independently, want a calibrated error rate&lt;/td&gt;
&lt;td&gt;Standard fixed-horizon A/B test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Users interact with each other (marketplace, shared supply)&lt;/td&gt;
&lt;td&gt;Switchback test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Goal is maximizing outcome during the test, not a calibrated verdict&lt;/td&gt;
&lt;td&gt;Multi-armed bandit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Only one treated unit exists, no randomization possible&lt;/td&gt;
&lt;td&gt;Pre/post, ideally with synthetic control&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Effect is slow, long-term, or compounding&lt;/td&gt;
&lt;td&gt;Hold-out test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;More than two variants on the table&lt;/td&gt;
&lt;td&gt;Decompose into a sequence of A/B tests&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;None of these replace the fixed-horizon test as a default. Reach for them once you can name the specific way your situation breaks its assumptions - not before.&lt;/p&gt;

&lt;h2&gt;
  
  
  Elsewhere
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/4thwithme" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; · &lt;a href="https://www.linkedin.com/in/andrii-popenko-3331b3193/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://dev.to/4thwithme"&gt;Dev.to&lt;/a&gt; · &lt;a href="https://substack.com/@andriipopenko124891" rel="noopener noreferrer"&gt;Substack&lt;/a&gt; · &lt;a href="https://4thwithme.dev/blog" rel="noopener noreferrer"&gt;4thwithme.dev/blog&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;May the &lt;code&gt;--force&lt;/code&gt; be with you. See you next week.&lt;/p&gt;

</description>
      <category>abtesting</category>
      <category>statistics</category>
      <category>datascience</category>
      <category>experimentation</category>
    </item>
    <item>
      <title>A/B testing: the stat-sig problem</title>
      <dc:creator>4thwithme</dc:creator>
      <pubDate>Sun, 30 Aug 2026 22:54:41 +0000</pubDate>
      <link>https://dev.to/4thwithme/ab-testing-the-stat-sig-problem-2mg1</link>
      <guid>https://dev.to/4thwithme/ab-testing-the-stat-sig-problem-2mg1</guid>
      <description>&lt;p&gt;Post source: &lt;a href="https://4thwithme.dev/blog/ab-testing-stat-sig-problem/" rel="noopener noreferrer"&gt;https://4thwithme.dev/blog/ab-testing-stat-sig-problem/&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  [ THE SETUP ]
&lt;/h2&gt;

&lt;p&gt;Generic scenario, no company specifics: you run an ecommerce storefront. Someone ships a change, a re-ranked product grid, a new checkout button color, doesn't matter. You split traffic 50/50, wait for data, and want to know: did this actually help, or did it just look like it helped?&lt;/p&gt;

&lt;p&gt;That's the whole job of an A/B test. It answers exactly one question: is there a difference between the two groups, yes or no. It does not tell you &lt;em&gt;why&lt;/em&gt;. It does not reliably tell you &lt;em&gt;how much&lt;/em&gt; the metric will move once you roll out to everyone. It rejects a null hypothesis or it fails to reject it. That's the entire output. Everything else people read into a test result, they're reading in themselves.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;[ DEF.1 ]&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Null hypothesis (H₀)&lt;/strong&gt; — the default, skeptical assumption that nothing changed: there is no real difference between A and B. An A/B test is built entirely to argue against this one claim. It either rejects it (evidence a difference exists) or fails to reject it (not enough evidence either way). It never proves the opposite outright, it just runs out of reasons to doubt it.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;the assumption every A/B test starts by trying to disprove&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  [ WHAT "STAT SIG" ACTUALLY MEANS ]
&lt;/h2&gt;

&lt;p&gt;Here's the misunderstanding almost everyone ships with: a p-value is not "the probability that B is better than A." It's the probability of seeing data this extreme (or more extreme) &lt;em&gt;if there were actually no difference at all&lt;/em&gt;. It's a statement about the data given no effect, not a statement about the effect given the data. Those sound similar. They are not the same claim, and the difference matters every time you're deciding whether to trust a result.&lt;/p&gt;

&lt;p&gt;Put another way: a p-value doesn't tell you "B is probably better." It tells you "if B and A were actually identical, how surprising would this result be?" A low p-value means the result would be surprising under the assumption of no real difference — that assumption starts to look shaky, so it's worth trusting. It does not tell you the odds that B is actually better; that's a different question the test never answers. It's the same logic as saying "if this coin were fair, getting 9 heads in 10 flips would be surprising" — that doesn't prove the coin is rigged, it just means the fair-coin explanation is hard to believe.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;[ DEF.2 ]&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;P-value&lt;/strong&gt; — the probability of seeing a result at least this extreme, &lt;em&gt;assuming the null hypothesis is true&lt;/em&gt;, assuming there's actually no difference between A and B. It is not the probability that B beats A, and it is not the probability that the null hypothesis itself is true. A small p-value just means "this would be a strange coincidence if nothing had actually changed," nothing more.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;a statement about the data given no effect, not about the effect given the data&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To get to a p-value you first have to set up two competing claims:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Null hypothesis (H₀)&lt;/strong&gt;: there is no difference between A and B. This is always the same claim, every single time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alternative hypothesis (H₁)&lt;/strong&gt;: there is a difference.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An A/B test never proves H₁. It only ever rejects H₀ or fails to reject it. Failing to reject H₀ doesn't mean "there's no effect," it means "we didn't detect one at the sensitivity we set up for." That distinction alone would save a lot of bad blog posts about "our test failed, this feature doesn't work."&lt;/p&gt;

&lt;p&gt;Then there are two ways to be wrong, and they have genuinely useful mnemonics:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;[ DEF.3 ]&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Type I error (false positive)&lt;/strong&gt; — you see a difference that isn't really there. Telling grandpa he's pregnant.&lt;br&gt;
&lt;strong&gt;Type II error (false negative)&lt;/strong&gt; — you miss a difference that is really there. Telling a pregnant woman she isn't.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;the two ways an A/B test can lie to you&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The industry default is 5% tolerance for Type I (α = 0.05) and 20% tolerance for Type II (β = 0.20, i.e. 80% power). Neither number is handed down by nature. They're both choices you make based on the cost of being wrong. A drug trial might demand α with eight zeroes before the first significant digit. A button-color test where you genuinely don't care much either way could reasonably run at α = 0.20. The 5% convention is a convention, not a law.&lt;/p&gt;

&lt;p&gt;Here's how the p-value and alpha actually connect: alpha is the surprise threshold you commit to &lt;em&gt;before&lt;/em&gt; you look at any data — your pre-agreed tolerance for crying wolf. The p-value is how surprising your actual result turned out to be. The decision rule is just a comparison: p-value &amp;lt; alpha means "significant." That's it. And alpha isn't just a cutoff, it IS your Type I error rate — a 5% alpha means you've pre-accepted a 5% chance of telling grandpa he's pregnant, across all the times you'd run this test. Type II (missing a real difference) isn't controlled by alpha at all — that's a separate knob, beta, driven mostly by how much data you collect.&lt;/p&gt;

&lt;p&gt;Beta works the opposite way from alpha: it's the rate at which you fail to catch a difference that's actually there — telling the pregnant woman she isn't. Beta isn't a tolerance you set directly like alpha; it falls out of how much data you collect, given the effect size you're trying to detect. Small sample, small true effect: beta is high, meaning you'll miss it most of the time and walk away concluding "no difference" when there really was one. More data pulls beta down. Power is just the flip side of the same number: power = 1 − β, so the industry-default β = 0.20 is the same fact stated as "80% power" — an 80% chance of actually seeing the effect if it's real, and a 20% chance of missing it outright.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ THE FOUR NUMBERS THAT SET THE BAR ]
&lt;/h2&gt;

&lt;p&gt;Before you run anything, four inputs determine how much data you actually need: your alpha threshold, your power, your baseline conversion rate, and your MDE, minimum detectable effect, the smallest lift you actually care about being able to see.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;[ DEF.4 ]&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MDE (minimum detectable effect)&lt;/strong&gt; — the smallest lift you actually care about being able to see. It's not a property of the data, it's a choice you make going in: below this size, a real effect might exist and your test still won't reliably catch it. Set it too small and your required sample size balloons past what your traffic can deliver in a reasonable timeframe; set it too large and you'll miss smaller wins that were real.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;the smallest lift you've decided is worth being able to see&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Change any one of these and your required sample size moves, usually a lot more than intuition suggests. Sample size explodes as your baseline moves away from 50%, or as your MDE shrinks. Concretely, at a 10% baseline conversion rate with the standard 5%/80% thresholds:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Detecting this relative lift&lt;/th&gt;
&lt;th&gt;Sample size needed, per arm&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;10% (10% → 11%)&lt;/td&gt;
&lt;td&gt;~14,313&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5% (10% → 10.5%)&lt;/td&gt;
&lt;td&gt;~56,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1% (10% → 10.01%)&lt;/td&gt;
&lt;td&gt;~1,414,681&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fma6xphewwit9d7onzpvi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fma6xphewwit9d7onzpvi.png" alt="Bar chart showing required sample size per arm exploding from ~14,313 at 10% relative lift to ~1,414,681 at 1% relative lift" width="800" height="462"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;FIG.1 — going from "detect a 10% lift" to "detect a 1% lift" costs you 100x the traffic&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Going from detecting a 10% lift to detecting a 1% lift costs you roughly one hundred times the traffic. This is why small-traffic products shouldn't be running color-tweak A/B tests: they simply don't have the volume to detect small effects reliably, and running one anyway doesn't make the math work, it just produces a noisy coin flip dressed up as a decision. Where does MDE actually come from in practice? There's no clean formula. In order of how people actually do it: gut feel on plausibility, historical results from similar past tests, or, the more honest method, work backwards from your constraints. How much traffic do you actually have, how many hypotheses are sitting in the backlog, how fast do you need to decide. Set your MDE from what's actually testable in your time and traffic budget, not from an abstract target someone wrote on a roadmap slide.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ USING THE CALCULATOR: TWO EXAMPLES ]
&lt;/h2&gt;

&lt;p&gt;You don't need to memorize the sample size formula to use any of this. &lt;a href="https://www.evanmiller.org/ab-testing/sample-size.html" rel="noopener noreferrer"&gt;Evan Miller's sample size calculator&lt;/a&gt; does the arithmetic for you: plug in your baseline conversion rate and your minimum detectable effect, and it hands you a number per variation. The one setting that trips people up is the toggle between &lt;strong&gt;Absolute&lt;/strong&gt; and &lt;strong&gt;Relative&lt;/strong&gt;. Absolute means "detect a change of this many percentage points" (a 20% baseline with a 5-point absolute MDE means detecting 20% → 25%). Relative means "detect a change of this many percent of the baseline itself" (a 20% baseline with a 10% relative MDE means detecting 20% → 22%, since 10% of 20 is 2). Relative is almost always the more honest way to think about it, because the same absolute point-move means something very different at a 2% baseline than at a 50% one. Here are two examples run through the same tool, same 5%/80% thresholds, to show how differently the same math treats two real spots on a storefront.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example 1: top of the funnel.&lt;/strong&gt; Say you're changing the color of the main call-to-action button on the homepage, high-traffic real estate that basically every visitor sees. Baseline click-through is 20%, and you want to detect at least a 10% relative lift, meaning you'd notice a move to roughly 22% or higher.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5bzv66w94fwnqznlg76q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5bzv66w94fwnqznlg76q.png" alt="Evan Miller's sample size calculator showing a 20% baseline conversion rate and 10% relative minimum detectable effect, resulting in a required sample size of 6,347 per variation" width="800" height="473"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;FIG.2 — 20% baseline, 10% relative MDE → 6,347 per variation&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That's 6,347 per variation, 12,694 total. If the homepage does something like 3,000 visits a day, split 50/50, that's 1,500 per arm per day, so you'd clear the bar in well under a week. The strategy here is straightforward because the traffic supports it: pick a real, meaningful MDE, run a standard fixed-horizon test, wait for the number, done.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example 2: deep in the site, low traffic.&lt;/strong&gt; Now say the change is on an advanced export feature buried three clicks into account settings, the kind of thing only a small slice of engaged users ever reaches. Baseline conversion on the action you're watching (say, clicking "export") is 3%, and because the effect would need to be large to matter at this baseline, you set a generous 20% relative MDE, detecting a move to roughly 3.6% or higher.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr92iyna04ha79nmb6866.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr92iyna04ha79nmb6866.png" alt="Evan Miller's sample size calculator showing a 3% baseline conversion rate and 20% relative minimum detectable effect, resulting in a required sample size of 13,050 per variation" width="800" height="473"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;FIG.3 — 3% baseline, 20% relative MDE (already a generous ask) → 13,050 per variation&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Even with an MDE more than double the first example's, you need 13,050 per variation, more than double the homepage number, because the low baseline is working against you the whole time. If that settings page gets 300 visits a week, that's 150 per arm per week. Reaching 13,050 per arm at that rate takes roughly 87 weeks, about a year and a half. That's not a test, that's a career milestone.&lt;/p&gt;

&lt;p&gt;The two examples share a calculator and a formula. They don't share a strategy. For the homepage button, run the standard test, it's cheap and fast. For the deep feature, a fixed-horizon 50/50 split is very likely the wrong tool entirely: either loosen your MDE further until the sample size matches your actual traffic (accepting you can only detect huge swings), or stop trying to force a classic A/B test onto traffic that can't support one, and reach for one of the alternatives from the next post in this series instead, a long-running hold-out, a qualitative read, or a Bayesian approach that doesn't demand a pre-fixed sample size to say something useful early.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ THE PROBLEM: PEEKING ]
&lt;/h2&gt;

&lt;p&gt;Here's where the dashboard-refreshing habit turns into an actual statistical problem, not just an impatience problem. Refreshing the page doesn't corrupt the underlying data, the numbers are whatever they are regardless of who's watching, but it does change your &lt;em&gt;stopping rule&lt;/em&gt;. "Stop and declare a winner the moment p &amp;lt; 0.05" isn't one test, it's a test you get to retry every time you check.&lt;/p&gt;

&lt;p&gt;And "check" has a specific meaning here: computing the p-value and using it to decide whether to stop, not just glancing at a chart. Opening the dashboard to see where the numbers stand is harmless on its own. It's treating that p-value as a decision point, "is this under 0.05, should I call it," that counts as a look.&lt;/p&gt;

&lt;p&gt;Your 5% alpha threshold is a guarantee about exactly one such decision point, made after your pre-calculated sample size is reached. Every time you check the result before that point and treat "p &amp;lt; 0.05 right now" as a stopping signal, you're giving yourself another shot at the threshold, another roll of the dice for a false alarm, no new data required. Do that ten times over the course of an experiment and your &lt;em&gt;real&lt;/em&gt; false-positive rate isn't 5% anymore. It's roughly 4x that, around 20%.&lt;/p&gt;

&lt;p&gt;Here's the simple version of why. A single 5%-threshold look means a 5% chance of a false alarm, and a 95% chance of correctly seeing nothing when there's nothing there. That 95% is your "safe" probability for one look. Now take another look. If those looks were fully independent, both looks correctly showing nothing would need 95% and then 95% again, multiplying down toward zero the more times you check, ten independent looks would already put you near a 40% false-alarm rate.&lt;/p&gt;

&lt;p&gt;Real peeking lands lower than that naive multiplication, because your looks aren't independent, each one is checking the same accumulating data as the last, just with a few more rows added. But the direction holds and the size is still large: more looks, more chances to get unlucky, and nobody told the dashboard to warn you which chance you're on.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj3lt5aviny1h8oagzrtf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj3lt5aviny1h8oagzrtf.png" alt="Line chart showing the true type I error rate climbing from 5% at one peek to roughly 25% at ten peeks" width="800" height="462"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;FIG.4 — your chosen 5% only holds if you look exactly once&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Intuitively: a p-value during a running test isn't a fixed property of your data, it's a noisy quantity that jumps around as samples accumulate. It can dip under 0.05 early by pure chance, climb back over, dip again. If you stop the moment it happens to be favorable, you're selectively sampling the noise, not the effect. The p-value only reflects the error rate you actually chose once the pre-calculated full sample size is reached. Everything before that is a preview, not a verdict, no matter how green the dashboard looks.&lt;/p&gt;

&lt;p&gt;This doesn't mean the textbook rule, "never look, always wait for full sample size," is the only correct answer. It's correct &lt;em&gt;if&lt;/em&gt; you need that specific 5%/20% guarantee, which you usually do in regulated or high-stakes contexts. It's adjustable if you're willing to trade away some of that precision on purpose: agree on a looser threshold up front, or use a testing method actually designed to handle repeated looks (more on that in a minute). What's not adjustable is pretending you didn't peek when you did, and reporting the original 5% anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ WHY IT KEEPS HAPPENING ]
&lt;/h2&gt;

&lt;p&gt;This isn't a knowledge problem. Most people who work near experimentation can recite "don't peek" if you ask them directly. It keeps happening anyway, because the incentives in the room are almost never aligned with statistical patience.&lt;/p&gt;

&lt;p&gt;A test trending toward a strong result after one week, against a five-week pre-registered horizon, puts real pressure on the room: waiting the textbook full duration for rigor has an actual cost in forgone revenue or forgone learning speed. A test trending the wrong direction creates the mirror pressure: cutting losses early feels responsible, and also inflates your error rate exactly the same way a favorable early peek does. Shipping pressure doesn't care which direction the needle is moving. It pushes toward stopping early either way.&lt;/p&gt;

&lt;p&gt;The organizational failure mode compounds from here. A single test tells you a difference was detected, not the true magnitude of that difference. Teams that build quarterly OKRs by summing the observed uplifts of "winning" tests are stacking a series of noisy, possibly-inflated point estimates and treating the sum as a forecast. It rarely survives contact with the next two quarters. The statistics didn't fail here. The organization asked the statistics a question they were never built to answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ ALTERNATIVES, BRIEFLY ]
&lt;/h2&gt;

&lt;p&gt;None of this means fixed-horizon testing is broken or that peeking is unforgivable. It means the standard frequentist test answers one specific question, "is there a detectable difference at a preset error tolerance," and if that's not actually the question you're trying to answer, there are other tools built for the question you do have.&lt;/p&gt;

&lt;p&gt;Sequential testing and always-valid p-values (mSPRT, group sequential designs) are built specifically to let you look as often as you want without the error-rate blowup, by spending your error budget across looks instead of assuming a single look. Bayesian A/B testing reframes the whole question away from p-values entirely, toward "what's the probability B is actually better, and by how much," and it doesn't require a fixed sample size to interpret at all. Both of these are real, useful, and deserve more than a paragraph each, which is exactly why they're getting their own posts in this series rather than a rushed footnote here.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ THE MANAGERIAL PROBLEM ]
&lt;/h2&gt;

&lt;p&gt;Even once the math is right, there's a second problem that's arguably harder: communicating a probabilistic result to a stakeholder who wants a yes-or-no answer by end of day.&lt;/p&gt;

&lt;p&gt;"We're 87% confident this is an improvement, with a plausible range of 0 to 4%" is a true, useful sentence. It is also not the sentence most rooms want to hear. The fix isn't to round it down into a fake yes-or-no, it's to make the underlying decision structure explicit before you ever run the test, so uncertainty has somewhere to land. A workable hypothesis has four parts: &lt;strong&gt;if we do X, then Y will change by Z, because W.&lt;/strong&gt; X is the change, Y is the metric you're watching, Z is the effect size you actually expect (which is also what feeds your sample size calculation), and W is your causal reasoning for why you expect it. W is the part almost everyone skips, and it's the part that compounds: a single test only tells you &lt;em&gt;that&lt;/em&gt; something changed. Your working theory of &lt;em&gt;why&lt;/em&gt; is what turns one test into a body of knowledge instead of an isolated anecdote you can't reuse.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ WHAT WE ACTUALLY DO NOW ]
&lt;/h2&gt;

&lt;p&gt;The practical version of all of this, generalized, not tied to any one team's exact setup:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pre-register your sample size before you look at results, driven by your actual alpha, power, baseline, and MDE, not by vibes&lt;/li&gt;
&lt;li&gt;If you must look early, decide that up front and either use a method designed for it (sequential testing, Bayesian) or explicitly accept a looser real error rate&lt;/li&gt;
&lt;li&gt;Run continuous A/A tests on any traffic that isn't allocated to something real, they're nearly free and they catch splitter and tracking bugs before those bugs corrupt an actual test&lt;/li&gt;
&lt;li&gt;Write the X/Y/Z/W hypothesis down before launch, so "why" survives even when the test result is ambiguous&lt;/li&gt;
&lt;li&gt;Treat a null result as information, not failure, log it and move to the next hypothesis instead of quietly re-running until something turns green&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this eliminates uncertainty. It just makes sure the uncertainty you're shipping with is the uncertainty you actually chose, instead of one that quietly grew while nobody was looking.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ USEFUL LINKS ]
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.evanmiller.org/ab-testing/sample-size.html" rel="noopener noreferrer"&gt;Evan Miller's A/B testing sample size calculator&lt;/a&gt; — the fastest gut-check on whether you have anywhere near enough traffic before you run anything&lt;/li&gt;
&lt;li&gt;"The ASA Statement on p-Values: Context, Process, and Purpose" (Wasserstein &amp;amp; Lazar, &lt;em&gt;The American Statistician&lt;/em&gt;, 2016) — the closest thing statistics has to an official correction of the p-value misunderstanding in this post, worth searching up directly&lt;/li&gt;
&lt;li&gt;mSPRT / "always-valid p-values" (Johari, Koomen, Pekelis, Walsh) — the paper behind modern sequential testing at scale, more detail in the next post in this series&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  [ ELSEWHERE ]
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/4thwithme" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; · &lt;a href="https://www.linkedin.com/in/andrii-popenko-3331b3193/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://dev.to/4thwithme"&gt;Dev.to&lt;/a&gt; · &lt;a href="https://substack.com/@andriipopenko124891" rel="noopener noreferrer"&gt;Substack&lt;/a&gt; · &lt;a href="https://4thwithme.dev/blog" rel="noopener noreferrer"&gt;4thwithme.dev/blog&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;May the &lt;code&gt;--force&lt;/code&gt; be with you. See you next week.&lt;/p&gt;

</description>
      <category>statistics</category>
      <category>abtesting</category>
      <category>datascience</category>
      <category>career</category>
    </item>
    <item>
      <title>My dev setup</title>
      <dc:creator>4thwithme</dc:creator>
      <pubDate>Sun, 23 Aug 2026 22:04:53 +0000</pubDate>
      <link>https://dev.to/4thwithme/my-dev-setup-3p29</link>
      <guid>https://dev.to/4thwithme/my-dev-setup-3p29</guid>
      <description>&lt;h2&gt;
  
  
  [ HARDWARE ]
&lt;/h2&gt;

&lt;h3&gt;
  
  
  [ Ecosystem ]
&lt;/h3&gt;

&lt;p&gt;Let's start from ecosystem. Windows, Linux, and macOS all have their strengths, but my main requirements are:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Stability and reliability, so I can focus on work instead of rebooting or troubleshooting Wi-Fi, Sound, or Bluetooth drivers - so Linux is out.&lt;/li&gt;
&lt;li&gt;Well-supported, and well-documented UNIX development environment - so Windows is out.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So macOS it is.&lt;/p&gt;

&lt;h3&gt;
  
  
  [ Hardware Requirements ]
&lt;/h3&gt;

&lt;p&gt;Let's go to the hardware requirements. I need a machine that can run multiple tools at once without slowing down, and that can handle large codebases and datasets. Supporting Machine Learning and AI workloads is also a plus.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;In 2026, to run all the tools I need, 16GB RAM is the absolute minimum.&lt;/li&gt;
&lt;li&gt;SSD storage is a must for fast boot times and quick access to files.&lt;/li&gt;
&lt;li&gt;A high-resolution display is important for productivity.&lt;/li&gt;
&lt;li&gt;GPU support is a plus for machine learning and AI workloads.&lt;/li&gt;
&lt;li&gt;Long battery life is important for working on the go.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Right now I work on a MacBook Pro 14" with M4 Pro chip and 24GB RAM, which fulfills all the requirements above. I also have a 4K external monitor, mechanical keyboard, and mouse for a more comfortable and efficient workflow.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Split is roughly 60% external monitor + mechanical keyboard + mouse&lt;/li&gt;
&lt;li&gt;40% laptop as-is. No trackball, no split keyboard — I have many sides, but not that many.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  [ Karabiner ]
&lt;/h3&gt;

&lt;p&gt;Where the OS actually gets tuned is keyboard shortcuts.&lt;br&gt;
It lets me add custom shortcuts, remap keys, and create complex modifications. For example, I have a rule that allows me to switch between English and Ukrainian keyboard layouts with a single keypress.&lt;/p&gt;

&lt;p&gt;Language switching is bound directly. No cycling, no counting presses — you press the key for the language you want and you're already there.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ctrl+E jumps straight to [E]nglish&lt;/li&gt;
&lt;li&gt;Ctrl+U straight to [U]krainian.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo06kdhxgkyzn6101wgz5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo06kdhxgkyzn6101wgz5.png" alt="Karabiner-Elements complex modifications list" width="800" height="173"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;FIG.1 — one keybind, one destination — no rotation involved&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Switch to English input source with Control + E"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"manipulators"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"from"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"key_code"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"e"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"modifiers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"mandatory"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"control"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"to"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"select_input_source"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"language"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"en"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"basic"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  [ Mission Control ]
&lt;/h3&gt;

&lt;p&gt;The same logic runs the desktop layout. I keep six Spaces, each hard-assigned to one thing. Ctrl+1 - Ctrl+6 jump straight to the matching Space. No "next/previous desktop" cycling here either, same reasoning as the language switch: a direct destination beats a direction you have to count.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;browser&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;IDE&lt;/strong&gt; - my main editor, where I do the bulk of my coding and debugging&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;terminal&lt;/strong&gt; - my main workspace, where I run Claude Code, git, and other CLI tools&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;free space&lt;/strong&gt; for DB/Redis GUIs and OrbStack - my dev tools, local servers, and containers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Slack&lt;/strong&gt; - work chat, async comms, and incident response&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Obsidian&lt;/strong&gt; - my 2nd brain, where I keep my notes and plans.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F998r2jqa5qfp6m3awi9o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F998r2jqa5qfp6m3awi9o.png" alt="Mission Control keybinds: Ctrl+1 through Ctrl+6" width="800" height="400"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;FIG.2 — Ctrl+[number] beats Ctrl+left/right, every single time&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Ctrl+A pulls up the Mission Control overview when I actually need to see all six at once. Every window fills its Space at 100%, but I keep them in windowed-fill mode rather than native macOS fullscreen — same usable area, but the menu bar and dock stay one motion away instead of a whole gesture away.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxt4wvwiqgx6lgixq3g43.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxt4wvwiqgx6lgixq3g43.png" alt="Six macOS Spaces, one app per desktop" width="800" height="116"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;FIG.3 — six desktops, six fixed jobs — the same one is always in the same place&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  [ TERMINAL ]
&lt;/h2&gt;

&lt;p&gt;My main working tool and my main terminal is Warp. No &lt;code&gt;tmux&lt;/code&gt;, no &lt;code&gt;zellij&lt;/code&gt;, no &lt;code&gt;herdr&lt;/code&gt;, no multiplexer at all, even though I've used a few of them and they all worked fine. But we are living in 2026, not in the wild 2010s. Warp already has tabs, workspaces, and project-grouped sessions natively, with an AI layer on top. I have no nostalgia for multiplexers, sorry ;)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi2x9owuprtmzs9i7az2c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi2x9owuprtmzs9i7az2c.png" alt="Warp terminal with project-grouped tabs" width="799" height="519"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;FIG.4 — tabs grouped by project, not by panes I have to remember the layout of (branch names blurred, not a leak)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;One deliberate exception: Neovim gets its own terminal app entirely — Rio or Ghostty, kept separate from Warp. Not because Warp can't run it fine, but because I don't want my editor's session sharing chrome, history, and AI panels with everything else I'm doing. It's not the 90s, my machine can afford a second terminal app just so nvim gets a clean box to live in.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ EDITOR, BRIEFLY ]
&lt;/h2&gt;

&lt;p&gt;Three editors plus one thing that isn't really an editor. We will dive deep into Claude Code and &lt;code&gt;nvim&lt;/code&gt; in future posts. In the meantime, here's the rough split of my time:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Neovim gets 20% of my time.&lt;/li&gt;
&lt;li&gt;Zed - 15% of time, because it's fast and occasionally I have a mood for it.&lt;/li&gt;
&lt;li&gt;My old buddy VS Code - 15%. It was my first IDE and some habits just don't leave.&lt;/li&gt;
&lt;li&gt;The other 50% is Claude Code, and for that half of my week I don't open an editor at all. The work happens in the terminal, in diffs, in review — the editor becomes optional infrastructure instead of the default starting point - this is the reason why terminal has to be modern, fast, and AI-enabled.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F423edn8u9b25s3lialg1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F423edn8u9b25s3lialg1.png" alt="Neovim editing a Lua plugin config, cloak.nvim setup" width="800" height="502"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;FIG.5 — nvim, mid-config — fittingly, the plugin on screen is the one that hides my secrets&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Almost all my IDEs are configured with the same philosophy: minimal, fast, and keyboard-driven. I don't need a million panels and toolbars, no terminal embedded in the IDE, no DB connectors, no Git GUI, no debugger panels — I have other tools for all of that. I just need a fast editor that can handle large files and projects. If I cannot find a plugin or theme that fits my needs, I write one myself. I have a few of those.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ AI TOOLING ]
&lt;/h2&gt;

&lt;p&gt;I've been using Claude Code for almost two years now — long enough that opening a traditional IDE feels like the exception, not the default. I run it straight in the terminal, doing the actual work: planning, multi-file changes, the stuff that used to mean opening an IDE and doing it by hand. To stay aware of what a session is actually doing, I use the &lt;code&gt;ccstatusline&lt;/code&gt; plugin, which keeps session cost, model, and context usage visible so I'm not flying blind on a long-running task. GitHub Copilot's inline autocomplete still runs alongside it for the moments I'm typing directly — the two aren't competing, they're solving different-sized problems: Copilot finishes your line, Claude Code finishes your ticket.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ THE ACTUAL WORKFLOW ]
&lt;/h2&gt;

&lt;p&gt;My working day:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Review Claude Code — whatever overnight/async sessions did while I wasn't looking, leaving feedback or new instructions where needed.&lt;/li&gt;
&lt;li&gt;Then Slack: blockers, incidents, and whatever business or engineering questions came in overnight.&lt;/li&gt;
&lt;li&gt;Then New Relic and Rollbar for anything performance- or error-shaped, plus a daily AI-generated report that's already combed through the logs for me so I'm reading a summary instead of raw noise.&lt;/li&gt;
&lt;li&gt;Then PR review. I prefer to do it in the browser default GitHub format. When I see GitHub, my brain switches to "code review mode" and I can focus on the diff, and it is easier for me to think about bottlenecks, edge cases, and the overall quality of the code.&lt;/li&gt;
&lt;li&gt;Then JIRA for tickets, planning, and triage.&lt;/li&gt;
&lt;li&gt;Then finally I check my own plan for the day written in Obsidian — which is a whole separate post, because "second brain" is a real thing, and it deserves a real writeup.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  [ EVOLUTION ]
&lt;/h2&gt;

&lt;p&gt;My setup has been stable for years at its core, but I'm not precious about it — I'll happily give a new tool a real two-to-three month trial to see if it actually brings something new and cool to the workflow. Most of them don't survive contact with what I already have. Raycast got a fair shot and never earned a place — Spotlight already does what I need, and a second launcher just for the sake of having a "better" one wasn't worth the context-switching. Same story with a git GUI — not VS Code's, not LazyVim's — terminal git stays, mostly because I trust what I can see happening over what a panel summarizes for me. My latest churn: Rio → Ghostty → Rio, right back to where I started after giving Ghostty a genuine try. That's the pattern — I'm open to switching, the current setup just keeps winning.&lt;/p&gt;

&lt;h2&gt;
  
  
  [ ELSEWHERE ]
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/4thwithme" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; · &lt;a href="https://www.linkedin.com/in/andrii-popenko-3331b3193/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://dev.to/4thwithme"&gt;Dev.to&lt;/a&gt; · &lt;a href="https://substack.com/@andriipopenko124891" rel="noopener noreferrer"&gt;Substack&lt;/a&gt; · &lt;a href="https://4thwithme.dev/blog" rel="noopener noreferrer"&gt;4thwithme.dev/blog&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;May the &lt;code&gt;--force&lt;/code&gt; be with you. See you next week.&lt;/p&gt;

</description>
      <category>macos</category>
      <category>productivity</category>
      <category>neovim</category>
      <category>cli</category>
    </item>
    <item>
      <title>Starting my weekly blog (ML, agentic dev)</title>
      <dc:creator>4thwithme</dc:creator>
      <pubDate>Mon, 17 Aug 2026 10:01:42 +0000</pubDate>
      <link>https://dev.to/4thwithme/who-i-am-what-i-do-starting-a-weekly-blog-physics-to-em-ml-agentic-dev-5hnh</link>
      <guid>https://dev.to/4thwithme/who-i-am-what-i-do-starting-a-weekly-blog-physics-to-em-ml-agentic-dev-5hnh</guid>
      <description>&lt;p&gt;Post source: &lt;a href="https://4thwithme.dev/blog/intro/" rel="noopener noreferrer"&gt;https://4thwithme.dev/blog/intro/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nine years ago I was calculating neutron flux in a reactor core simulation. Today I spend my mornings in 1:1s and my evenings building MVPs, and ML recommendations models. Somewhere between those two sentences is the short version of how I got here. The reactor, for the record, did not explode.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1d6atzm9crq2zfj9jlne.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1d6atzm9crq2zfj9jlne.gif" alt="Springfield Nuclear Power Plant, the actual level of oversight I had, most days" width="450" height="264"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'm Andrii - an Engineering Manager based in Barcelona, currently leading a team of engineers at a promo-products ecommerce company. My path here was not exactly linear, it started with a physics degree instead of a computer science one: a Master's in Nuclear Power Engineering, a couple of years as a neutron physics engineer running Fortran and Python for reactor calculations, which is where the programming itch actually started, at the same time a few years teaching robotics and programming to kids who debugged with more patience than most senior engineers I've since worked with. Then a full pivot into software, no looking back, mostly because there was nothing to look back at except radiation shielding calculations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who I am
&lt;/h2&gt;

&lt;p&gt;If you want the honest ratio: I'm about 70% engineer, 30% manager. I still want to be in the code, still want to know why the model regressed and not just that it did, and I'd rather spend an evening tinkering with a training run or working through the math behind it than sit through another roadmap slide. Management is real work and I take it seriously, but the itch that got me here in the first place is still curiosity about how things actually work, down to the math.&lt;/p&gt;

&lt;h3&gt;
  
  
  day job
&lt;/h3&gt;

&lt;p&gt;I lead the team responsible for recommendation and search on a storefront that has strong opinions about branded mugs, the kind of place you order 200 branded pens from and never think about again until the next company retreat. Lately my job is less "ship the feature" and more "figure out how a whole engineering org adopts AI tooling without either turning it into a personality or pretending it doesn't exist." Both camps are loud. I try to be the annoying third option that just ships things.&lt;/p&gt;

&lt;h3&gt;
  
  
  outside the day job
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9i2jdl8ucl6tbigtfxqb.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9i2jdl8ucl6tbigtfxqb.gif" alt="Claude Code meme: nature-documentary footage of someone coding manually" width="360" height="337"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'm an AI/ML enthusiast in the fullest, most unqualified-to-stop-talking-about-it sense, hands-on with code, and hands-on with hardware too. I like tinkering with models, sitting with the math until it actually clicks instead of just pattern-matching the notation, and reading research papers slower than is efficient because skimming defeats the point. I write Neovim plugins nobody asked for, I'm working through computer vision fundamentals, and I have a Three.js habit that has produced zero shipped products and several very nice spinning cubes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building / learning
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Computer vision, slowly, deliberately, mostly on Sundays&lt;/li&gt;
&lt;li&gt;The math under the models I use daily, not just the API surface&lt;/li&gt;
&lt;li&gt;A Neovim plugin nobody asked for (&lt;code&gt;ss.nvim&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;An unreasonable Three.js habit with a 0% shipped-product conversion rate&lt;/li&gt;
&lt;li&gt;Whatever this blog forces me to actually finish, since public commitment works better than a private todo list&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why this blog
&lt;/h2&gt;

&lt;p&gt;"Cold start" is a recommender-systems term. No data on a new user or item, and you still have to guess something useful, confidently, with nothing to go on. That's also just what most Mondays feel like. So the rule is simple: &lt;strong&gt;one honest post a week, no excuses&lt;/strong&gt;, a threat I'm making primarily to myself, and no quiet retirement of this blog in week six like every New Year's resolution I've ever made.&lt;/p&gt;

&lt;p&gt;Expect posts on ML, agentic development, Claude Code, nvim, new tools I'm evaluating, and the occasional controversial take I'll defend confidently for about a week.&lt;/p&gt;

&lt;h2&gt;
  
  
  Elsewhere
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/4thwithme" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; · &lt;a href="https://www.linkedin.com/in/andrii-popenko-3331b3193/" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://substack.com/@andriipopenko124891" rel="noopener noreferrer"&gt;Substack&lt;/a&gt; · &lt;a href="https://4thwithme.dev/blog" rel="noopener noreferrer"&gt;4thwithme.dev/blog&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;May the &lt;code&gt;--force&lt;/code&gt; be with you. See you next week.&lt;/p&gt;

</description>
      <category>career</category>
      <category>ai</category>
      <category>machinelearning</category>
      <category>watercooler</category>
    </item>
  </channel>
</rss>
