<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Samuel Mutemi</title>
    <description>The latest articles on DEV Community by Samuel Mutemi (@mtsammy40).</description>
    <link>https://dev.to/mtsammy40</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1368344%2Fdc35f5bb-1ed9-4b80-a7e4-e76041372894.jpg</url>
      <title>DEV Community: Samuel Mutemi</title>
      <link>https://dev.to/mtsammy40</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mtsammy40"/>
    <language>en</language>
    <item>
      <title>The Two Joints: Where Agentic Engineering Breaks</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Tue, 18 Aug 2026 18:30:35 +0000</pubDate>
      <link>https://dev.to/mtsammy40/the-two-joints-where-agentic-engineering-breaks-52lm</link>
      <guid>https://dev.to/mtsammy40/the-two-joints-where-agentic-engineering-breaks-52lm</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge" class="crayons-story__hidden-navigation-link"&gt;Agentic Engineering Is Not Vibe Coding: The Three-Skill Loop I Use to Ship Distributed Systems&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/mtsammy40" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1368344%2Fdc35f5bb-1ed9-4b80-a7e4-e76041372894.jpg" alt="mtsammy40 profile" class="crayons-avatar__image" width="800" height="801"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/mtsammy40" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Samuel Mutemi
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Samuel Mutemi
                
                
              
              &lt;div id="story-author-preview-content-4294299" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/mtsammy40" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1368344%2Fdc35f5bb-1ed9-4b80-a7e4-e76041372894.jpg" class="crayons-avatar__image" alt="" width="800" height="801"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Samuel Mutemi&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 2&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge" id="article-link-4294299"&gt;
          Agentic Engineering Is Not Vibe Coding: The Three-Skill Loop I Use to Ship Distributed Systems
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/webdev"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;webdev&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/programming"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;programming&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/automation"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;automation&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;2&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              1&lt;span class="hidden s:inline"&gt;&amp;nbsp;comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            11 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;



&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/mtsammy40/where-does-truth-live-a-field-guide-to-agentic-engineering-methodologies-23hg" class="crayons-story__hidden-navigation-link"&gt;Where Does Truth Live? A Field Guide to Agentic Engineering Methodologies&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/mtsammy40" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1368344%2Fdc35f5bb-1ed9-4b80-a7e4-e76041372894.jpg" alt="mtsammy40 profile" class="crayons-avatar__image" width="800" height="801"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/mtsammy40" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Samuel Mutemi
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Samuel Mutemi
                
                
              
              &lt;div id="story-author-preview-content-4313587" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/mtsammy40" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1368344%2Fdc35f5bb-1ed9-4b80-a7e4-e76041372894.jpg" class="crayons-avatar__image" alt="" width="800" height="801"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Samuel Mutemi&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/mtsammy40/where-does-truth-live-a-field-guide-to-agentic-engineering-methodologies-23hg" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 4&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/mtsammy40/where-does-truth-live-a-field-guide-to-agentic-engineering-methodologies-23hg" id="article-link-4313587"&gt;
          Where Does Truth Live? A Field Guide to Agentic Engineering Methodologies
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/agents"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;agents&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/programming"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;programming&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/productivity"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;productivity&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/mtsammy40/where-does-truth-live-a-field-guide-to-agentic-engineering-methodologies-23hg" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;1&lt;span class="hidden s:inline"&gt;&amp;nbsp;reaction&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/mtsammy40/where-does-truth-live-a-field-guide-to-agentic-engineering-methodologies-23hg#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            16 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


&lt;p&gt;Intent — the spec says what correct means.&lt;/p&gt;

&lt;p&gt;Verification — tests check whether you got it.&lt;/p&gt;

&lt;p&gt;Correction — a loop closes the gap.&lt;/p&gt;

&lt;p&gt;This post is what happened when I stopped running those as three separate practices and wired them into one system.&lt;/p&gt;

&lt;p&gt;The layers were the easy part.&lt;/p&gt;

&lt;p&gt;The hard part is where they join, and there are &lt;strong&gt;exactly two joints&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Real Problem: I Was the Wiring
&lt;/h2&gt;

&lt;p&gt;Running three methodologies "concurrently" looked like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The spec lived in GitHub Issues.&lt;/li&gt;
&lt;li&gt;The tests lived in the repo.&lt;/li&gt;
&lt;li&gt;The loop lived in a skill that re-ran an agent until a task list emptied.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All three were real.&lt;/p&gt;

&lt;p&gt;The connection between them was &lt;strong&gt;me, noticing things&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I noticed that a failing test meant a PRD assumption was wrong.&lt;/p&gt;

&lt;p&gt;I noticed that slices five through nine came from a version of the spec that no longer existed.&lt;/p&gt;

&lt;p&gt;I noticed that an agent had "fixed" a failing test by making the test weaker.&lt;/p&gt;

&lt;p&gt;That's not a system.&lt;/p&gt;

&lt;p&gt;That's three components and an operator, and &lt;strong&gt;the operator is the part that doesn't scale&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Each layer had a trigger, an input and an output.&lt;/p&gt;

&lt;p&gt;None of them had a defined interface to the layer next door, so the interface defaulted to my attention.&lt;/p&gt;

&lt;p&gt;So the goal isn't:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Adopt all three."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Give the handoffs between layers a name, a trigger and an artifact, so they happen without me.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  The Flow
&lt;/h2&gt;

&lt;p&gt;The system looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                         SPEC
                    (what correct means)
                           |
                           v
                 +---------+---------+
                 |                   |
                 v                   v
            work slices        tests written
                              and FROZEN first
                 |                   |
                 v                   |
          agent implements           |
                 |                   |
                 +---------+---------+
                           |
                           v
                      run the tests
                           |
                 +---------+---------+
                 |                   |
                 v                   v
             code is wrong       spec looks wrong
             INNER LOOP          OUTER LOOP
             retry N times       STOP, no retries
             agent fixes it      human decides
                 |                   |
                 v                   v
             human gate,        amend the spec,
             by blast radius    re-derive the rest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two loops.&lt;/p&gt;

&lt;p&gt;The difference between them is the whole design:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The inner loop fixes the code.&lt;/strong&gt;&lt;br&gt;
Runs unsupervised, on a retry budget, allowed to act.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The outer loop fixes the spec.&lt;/strong&gt;&lt;br&gt;
Runs on escalation, no retry budget, allowed only to propose.&lt;/p&gt;

&lt;p&gt;Changing the spec means changing intent, and intent isn't something an agent gets to change by itself.&lt;/p&gt;

&lt;p&gt;A loop that can edit its own target will move the target instead of solving the problem.&lt;/p&gt;

&lt;p&gt;Every time.&lt;/p&gt;

&lt;p&gt;The rest of this post is those two joints.&lt;/p&gt;


&lt;h1&gt;
  
  
  Joint One: Your Tests Can't Just Be a Copy of Your Spec
&lt;/h1&gt;

&lt;p&gt;Part two ended on this claim, and it's harder than it sounds:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A feedback loop with a bad sensor is worse than no loop, because it converges, confidently, at machine speed, on the wrong thing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The obvious move is to generate your tests from your spec.&lt;/p&gt;

&lt;p&gt;That's the spec-driven pitch, and it works.&lt;/p&gt;

&lt;p&gt;But it buys you exactly one thing:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It catches the code disagreeing with the spec.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Now notice what it can't catch.&lt;/p&gt;

&lt;p&gt;If the spec is wrong about the real world, a test generated from that spec is wrong in the same direction.&lt;/p&gt;

&lt;p&gt;It passes.&lt;/p&gt;

&lt;p&gt;Green checkmark.&lt;/p&gt;

&lt;p&gt;Everyone goes home.&lt;/p&gt;

&lt;p&gt;Safety engineering has a name for this: &lt;strong&gt;common-cause failure&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Two backups don't help if they share a design flaw, which is why aircraft use sensors built on different physical principles instead of just duplicating one.&lt;/p&gt;

&lt;p&gt;Your spec-generated test suite is one channel, and its design is the spec.&lt;/p&gt;
&lt;h2&gt;
  
  
  What Can the Test Actually Catch?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What went wrong&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;th&gt;Can a spec-generated test catch it?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Code doesn't match the spec&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;PRD says retry 3 times, code retries once&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Yes.&lt;/strong&gt; This is its job&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Spec contradicts itself&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Requires strict ordering on a queue that doesn't guarantee it&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Yes, cheaply&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Spec is silent&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Says nothing about ordering, so the agent picks one&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;No.&lt;/strong&gt; No test exists for an unstated rule&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Spec is wrong about the world&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"The provider's charge endpoint is idempotent." It isn't&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;No.&lt;/strong&gt; The test asserts the same falsehood&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Spec is right, the requirement was bad&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Flawless build of the wrong feature&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;No&lt;/strong&gt;, and no test ever will&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Rows three and four are where money gets lost.&lt;/p&gt;

&lt;p&gt;The fix isn't a better prompt.&lt;/p&gt;

&lt;p&gt;It's a &lt;strong&gt;second kind of test&lt;/strong&gt;.&lt;/p&gt;


&lt;h2&gt;
  
  
  Two Kinds of Tests, and Only One Can Challenge the Spec
&lt;/h2&gt;

&lt;p&gt;I now split tests by what they're measured against, not by unit/integration/e2e.&lt;/p&gt;
&lt;h3&gt;
  
  
  Spec Tests
&lt;/h3&gt;

&lt;p&gt;Generated from the spec.&lt;/p&gt;

&lt;p&gt;Measured against intent.&lt;/p&gt;

&lt;p&gt;Property tests over the stated invariants, assertions on the out-of-scope list.&lt;/p&gt;

&lt;p&gt;Cheap, generated, and I make a lot of them.&lt;/p&gt;

&lt;p&gt;They ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Does the code do what we said?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  Reality Tests
&lt;/h3&gt;

&lt;p&gt;Not generated from the spec.&lt;/p&gt;

&lt;p&gt;Measured against the world.&lt;/p&gt;

&lt;p&gt;Provider sandbox checks, fault injection, replays of real incidents, reconciliation against the provider's ledger.&lt;/p&gt;

&lt;p&gt;Expensive, mostly hand-written, they accumulate slowly.&lt;/p&gt;

&lt;p&gt;They ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Does this survive contact?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The line that reorganized my thinking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Only a test that didn't come from the spec can prove the spec wrong.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Anything derived from the target can confirm the target.&lt;/p&gt;

&lt;p&gt;Nothing derived from it can contradict it.&lt;/p&gt;

&lt;p&gt;Which gives a hard rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A failing spec test is never grounds for changing the spec. It means the code is wrong. Only a failing reality test is allowed to say the spec is wrong.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sounds like bureaucracy until the first time an agent responds to a hard failing test by suggesting you amend the PRD.&lt;/p&gt;

&lt;p&gt;Which it will, because that's the cheapest path to green.&lt;/p&gt;


&lt;h2&gt;
  
  
  Three Habits That Keep the Spec Tests Honest
&lt;/h2&gt;
&lt;h3&gt;
  
  
  1. Write the Tests Before the Code — and Freeze Them
&lt;/h3&gt;

&lt;p&gt;Not the same session.&lt;/p&gt;

&lt;p&gt;Not after there's an implementation to look at.&lt;/p&gt;

&lt;p&gt;Sequencing buys independence almost free.&lt;/p&gt;

&lt;p&gt;The agent writing the code may not edit its own tests.&lt;/p&gt;

&lt;p&gt;If a slice genuinely needs a test changed, that's not a code change.&lt;/p&gt;

&lt;p&gt;It's a claim about intent, and it escalates.&lt;/p&gt;

&lt;p&gt;This one rule killed my worst failure mode.&lt;/p&gt;
&lt;h3&gt;
  
  
  2. Hunt the Silences
&lt;/h3&gt;

&lt;p&gt;The clarifying-questions pass from &lt;code&gt;/to-prd&lt;/code&gt; was aimed at my rough prompt.&lt;/p&gt;

&lt;p&gt;I now run it again on the finished spec, asking only:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What does this not say that an implementer would have to guess?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Each answer becomes a spec edit or a recorded "don't care."&lt;/p&gt;

&lt;p&gt;Both are durable.&lt;/p&gt;

&lt;p&gt;An agent's undocumented assumptions are the most dangerous thing in the system, because they're the only part that isn't written down anywhere.&lt;/p&gt;


&lt;h1&gt;
  
  
  Joint Two: What Happens When the Loop Decides the Target Was Wrong
&lt;/h1&gt;
&lt;h2&gt;
  
  
  The Failure This Exists to Prevent
&lt;/h2&gt;

&lt;p&gt;A loop with a retry budget and no escape hatch doesn't stop when the spec is wrong.&lt;/p&gt;

&lt;p&gt;It keeps going, because that's what loops do.&lt;/p&gt;

&lt;p&gt;And since it can't reach the spec, it reduces the error the only other way available:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;by weakening the test.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Relax the assertion.&lt;/p&gt;

&lt;p&gt;Widen the tolerance.&lt;/p&gt;

&lt;p&gt;Skip the case with a plausible comment.&lt;/p&gt;

&lt;p&gt;Special-case the failing input.&lt;/p&gt;

&lt;p&gt;Each move is locally reasonable, and each one lies to you afterward.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A loop that can weaken its own tests will always converge. That's not a feature, that's the bug.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So the outer loop isn't mainly a correction mechanism.&lt;/p&gt;

&lt;p&gt;It's a &lt;strong&gt;stop button&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;And stopping is its most important capability.&lt;/p&gt;


&lt;h2&gt;
  
  
  The Stop, and the Four Verdicts
&lt;/h2&gt;

&lt;p&gt;When a reality test fails in a way the inner loop can't fix, the loop halts that branch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zero retries.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because retrying a spec error just buys more attempts at the wrong problem.&lt;/p&gt;

&lt;p&gt;Then it files a spec challenge that blocks the parent PRD.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;spec_challenge&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;PRD-412 @ a3f19c2&lt;/span&gt;

  &lt;span class="na"&gt;assumption&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;provider&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;charge&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;endpoint&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;is&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;idempotent&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;our&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;key"&lt;/span&gt;

  &lt;span class="na"&gt;contradicted_by&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;reality/provider-sandbox-conformance&lt;/span&gt;
    &lt;span class="c1"&gt;# must be a reality test&lt;/span&gt;

    &lt;span class="na"&gt;observed&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;duplicate&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;charge&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;retry&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;identical&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;key"&lt;/span&gt;

    &lt;span class="na"&gt;reproduced&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;5&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;of&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;5&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;runs"&lt;/span&gt;

  &lt;span class="na"&gt;blast_radius&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;slices_blocked&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;4&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;5&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;7&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;slices_already_merged_on_this_assumption&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;1&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;2&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

  &lt;span class="na"&gt;smallest_fix_proposed&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;key on provider_ref only; add 15m reconciliation sweep&lt;/span&gt;
    &lt;span class="s"&gt;against provider ledger.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four fields do the work.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Test Class
&lt;/h3&gt;

&lt;p&gt;Because it decides who's even allowed to file this.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Slices Already Merged
&lt;/h3&gt;

&lt;p&gt;Because that's the expensive number and I want it before I decide anything.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Smallest Fix
&lt;/h3&gt;

&lt;p&gt;Because "smallest" is load-bearing.&lt;/p&gt;

&lt;p&gt;An unconstrained agent will propose rewriting the spec around its discovery.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Reproduced
&lt;/h3&gt;

&lt;p&gt;Because "failed once" and "failed five out of five" are different conversations.&lt;/p&gt;

&lt;p&gt;Then a human decides.&lt;/p&gt;

&lt;p&gt;There are exactly four answers, and naming them turned a judgment call into triage:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The spec was wrong.&lt;/strong&gt; Amend it, cascade downstream. The intended path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The test was wrong.&lt;/strong&gt; Fix the test. If it was a spec test, that usually means the requirement was ambiguous, so the spec gets touched too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The world is wrong.&lt;/strong&gt; The provider's docs lie. The spec gets a workaround with an expiry date and a ticket, because undated workarounds become permanent architecture.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The requirement was wrong.&lt;/strong&gt; Not the wording, the requirement. This leaves engineering entirely. Good systems just surface it early and cheaply.&lt;/li&gt;
&lt;/ol&gt;




&lt;h1&gt;
  
  
  Stamp Everything With the Spec Version
&lt;/h1&gt;

&lt;p&gt;An amendment describes what changed, not the whole world again.&lt;/p&gt;

&lt;p&gt;Then it propagates, in three tiers with very different costs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Not Started Yet
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Regenerate.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Free.&lt;/p&gt;

&lt;p&gt;This is the entire payoff of spec-driven work.&lt;/p&gt;

&lt;h3&gt;
  
  
  In Flight
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Halt, throw away, re-derive.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Costs tokens and nothing else.&lt;/p&gt;

&lt;p&gt;Implementations are disposable; that was the bet.&lt;/p&gt;

&lt;h3&gt;
  
  
  Already Merged
&lt;/h3&gt;

&lt;p&gt;The one nobody warns you about.&lt;/p&gt;

&lt;p&gt;It isn't automatically wrong, but it is now unverified and has to be re-checked against the amended spec.&lt;/p&gt;

&lt;p&gt;So every derived artifact carries the spec revision it came from:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Issues&lt;/li&gt;
&lt;li&gt;Generated tests&lt;/li&gt;
&lt;li&gt;Commit trailers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Derived-From: PRD-412@a3f19c2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Unglamorous, and it converts the worst question in the system from an archaeology dig into a search.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Which merged work came from a spec that has since changed?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Without stamps, the honest answer after three amendments is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Some of it, let me read the git log."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;With them, you get a list.&lt;/p&gt;

&lt;p&gt;And lists can be worked.&lt;/p&gt;




&lt;h2&gt;
  
  
  One More Guardrail
&lt;/h2&gt;

&lt;p&gt;If the same assumption gets amended twice in one feature, I stop.&lt;/p&gt;

&lt;p&gt;Two amendments mean I'm not correcting my model of the system.&lt;/p&gt;

&lt;p&gt;I'm searching for one.&lt;/p&gt;

&lt;p&gt;That's a spike, run deliberately in a throwaway branch, with nothing derived from it until it finishes.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A target that moves every time a test complains isn't a target.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  What Got Worse
&lt;/h1&gt;

&lt;p&gt;It's worth being honest about what this architecture doesn't solve.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reality Tests Are Expensive
&lt;/h2&gt;

&lt;p&gt;Reality tests are expensive and don't generate well.&lt;/p&gt;

&lt;p&gt;That's the honest asymmetry here.&lt;/p&gt;

&lt;p&gt;The channel I need most for correctness is the one that can't be derived from the artifact I have, by definition.&lt;/p&gt;

&lt;p&gt;Fault injection.&lt;/p&gt;

&lt;p&gt;Real sandboxes.&lt;/p&gt;

&lt;p&gt;Recorded incidents.&lt;/p&gt;

&lt;p&gt;It accumulates at human speed.&lt;/p&gt;

&lt;p&gt;Anyone selling fully generated verification is selling you one channel and calling it two.&lt;/p&gt;

&lt;h2&gt;
  
  
  Escalation Fatigue
&lt;/h2&gt;

&lt;p&gt;Which is the same bottleneck again.&lt;/p&gt;

&lt;p&gt;If the loop escalates too readily, spec challenges pile up and get rubber-stamped exactly the way oversized diffs do.&lt;/p&gt;

&lt;p&gt;I've moved the constraint from reviewing code to deciding about intent.&lt;/p&gt;

&lt;p&gt;That's a better place for it, far more leverage per decision, but part one's point holds and it's recursive:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;You don't eliminate a bottleneck, you relocate it. Then you go find it again.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  More Artifacts to Keep True
&lt;/h2&gt;

&lt;p&gt;Spec.&lt;/p&gt;

&lt;p&gt;Spec tests.&lt;/p&gt;

&lt;p&gt;Reality tests.&lt;/p&gt;

&lt;p&gt;Version stamps.&lt;/p&gt;

&lt;p&gt;Part two warned that a hybrid fails as a stale spec that agents still trust.&lt;/p&gt;

&lt;p&gt;Version stamps make that detectable, not impossible.&lt;/p&gt;




&lt;h1&gt;
  
  
  What I Still Don't Have
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Provenance Tooling
&lt;/h2&gt;

&lt;p&gt;No tooling for any of the provenance.&lt;/p&gt;

&lt;p&gt;Commit trailers and issue labels, which is to say conventions and discipline, which is to say it will decay.&lt;/p&gt;

&lt;p&gt;This feels like something that should exist.&lt;/p&gt;

&lt;p&gt;If it does, tell me.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Next Generation of Senior Engineers
&lt;/h2&gt;

&lt;p&gt;No answer to part one's open question.&lt;/p&gt;

&lt;p&gt;If the mechanical work is where judgment used to get manufactured, and that work is now agentic, where do the next senior engineers come from?&lt;/p&gt;

&lt;p&gt;Deciding spec challenges is excellent practice for exactly the skill that matters, and it's also the task I'd hand to the most experienced person in the room.&lt;/p&gt;

&lt;p&gt;Which means it isn't a training ground.&lt;/p&gt;

&lt;p&gt;I'm suspicious of anyone who claims to have solved this.&lt;/p&gt;




&lt;h1&gt;
  
  
  Where This Leaves the Series
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Part one:&lt;/strong&gt; your job moved from building the system to designing the system that builds the system.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part two:&lt;/strong&gt; sort the methodologies by where truth lives between runs, and you get three layers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part three:&lt;/strong&gt; the layers were the easy part, and a hybrid is defined by its joints, not its parts.&lt;/p&gt;

&lt;p&gt;If one thing survives out of all of it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Both joints are the same problem stated twice.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Independence has to be engineered in.&lt;/p&gt;

&lt;p&gt;It is never the default.&lt;/p&gt;

&lt;p&gt;A test derived from the spec can't correct the spec, and a loop that can edit its target will move the target instead of doing the work.&lt;/p&gt;

&lt;p&gt;And part two's caveat applies here more than anywhere.&lt;/p&gt;

&lt;p&gt;A lot of this is model-specific error correction with a good name on it.&lt;/p&gt;

&lt;p&gt;The rule about tests exists because today's models will happily weaken one to reach green.&lt;/p&gt;

&lt;p&gt;The zero-retry rule exists because they'll happily rewrite a target to hit it.&lt;/p&gt;

&lt;p&gt;If a future model reliably refuses both, some of this becomes scar tissue and should be cut.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Notice which of your rituals are load-bearing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Retune when you change models.&lt;/p&gt;

&lt;p&gt;And if you've built the same shape and hit a different joint, tell me, because I'm fairly sure I've only found two of them.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part three of a series on agentic engineering in production. Part one covered the three-skill workflow and where the bottleneck goes when code stops being the constraint. Part two mapped the methodology landscape by asking where truth lives between agent runs.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>agents</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Where Does Truth Live? A Field Guide to Agentic Engineering Methodologies</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Tue, 04 Aug 2026 13:48:42 +0000</pubDate>
      <link>https://dev.to/mtsammy40/where-does-truth-live-a-field-guide-to-agentic-engineering-methodologies-23hg</link>
      <guid>https://dev.to/mtsammy40/where-does-truth-live-a-field-guide-to-agentic-engineering-methodologies-23hg</guid>
      <description>&lt;p&gt;Last time I wrote about the three-skill workflow I use to ship payment infrastructure, and about where the bottleneck went once code stopped being the bottleneck.&lt;/p&gt;


&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge" class="crayons-story__hidden-navigation-link"&gt;Agentic Engineering Is Not Vibe Coding: The Three-Skill Loop I Use to Ship Distributed Systems&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/mtsammy40" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1368344%2Fdc35f5bb-1ed9-4b80-a7e4-e76041372894.jpg" alt="mtsammy40 profile" class="crayons-avatar__image" width="800" height="801"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/mtsammy40" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Samuel Mutemi
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Samuel Mutemi
                
              
              &lt;div id="story-author-preview-content-4294299" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/mtsammy40" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1368344%2Fdc35f5bb-1ed9-4b80-a7e4-e76041372894.jpg" class="crayons-avatar__image" alt="" width="800" height="801"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Samuel Mutemi&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 2&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge" id="article-link-4294299"&gt;
          Agentic Engineering Is Not Vibe Coding: The Three-Skill Loop I Use to Ship Distributed Systems
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/webdev"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;webdev&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/programming"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;programming&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/automation"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;automation&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;2&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              1&lt;span class="hidden s:inline"&gt;&amp;nbsp;comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            11 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success crayons-icon c-btn__icon"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


&lt;p&gt;I thought it was prudent to answer a question I often get when discussing my workflow; 'Why this approach?'. There are a dozen named methodologies in this space now, several of them with more GitHub stars than my entire company has customers, and I picked a particular combination for particular reasons.&lt;/p&gt;

&lt;p&gt;So this post is the survey I wish I'd had when I started. I'll walk the landscape first, honestly, including the parts of it I rejected. Then I'll tell you why I landed on spec-driven development plus loop engineering, and what each half is actually doing.&lt;/p&gt;




&lt;h2&gt;
  
  
  The organizing question
&lt;/h2&gt;

&lt;p&gt;Here's the lens that made this landscape legible to me.&lt;/p&gt;

&lt;p&gt;Context windows are volatile memory. They fill, they compact, they lose the thread. Every session ends and takes its understanding with it. So every agentic methodology, every single one, whether or not its authors frame it this way, is an answer to one question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where does truth live between agent runs?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's it. That's the whole taxonomy. Sort the methodologies by where they park durable state and the differences between them stop being marketing and start being architecture.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Methodology&lt;/th&gt;
&lt;th&gt;Truth lives in&lt;/th&gt;
&lt;th&gt;Fails when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Vibe coding&lt;/td&gt;
&lt;td&gt;The chat session&lt;/td&gt;
&lt;td&gt;The session ends&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context engineering&lt;/td&gt;
&lt;td&gt;The assembled context / repo files&lt;/td&gt;
&lt;td&gt;Assembly is wrong or stale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spec-driven development&lt;/td&gt;
&lt;td&gt;A durable spec artifact&lt;/td&gt;
&lt;td&gt;The spec drifts from reality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Loop-based autonomy&lt;/td&gt;
&lt;td&gt;The filesystem + git history&lt;/td&gt;
&lt;td&gt;The stop condition is wrong&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TDD governance&lt;/td&gt;
&lt;td&gt;The test suite&lt;/td&gt;
&lt;td&gt;Tests encode the wrong contract&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-agent orchestration&lt;/td&gt;
&lt;td&gt;The orchestrator's task graph&lt;/td&gt;
&lt;td&gt;A handoff silently corrupts&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Let's go through them.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Prompt-native development (vibe coding)
&lt;/h2&gt;

&lt;p&gt;Truth lives in the chat. You describe, you get code, you eyeball it, you continue.&lt;/p&gt;

&lt;p&gt;I covered this in part one so I won't relitigate it. I'll just note that it remains the correct tool for spikes, throwaway scripts, and exploring an unfamiliar API, and that its defining property is that nothing survives the session. When the context compacts, your intent compacts with it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where it lives on the map:&lt;/strong&gt; the origin point. Everything else is a strategy for surviving what this one loses.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Context engineering (and its cousin, harness engineering)
&lt;/h2&gt;

&lt;p&gt;Truth lives in the assembled context: &lt;code&gt;AGENTS.md&lt;/code&gt;, &lt;code&gt;CLAUDE.md&lt;/code&gt;, repo conventions, retrieval strategy, file layout designed for agent navigation.&lt;/p&gt;

&lt;p&gt;This is the discipline of making sure that whatever the agent reads at the top of a run actually orients it correctly. There's now real research on it, a body of work on structured context engineering for file-native agentic systems, and an emerging line on &lt;a href="https://arxiv.org/abs/2602.14690" rel="noopener noreferrer"&gt;"harness engineering"&lt;/a&gt;, which studies the scaffolding around the model rather than the model itself.&lt;/p&gt;

&lt;p&gt;I don't think context engineering is a &lt;em&gt;rival&lt;/em&gt; methodology. I think it's a substrate. Every other approach on this list is doing context engineering whether it admits it or not, a spec is context, a task list is context, a test failure is context. The people who talk about it explicitly are just being honest about the mechanism.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My take:&lt;/strong&gt; necessary, not sufficient. A perfectly engineered context still doesn't tell you what to build, or how you'll know it worked.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Spec-driven development
&lt;/h2&gt;

&lt;p&gt;Truth lives in a specification artifact that outlives every session.&lt;/p&gt;

&lt;p&gt;This is the biggest cluster by far, and it's worth being precise, because "SDD" now names at least four meaningfully different things. The &lt;a href="https://en.wikipedia.org/wiki/Spec-driven_development" rel="noopener noreferrer"&gt;common thread&lt;/a&gt; is that a structured specification, usually Markdown, sometimes machine-readable, becomes the authoritative source from which implementation, tests, and docs are derived, rather than documentation written retrospectively.&lt;/p&gt;

&lt;p&gt;The variants differ mostly in how much ceremony they impose and how far they push the spec toward being literal source code.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lean and constitution-driven, GitHub Spec Kit
&lt;/h3&gt;

&lt;p&gt;A four-phase workflow: specify, plan, tasks, implement. Every spec inherits from a project-wide "constitution" that encodes durable rules, stack, conventions, non-negotiables. GitHub's distribution has made it the highest-star option in the category.&lt;/p&gt;

&lt;p&gt;Good: minimal ceremony, tool-agnostic, the constitution idea is genuinely excellent. Less good: it's optimized for greenfield and for reasonably-sized changes; small edits fight the phase structure.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lightweight and brownfield-first, OpenSpec
&lt;/h3&gt;

&lt;p&gt;Purpose-built for &lt;em&gt;modifications&lt;/em&gt; rather than net-new work, using delta markers, ADDED, MODIFIED, REMOVED, so a change proposal describes what shifts rather than restating the world.&lt;/p&gt;

&lt;p&gt;If your reality is a mature codebase where every task is an amendment, this is the most honest model on the list. It's also the cheapest to run.&lt;/p&gt;

&lt;h3&gt;
  
  
  Full agile simulation, BMAD-METHOD
&lt;/h3&gt;

&lt;p&gt;The maximalist option: a dozen-plus specialized agent personas, Analyst, PM, Architect, Scrum Master, Developer, QA, passing artifacts down a pipeline that mirrors a full SDLC.&lt;/p&gt;

&lt;p&gt;I want to be fair to BMAD, because it's genuinely impressive and it solves a real problem for teams that already think in PRDs and sprint stories. But the critique I find most convincing is structural: &lt;strong&gt;a persona pipeline is only as good as its weakest handoff.&lt;/strong&gt; When the Architect makes an assumption the PM never documented, the Scrum Master faithfully propagates it into stories and the Developer implements it with total confidence. You discover it in QA, or in production. More personas means more handoff surface, and handoff failures are a nasty debugging class because every individual agent behaved reasonably.&lt;/p&gt;

&lt;p&gt;The cost is also non-trivial. One consultancy &lt;a href="https://reenbit.com/bmad-vs-spec-kit-vs-openspec-choosing-your-spec-driven-ai-framework/" rel="noopener noreferrer"&gt;reports&lt;/a&gt; BMAD runs averaging in the low tens of thousands of tokens per workflow and monthly frontier-model bills in the high hundreds to low thousands of dollars per developer. Your mileage will vary enormously, but it's not a rounding error.&lt;/p&gt;

&lt;h3&gt;
  
  
  Environment-native, Kiro, and the platform tier
&lt;/h3&gt;

&lt;p&gt;AWS Kiro is a full IDE with the spec workflow built in from the ground up rather than layered on top. Same for various "agentic development environment" platforms that coordinate agents around a shared living spec.&lt;/p&gt;

&lt;p&gt;The trade is explicit: tighter integration in exchange for moving into someone's environment. That's a real cost if your team's tooling is already settled, and a real benefit if it isn't.&lt;/p&gt;

&lt;h3&gt;
  
  
  Spec-as-source, Tessl and friends
&lt;/h3&gt;

&lt;p&gt;The radical position: the spec is the &lt;em&gt;primary&lt;/em&gt; artifact and the code is a build output. Edit the spec, regenerate. The analogy people reach for is Terraform or a SQL query planner, you write intent, the system produces the plan.&lt;/p&gt;

&lt;p&gt;I find this directionally correct and practically premature, and I'll come back to why.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Loop-based autonomy, the Ralph technique
&lt;/h2&gt;

&lt;p&gt;Truth lives on disk: a task list, per-task specs, logs, and git history.&lt;/p&gt;

&lt;p&gt;Geoffrey Huntley named this one in mid-2025 after Ralph Wiggum, on the theory that any single iteration is a plain agent run that might get something slightly wrong, and the loop wins through persistence and a stable source of truth rather than one brilliant prompt. The minimal form is almost insultingly simple, a bash &lt;code&gt;while&lt;/code&gt; loop that re-invokes a non-interactive agent until a todo list is exhausted.&lt;/p&gt;

&lt;p&gt;The insight that makes it work is the one people miss: &lt;strong&gt;each iteration gets a fresh context.&lt;/strong&gt; A single long session accumulates fatigue, the window fills, compaction drops details, the model loses the thread. A series of short sessions that re-orient from files on disk stays sharp. Progress doesn't live in the chat; it lives in committed code, a task file, and a log the next iteration reads on boot.&lt;/p&gt;

&lt;p&gt;There are now many implementations, TUIs for visibility, a Vercel Labs wrapper that loops until a &lt;code&gt;verifyCompletion&lt;/code&gt; function passes, and enterprise write-ups running Ralph across dependency-ordered backlogs inside test-gated SDLCs.&lt;/p&gt;

&lt;p&gt;The catch, and it's a big one: canonical Ralph runs in &lt;code&gt;--yolo&lt;/code&gt; mode. Full permissions, no confirmations. That mandates real sandboxing, and it means the technique in its pure form is unusable on anything where a bad iteration can move money or drop a table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;But note what Ralph actually is.&lt;/strong&gt; It isn't a replacement for structured development, it's an &lt;em&gt;execution engine&lt;/em&gt; for a structure you already have. Its own documentation says as much. That distinction turns out to matter a lot.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. TDD governance
&lt;/h2&gt;

&lt;p&gt;Truth lives in the test suite. Tests are written first and act as executable specifications, so correctness is defined before generation begins.&lt;/p&gt;

&lt;p&gt;This has picked up serious academic attention, &lt;a href="https://arxiv.org/abs/2604.26615" rel="noopener noreferrer"&gt;work presented at EASE 2026&lt;/a&gt; on TDD governance for multi-agent code generation, and systems like TDFlow building agentic workflows around test-driven cycles. The motivating argument is one every distributed-systems person will recognize: LLMs are non-deterministic, identical prompts produce different outputs, and in multi-agent settings a small logic error propagates across the whole workflow. Automated guardrails aren't a nicety.&lt;/p&gt;

&lt;p&gt;The research also finds that test cases reduce ambiguity by functioning as executable specs, which is the same claim SDD makes about prose specs, arrived at from a different direction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where it's weak:&lt;/strong&gt; tests specify behavior beautifully and architecture terribly. A test suite cannot tell you that a module shouldn't know about another module, or that this queue is at-least-once, or &lt;em&gt;why&lt;/em&gt; a constraint exists. It's a verification layer, not an intent layer.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Multi-agent orchestration
&lt;/h2&gt;

&lt;p&gt;Truth lives in the orchestrator's task graph and shared memory.&lt;/p&gt;

&lt;p&gt;Anthropic's 2026 agentic coding trends work describes the shape: an orchestrator coordinating specialized agents in parallel, each with dedicated context, results synthesized into integrated output, as against single-agent workflows that process sequentially through one window.&lt;/p&gt;

&lt;p&gt;Addy Osmani's &lt;a href="https://addyosmani.com/blog/code-agent-orchestra/" rel="noopener noreferrer"&gt;tiering&lt;/a&gt; is the most useful practical framing I've found:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Conductor tier&lt;/strong&gt;, subagents inside one session. No extra tooling. Your context window is the ceiling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local orchestration tier&lt;/strong&gt;, multiple agents in isolated git worktrees, with dashboards and merge control. Best at roughly 3–10 agents on a codebase you know.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud async tier&lt;/strong&gt;, assign, close laptop, return to a pull request.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The honest counterweight: production-scale orchestration is a multi-quarter engineering effort, observability tooling doesn't work on non-deterministic flows out of the box, and &lt;a href="https://www.thoughtworks.com/radar" rel="noopener noreferrer"&gt;Thoughtworks' radar&lt;/a&gt; has been flagging the accumulation of &lt;em&gt;cognitive debt&lt;/em&gt; as AI-assisted development scales, the organizational version of the review bottleneck I wrote about in part one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My take:&lt;/strong&gt; orchestration is an execution strategy, not a methodology. It answers "how many at once," not "toward what."&lt;/p&gt;




&lt;h2&gt;
  
  
  The axes that actually matter
&lt;/h2&gt;

&lt;p&gt;Strip the branding and I think you're choosing along five axes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Axis&lt;/th&gt;
&lt;th&gt;The question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Durability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Does intent survive a context compaction?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Re-derivability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Can you regenerate the work from the artifact, or only document it?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Correction cost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;When you're wrong, do you fix one thing or N things?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Human position&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Where does judgment enter, before, during, or after generation?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Overhead&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;What does one cycle cost in tokens, wall-clock, and ceremony?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Re-derivability is the one I'd underlined three times if this were paper. Most teams evaluating SDD are asking whether the spec is good &lt;em&gt;documentation&lt;/em&gt;. That's the wrong question, and it's the question that gets you a spec nobody maintains.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why I landed on spec-driven plus loop engineering
&lt;/h2&gt;

&lt;p&gt;Here's my actual reasoning, and it's less about picking winners than about noticing that two of these categories are answering different questions and therefore compose.&lt;/p&gt;

&lt;h3&gt;
  
  
  The spec is an anchor, and anchors let you re-derive
&lt;/h3&gt;

&lt;p&gt;The property I care about most is &lt;strong&gt;durability that enables regeneration.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When I have a well-specified PRD, I'm not holding a description of what got built. I'm holding the thing the work is &lt;em&gt;generated from&lt;/em&gt;. That gives me three capabilities I refuse to give up:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One, I can iterate on the same requirement many times.&lt;/strong&gt; Not "revise the code," revise the &lt;em&gt;requirement&lt;/em&gt;, and let everything downstream follow. The unit of change moves up a level.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two, I can fan out.&lt;/strong&gt; One artifact spawns many tasks. The PRD is the source; the issues are derived; the implementations are derived from those. That's a tree with a single root I control, not a pile of chat sessions I have to reconcile.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three, and this is the one that sold me, I can throw implementations away.&lt;/strong&gt; When an agent goes off the rails on slice four because it invented an assumption nobody wrote down, I don't debug the agent's reasoning or nurse the code back to health. I go to the document, fix the assumption, and regenerate. Completely. As many times as I need to.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The mental shift: &lt;strong&gt;implementations became disposable and the specification became the asset.&lt;/strong&gt; That inverts the relationship I was trained on, where code was the thing you protected and documentation was the thing that rotted.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Notice this is the spec-as-source thesis, but I take it partway, deliberately. I don't accept "the spec compiles to code and you never read the output." In payments I have to be able to read, own, and be on call for what shipped. What I accept is the weaker, more useful claim: the spec is what work is &lt;em&gt;re-derived&lt;/em&gt; from. Code stays reviewable and owned. It just stops being precious.&lt;/p&gt;

&lt;h3&gt;
  
  
  The loop is how the anchor gets corrected
&lt;/h3&gt;

&lt;p&gt;Here's what spec-driven development does not do: guarantee that the spec is right.&lt;/p&gt;

&lt;p&gt;A specification is a hypothesis. It's a good hypothesis, it's been through clarifying questions and human review, but it is written before contact with the actual system, and mine are regularly wrong in ways I could not have predicted. Spec-driven development on its own, with no feedback mechanism, is waterfall with better tooling. Same failure mode, faster.&lt;/p&gt;

&lt;p&gt;That's what loop engineering is for, and I want to be precise about what I take from it. Not &lt;code&gt;--yolo&lt;/code&gt; autonomy. I took the structural insights:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fresh context per iteration&lt;/strong&gt; beats one long degrading session&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;State on disk&lt;/strong&gt;, task files, logs, git history, is the memory layer, not the chat&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An explicit stop condition&lt;/strong&gt; so the loop ends on a signal rather than a guess&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verification delegated into the loop&lt;/strong&gt; so the agent finds its own errors before I do&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Applied on top of a spec, the loop stops being brute force and becomes something much more useful: a mechanism that &lt;em&gt;invalidates spec assumptions and pushes corrections back upstream&lt;/em&gt;. Validation fails on slice two, the assumption behind it was wrong, the PRD gets edited, and slices three through seven re-derive before an agent ever touches them.&lt;/p&gt;

&lt;p&gt;That's the cascade I described in part one. It only works because both halves are present.&lt;/p&gt;

&lt;h3&gt;
  
  
  Stated as control theory, because that's what it is
&lt;/h3&gt;

&lt;p&gt;For anyone who builds distributed systems, this framing will land immediately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;        spec  ────────▶  SETPOINT     (what "correct" means)
                              │
                              ▼
       agent  ────────▶  ACTUATOR     (does the thing)
                              │
                              ▼
        loop  ────────▶  FEEDBACK     (measures the gap, corrects)
                              │
                              └──────▶ error signal updates the SETPOINT
                                       when the setpoint was the problem
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Spec without loop&lt;/strong&gt; is open-loop control. Dead reckoning. Fine until reality diverges from your model, and you find out late.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Loop without spec&lt;/strong&gt; is a controller with no setpoint. It will converge on &lt;em&gt;something&lt;/em&gt;, reliably, persistently, cheerfully, just not necessarily on what you wanted. Or it oscillates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Both&lt;/strong&gt; is closed-loop control with a defined target and an error signal that can correct the target itself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's not a novel idea. It's the oldest idea in engineering, applied to a new actuator. Which is roughly why I trust it more than the approaches that feel newer.&lt;/p&gt;

&lt;h3&gt;
  
  
  What I took from the others anyway
&lt;/h3&gt;

&lt;p&gt;Methodologies aren't teams. You don't have to pick one and wear the jersey:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;From Spec Kit&lt;/strong&gt;, the &lt;em&gt;constitution&lt;/em&gt; idea. Durable project-level rules every spec inherits, so I'm not restating non-negotiables in every PRD.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;From OpenSpec&lt;/strong&gt;, delta thinking. Most of my work is amendment, not creation, and specs that describe changes beat specs that restate the world.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;From TDD governance&lt;/strong&gt;, tests as the loop's ground truth. Property-based and contract tests are what make the agent's self-correction meaningful rather than self-congratulatory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;From orchestration&lt;/strong&gt;, parallelism bounded by the dependency graph, and worktree isolation. As an execution strategy underneath the loop, not as the loop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;From the security research&lt;/strong&gt;, this one's my favourite. There's &lt;a href="https://arxiv.org/abs/2602.02584" rel="noopener noreferrer"&gt;work on "constitutional" SDD&lt;/a&gt; that embeds non-negotiable security constraints, derived from CWE/MITRE Top 25 and regulatory frameworks, into the specification layer, so generated code satisfies them &lt;em&gt;by construction&lt;/em&gt; rather than by inspection. If you read part one, you know why that lands for me. It's the compliance bottleneck, addressed at the only place it can be addressed cheaply: upstream, in the artifact.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  And what I rejected, plainly
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;BMAD&lt;/strong&gt;, the handoff surface. In a domain where a silently propagated wrong assumption means a double-charge, I want fewer inter-agent boundaries, not more. Also the token bill.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Canonical Ralph&lt;/strong&gt;, I took the pattern and left the &lt;code&gt;--yolo&lt;/code&gt;. Human gates on money-moving paths are non-negotiable, which is incompatible with the pure form.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kiro and the platform tier&lt;/strong&gt;, I'm not moving environments for a brownfield payments codebase. That's a fine trade for other people; it isn't for me.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TDD governance as a &lt;em&gt;whole&lt;/em&gt; methodology&lt;/strong&gt;, rejected, because tests can't carry architectural intent. But hold this one loosely; I come back to it at the end, and my position on it has changed.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The strongest objection to my position
&lt;/h2&gt;

&lt;p&gt;I'd rather state this than have it stated at me.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/cameronsjo/spec-compare" rel="noopener noreferrer"&gt;most rigorous comparison&lt;/a&gt; I've found of these tools flags two things that should give any SDD advocate pause. First: it's genuinely unclear when SDD adds value versus overhead, for trivial changes, the heavier frameworks are a sledgehammer cracking a nut. Second, and sharper: the historical parallel to Model-Driven Development, which made structurally similar promises in the 2000s and did not survive contact with real software.&lt;/p&gt;

&lt;p&gt;I think the counterargument is that MDD failed partly because the generation step was rigid and the abstraction leaked badly, you got code you couldn't touch from models you couldn't express real systems in. LLM-generated implementations are readable, editable, and idiomatic, and specs are prose. That's a materially different failure surface.&lt;/p&gt;

&lt;p&gt;But I hold it loosely. The honest version is that the Stack Overflow numbers keep telling us adoption isn't the problem, most developers use these tools and far fewer trust their output. Every methodology on this list is a bet about how to close that gap. Mine is a bet too.&lt;/p&gt;




&lt;h2&gt;
  
  
  There is no correct answer, and that isn't a cop-out
&lt;/h2&gt;

&lt;p&gt;I want to be careful here, because "it depends" is what people say when they haven't thought about something. I have thought about this, and it still depends, for two reasons that are worth separating.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The first is that your constraints genuinely differ from mine.&lt;/strong&gt; Blast radius, brownfield versus greenfield, team size, regulatory posture, how much of your work is amendment. Those are real inputs and they point at different answers.&lt;/p&gt;

&lt;p&gt;But there's a softer input that I think gets dismissed too easily: &lt;strong&gt;your implementation preferences are a real engineering constraint, not a personality quirk.&lt;/strong&gt; A methodology you abandon in week three has negative value, you paid the setup cost and got none of the compounding. If a seven-persona pipeline makes you want to close the laptop, that's data. Ceremony you won't sustain is worse than ceremony you never adopted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The second reason is that the ground moves, and it moves fast.&lt;/strong&gt; Not just the tooling, the &lt;em&gt;models&lt;/em&gt;. They differ from each other in ways that matter to methodology design: instruction-following, how gracefully they degrade over long context, whether they ask or assume when a spec is ambiguous, what their characteristic failure modes even are.&lt;/p&gt;

&lt;p&gt;Which leads to an uncomfortable observation about this entire field guide:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A large fraction of what we call methodology is really model-specific error correction with a good name on it. Some of these practices exist to compensate for a particular failure mode of a particular generation of model. When the model changes, the compensation may become unnecessary, or insufficient.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So treat everything above, mine included, as a snapshot. Retune when you switch models. Notice which of your rituals are load-bearing and which are scar tissue.&lt;/p&gt;




&lt;h2&gt;
  
  
  But the shape of the near-optimum is visible
&lt;/h2&gt;

&lt;p&gt;Here's where I'll be less hedging. I don't think there's a perfect methodology. I do think the &lt;em&gt;shape&lt;/em&gt; of a very good one is now clear, and it isn't any single entry on this list.&lt;/p&gt;

&lt;p&gt;It's three layers doing three genuinely distinct jobs, with no overlap between them:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Methodology&lt;/th&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Intent&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Spec-driven development&lt;/td&gt;
&lt;td&gt;Defines what correct means, and why. Durable, re-derivable.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Verification&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;TDD governance&lt;/td&gt;
&lt;td&gt;Determines whether it's correct, mechanically, in a way no agent can talk past.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Correction&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Loop engineering&lt;/td&gt;
&lt;td&gt;Closes the gap, and pushes errors back up to Intent when the intent was the problem.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Earlier in this post I drew the control loop with two components. That was incomplete, and deliberately so, because I wanted to arrive here. The full picture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;        spec  ────────▶  SETPOINT      what "correct" means
                              │
                              ▼
       agent  ────────▶  ACTUATOR      does the thing
                              │
                              ▼
       tests  ────────▶  SENSOR        measures reality, honestly
                              │
                              ▼
        loop  ────────▶  CONTROLLER    computes the error, acts on it
                              │
                              └───────▶ and when the error is in the SETPOINT
                                        itself, corrects upstream
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Adding the sensor is not a garnish, and this is the part I got wrong for a while. &lt;strong&gt;A feedback loop with an untrustworthy sensor is worse than no loop at all&lt;/strong&gt;, because it converges, confidently, persistently, at machine speed, on the wrong thing. That is precisely the failure mode of an agent writing its own tests from the same misunderstanding that produced its code. It measures, it agrees with itself, it proceeds.&lt;/p&gt;

&lt;p&gt;That's why I've moved TDD governance from "component I borrowed" to co-equal layer. Spec-driven development gives the loop something to converge &lt;em&gt;toward&lt;/em&gt;. Test governance gives it something to converge &lt;em&gt;by&lt;/em&gt;. Remove either and the third stops working.&lt;/p&gt;




&lt;h2&gt;
  
  
  Composing rather than choosing
&lt;/h2&gt;

&lt;p&gt;If the three layers are the skeleton, the individual frameworks become a parts bin rather than a set of competing religions:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Take&lt;/th&gt;
&lt;th&gt;From&lt;/th&gt;
&lt;th&gt;Because&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Project constitution&lt;/td&gt;
&lt;td&gt;Spec Kit&lt;/td&gt;
&lt;td&gt;Non-negotiables shouldn't be restated per feature&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delta specs&lt;/td&gt;
&lt;td&gt;OpenSpec&lt;/td&gt;
&lt;td&gt;Most real work is amendment, not creation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fresh context per iteration, state on disk&lt;/td&gt;
&lt;td&gt;Ralph&lt;/td&gt;
&lt;td&gt;Long sessions degrade; files don't&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Test-first gating, adversarially derived&lt;/td&gt;
&lt;td&gt;TDD governance&lt;/td&gt;
&lt;td&gt;The sensor has to be independent of the actuator&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Worktree parallelism&lt;/td&gt;
&lt;td&gt;Orchestration tooling&lt;/td&gt;
&lt;td&gt;Execution strategy under the loop, not the loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Constraints in the spec layer&lt;/td&gt;
&lt;td&gt;Constitutional SDD&lt;/td&gt;
&lt;td&gt;Compliance is cheapest upstream&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two warnings before you go shopping.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every layer is an artifact you now have to keep true.&lt;/strong&gt; Composition has a maintenance cost that compounds, and the failure mode of a hybrid is a spec that no longer describes the system, gating agents that trust it. Adopt a layer only if you'll maintain it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And don't confuse composition with accumulation.&lt;/strong&gt; Bolting six frameworks together gets you the union of their ceremony and the intersection of their benefits. The point is to take one thing from each that does a job nothing else does, and to be able to say what that job is.&lt;/p&gt;

&lt;p&gt;So, four questions, and they're about composing rather than picking:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Amendment or creation?&lt;/strong&gt; → delta specs versus phase-structured ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What's your cost of being wrong?&lt;/strong&gt; → determines how much sensor you need and where the human gates go.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-derive or just document?&lt;/strong&gt; → if you'll never regenerate from the artifact, you're maintaining documentation, and it will rot the way documentation always has.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What ceremony budget will you actually sustain?&lt;/strong&gt; → answered honestly, not aspirationally.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What I'm building next
&lt;/h2&gt;

&lt;p&gt;The obvious move, once you see the three layers, is to stop treating them as three methodologies you're running concurrently and start treating them as one strategy with a single control flow. Intent, verification, and correction as first-class stages in one system rather than three practices you're manually keeping in sync.&lt;/p&gt;

&lt;p&gt;That's what I've been working on. It's a hybrid, spec-driven at the anchor, test-governed at the gate, loop-driven at the correction step, and the interesting problems turn out to be at the seams: how the sensor gets derived from the spec without inheriting its blind spots, and what exactly happens when the loop concludes the setpoint was wrong.&lt;/p&gt;

&lt;p&gt;That's the next post.&lt;/p&gt;

&lt;p&gt;In the meantime: if you're running a combination I haven't covered, especially if you've made a persona-pipeline approach work at scale, because I would like to be wrong about that, tell me what broke and what held.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part two of a series on agentic engineering in production. Part one covered the three-skill workflow and where the bottleneck goes when code stops being the constraint. Part three is the hybrid.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>programming</category>
      <category>ai</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Agentic Engineering Is Not Vibe Coding: The Three-Skill Loop I Use to Ship Distributed Systems</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Sun, 02 Aug 2026 17:59:55 +0000</pubDate>
      <link>https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge</link>
      <guid>https://dev.to/mtsammy40/agentic-engineering-is-not-vibe-coding-the-three-skill-loop-i-use-to-ship-distributed-systems-2bge</guid>
      <description>&lt;p&gt;There is a category error running loose in our industry right now, and it is costing teams real money.&lt;/p&gt;

&lt;p&gt;The error is treating &lt;em&gt;vibe coding&lt;/em&gt; and &lt;em&gt;agentic engineering&lt;/em&gt; as the same activity performed at different levels of enthusiasm. They are not the same activity. They have different units of work, different failure modes, different artifacts, and, this is the part that matters, different economics when you point them at a production payment system that moves other people's money.&lt;/p&gt;

&lt;p&gt;I build distributed systems for a living. Lately most of that has been payment infrastructure, which is the least forgiving place I know of to be wrong. Over the last while I have converged on a workflow that is roughly two-thirds spec-driven development and one-third loop engineering, implemented as three Claude Code skills that hand off to each other.&lt;/p&gt;

&lt;p&gt;This post is that workflow, plus the things it broke that nobody warned me about.&lt;/p&gt;




&lt;h2&gt;
  
  
  The distinction, stated sharply
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Vibe coding&lt;/th&gt;
&lt;th&gt;Agentic engineering&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Optimizes for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Time-to-first-output&lt;/td&gt;
&lt;td&gt;Time-to-verified-increment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Unit of work&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A file, a function, "the thing I asked for"&lt;/td&gt;
&lt;td&gt;An independently testable vertical slice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Where intent lives&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;In a prompt that scrolls away&lt;/td&gt;
&lt;td&gt;In a durable, reviewed artifact&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Verification&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Visual, manual, by you&lt;/td&gt;
&lt;td&gt;Automated, delegated, adversarial&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Context&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Hoped for&lt;/td&gt;
&lt;td&gt;Engineered from a source of truth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Characteristic failure&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Shipping something you can't explain&lt;/td&gt;
&lt;td&gt;Caught by the loop, not by you at 2am&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;Vibe coding is a legitimate mode. I use it constantly for spikes, throwaway scripts, and exploring an unfamiliar API. It is a terrible mode for anything with an idempotency requirement.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The rest of this post is about the second column.&lt;/p&gt;




&lt;h2&gt;
  
  
  Spec-driven development plus loop engineering
&lt;/h2&gt;

&lt;p&gt;Two ideas are doing the heavy lifting here, so let me define them the way I actually use them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Spec-driven development&lt;/strong&gt; means intent lives in an artifact that outlives any single agent session. Not in a prompt. Not in your head. Not in a Slack thread. In a document an agent can re-read at the top of every run, that a human reviewed and approved, and that changes through a visible edit rather than a vibe shift.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Loop engineering&lt;/strong&gt; means deliberately designing the feedback loop the agent operates inside. What can it execute? What tells it that it is wrong, quickly and unambiguously? Where must it stop and escalate to a human?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A model without a tight loop is a very expensive autocomplete. A model with a tight loop is a colleague who never gets bored of running the test suite.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Neither is a prompting trick. Both are systems design applied one level up from the code.&lt;/p&gt;




&lt;h2&gt;
  
  
  The workflow: three skills, one handoff chain
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  rough prompt
       │
       ▼
┌──────────────┐   clarifying questions ──▶ human
│   /to-prd     │◀──────────────────────────  ┘
└──────┬───────┘
       │  PRD: user stories · outcomes · invariants · OUT OF SCOPE
       ▼        (posted to GitHub Issues, reviewed before any code exists)
┌──────────────┐
│  /to-issues   │
└──────┬───────┘
       │  vertical slices, each independently testable, each linking the PRD
       ▼
┌──────────────┐        ┌─────────────────┐
│ engineering  │───────▶│ parallel slices │
│    skill     │        │  1   2   3   4  │
└──────┬───────┘        └────────┬────────┘
       │                         │
       │      ┌──────────────────┘
       │      ▼
       │  validation fails ──▶ assumption was wrong
       │      │
       │      └──▶ PRD updated ──▶ UNDONE slices re-derived  ← the outer loop
       ▼
  human gate (blast radius, not difficulty)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  1. &lt;code&gt;/to-prd&lt;/code&gt;, turn a vague intent into a specification
&lt;/h3&gt;

&lt;p&gt;Every feature and every non-trivial bug starts here. I give it a rough prompt. Sometimes embarrassingly rough, a sentence and a link to a Sentry issue.&lt;/p&gt;

&lt;p&gt;The skill then does four things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Extrapolates the prompt into candidate requirements&lt;/li&gt;
&lt;li&gt;Reads the actual codebase to ground those requirements in what exists&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Asks me clarifying questions&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Emits a full PRD, user stories, outcomes, explicit out-of-scope, and posts it to GitHub Issues&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Step three is the entire value proposition, and it took me a while to see that.&lt;/p&gt;

&lt;p&gt;The clarifying questions convert ambiguity into a decision at minute three instead of a defect at week six. When I write &lt;em&gt;"add retry logic to the settlement webhook,"&lt;/em&gt; a vibe-coded session gives me retry logic. &lt;code&gt;/to-prd&lt;/code&gt; comes back and asks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Idempotent on the provider's key, or ours?&lt;/li&gt;
&lt;li&gt;What happens to ordering when attempt two lands before attempt one?&lt;/li&gt;
&lt;li&gt;Does a poisoned message go to a DLQ, or block the partition?&lt;/li&gt;
&lt;li&gt;Is partial settlement a valid terminal state?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I did not know the answer to two of those. &lt;strong&gt;That is the point.&lt;/strong&gt; The gap between "what I asked for" and "what I meant" is where production incidents are manufactured, and a specification pass is a machine for finding that gap early.&lt;/p&gt;

&lt;h4&gt;
  
  
  The out-of-scope section is load-bearing
&lt;/h4&gt;

&lt;p&gt;I want to be emphatic about this one. Agents are enthusiastic. Given a settlement webhook to fix, an unconstrained agent will refactor your retry utility, introduce a circuit breaker, and rename three things on the way past.&lt;/p&gt;

&lt;p&gt;An explicit out-of-scope list is the cheapest guardrail in the entire workflow. It is the difference between a 200-line diff and a 2,000-line diff that nobody can review.&lt;/p&gt;

&lt;p&gt;Posting to GitHub Issues rather than a local file is deliberate: it puts the spec where the team already lives, gives it a URL, makes it commentable, and makes review of the &lt;strong&gt;specification&lt;/strong&gt; a first-class event, before a single line of implementation exists.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. &lt;code&gt;/to-issues&lt;/code&gt;, decompose into vertical slices
&lt;/h3&gt;

&lt;p&gt;This skill takes the approved PRD and breaks it into implementation issues. The decomposition rule is the important part: &lt;strong&gt;vertical, never horizontal.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;HORIZONTAL (org-chart slicing)        VERTICAL (what I actually use)

┌───────────────────────────┐         ┌────┐┌────┐┌────┐┌────┐
│  frontend ticket          │         │ UI ││ UI ││ UI ││ UI │
├───────────────────────────┤         ├────┤├────┤├────┤├────┤
│  backend ticket           │         │ API││ API││ API││ API│
├───────────────────────────┤         ├────┤├────┤├────┤├────┤
│  migration ticket         │         │ DB ││ DB ││ DB ││ DB │
└───────────────────────────┘         └────┘└────┘└────┘└────┘
 nothing testable until all             each slice testable
 three land simultaneously              on its own, on merge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Horizontal slicing is an artifact of the org chart, not of engineering. Its defining property is that &lt;em&gt;nothing is testable until everything lands&lt;/em&gt;. You accumulate work in progress for two weeks and discover on integration day that the API contract you both agreed to meant two different things.&lt;/p&gt;

&lt;p&gt;Vertical slicing means every slice cuts through the whole stack: thin, but complete. Each push can be validated on its own. The next push iterates on top of it.&lt;/p&gt;

&lt;p&gt;This matters more with agents than it ever did with humans, for a specific reason:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The validation signal is the cheapest resource you have, and slicing determines how often you get one. If a slice is independently testable, the agent self-corrects and the loop closes without you. If it isn't, &lt;strong&gt;you&lt;/strong&gt; are the test suite, and you've just reintroduced yourself as the bottleneck in the exact place you were trying to remove yourself from.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The second benefit is the one I underrated at the start. When validating slice two reveals that a PRD assumption was wrong, that corrected context &lt;strong&gt;cascades downstream to slices three through seven before an agent ever touches them.&lt;/strong&gt; You are not fixing seven implementations. You are fixing one document and re-deriving the remaining work from it.&lt;/p&gt;

&lt;p&gt;A wrong assumption caught in slice two costs an edit. The same assumption caught at integration costs a rewrite.&lt;/p&gt;

&lt;p&gt;That cascade is the loop engineering part. The inner loop is &lt;em&gt;agent writes code, tests fail, agent fixes code.&lt;/em&gt; The outer loop is &lt;em&gt;validation invalidates a spec assumption, spec updates, undone work re-derives.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The engineering skill, execute, in parallel, with gates
&lt;/h3&gt;

&lt;p&gt;The third skill picks up the generated issues, each referencing the parent PRD, and executes them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Parallelism where the dependency graph allows it.&lt;/strong&gt; Independent slices run concurrently. This is where the wall-clock gains actually come from, not from the model typing faster than you, but from four slices progressing at once while you're in a meeting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Human-in-the-loop markers.&lt;/strong&gt; Some issues are flagged as requiring a human. The skill prepares, stops, presents, and waits. Everything else runs AFK.&lt;/p&gt;

&lt;p&gt;My gating rule is blast radius, not difficulty:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Runs unsupervised&lt;/th&gt;
&lt;th&gt;Requires me in the room&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gnarly reconciliation algorithm&lt;/td&gt;
&lt;td&gt;Schema migration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Complex parsing / transformation&lt;/td&gt;
&lt;td&gt;Any money-moving code path&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Test scaffolding, fixtures&lt;/td&gt;
&lt;td&gt;Authorization boundary changes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Refactors covered by property tests&lt;/td&gt;
&lt;td&gt;Anything whose failure is worse than a revert&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;An agent can write a hard reconciliation algorithm unsupervised, because a wrong one fails a property test. Difficulty is a bad proxy. &lt;strong&gt;Cost-of-being-wrong is the right one.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Why payments make all of this non-optional
&lt;/h2&gt;

&lt;p&gt;You can run a sloppy version of this on a CRUD app and mostly get away with it. Distributed systems that move money will not extend you that courtesy.&lt;/p&gt;

&lt;p&gt;Here is the failure mode that made me take specification seriously:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Looks fine. Passes review. Passes its tests.&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;settle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;payment&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Payment&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="nx"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;charge&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;payment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;currency&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;payment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;currency&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That code double-charges a customer under a network partition, because the retry has no idempotency key and the underlying delivery guarantee was at-least-once all along. A timeout is &lt;em&gt;ambiguous&lt;/em&gt;, the charge may well have committed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// The invariant, made explicit in code because it was explicit in the spec.&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;settle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;payment&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Payment&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;idempotencyKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;payment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;payment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;attemptEpoch&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="nx"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;charge&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;payment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;currency&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;payment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;currency&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;idempotencyKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// stable across retries, this is the whole ballgame&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;Nothing about the first version looks wrong on review. It is wrong at the level of &lt;strong&gt;invariants&lt;/strong&gt;, and invariants are invisible in a diff.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Which means the specification has to carry more than behavior. My PRDs for anything on a payment path encode non-functional requirements explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;invariants&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;idempotency&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;provider_ref + attempt_epoch&lt;/span&gt;
    &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ours&lt;/span&gt;
    &lt;span class="na"&gt;ttl&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;24h&lt;/span&gt;
    &lt;span class="na"&gt;on_collision&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;return prior result, do not re-execute&lt;/span&gt;
  &lt;span class="na"&gt;ordering&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;transport_guarantee&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;at-least-once, unordered&lt;/span&gt;
    &lt;span class="na"&gt;code_may_assume&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nothing&lt;/span&gt;
  &lt;span class="na"&gt;failure_modes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;timeout&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ambiguous, reconcile, never blind-retry&lt;/span&gt;
    &lt;span class="na"&gt;poison_message&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;DLQ after 3, alert, never block partition&lt;/span&gt;
    &lt;span class="na"&gt;partial_settlement&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;valid terminal state, requires ledger entry&lt;/span&gt;
  &lt;span class="na"&gt;reconciliation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;cadence&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;every 15m against provider ledger&lt;/span&gt;
    &lt;span class="na"&gt;drift_threshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;0, any delta pages&lt;/span&gt;
  &lt;span class="na"&gt;observability&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;trace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;payment_id propagated end to end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And verification has to be &lt;strong&gt;adversarial rather than confirmatory&lt;/strong&gt;. Example-based tests written from the same understanding that produced the code will faithfully confirm that understanding, including the parts that are wrong. Property-based tests, fault injection, and contract tests are ground truth an agent cannot talk its way past. They are what make self-correction meaningful instead of self-congratulatory.&lt;/p&gt;




&lt;h2&gt;
  
  
  What actually got better
&lt;/h2&gt;

&lt;p&gt;Being specific, because "10x productivity" is a claim I don't believe from anyone:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ambiguity surfaces before implementation, not after.&lt;/strong&gt; The clarifying-questions pass is the highest-leverage part of the workflow and costs about four minutes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wall-clock time on multi-slice features dropped substantially&lt;/strong&gt;, mostly parallelism, plus not context-switching myself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The spec exists.&lt;/strong&gt; For the first time in my career, the design doc isn't a thing we wrote after shipping to satisfy an auditor. It's load-bearing, so it gets maintained.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Onboarding is different.&lt;/strong&gt; A new engineer reads the PRD chain and understands &lt;em&gt;why&lt;/em&gt;, not just what.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;My worst days got better.&lt;/strong&gt; Tedious, well-understood, mechanically tiresome work now happens without my attention.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Now the honest part: I moved the bottleneck, I did not remove it
&lt;/h2&gt;

&lt;p&gt;This is the section I most want people to sit with, because the industry conversation is stuck on generation speed, and generation speed stopped being the interesting variable a while ago.&lt;/p&gt;

&lt;h3&gt;
  
  
  Code review does not scale with code generation
&lt;/h3&gt;

&lt;p&gt;Generation is now nearly free. Comprehension is exactly as expensive as it was in 2019.&lt;/p&gt;

&lt;p&gt;A senior engineer can meaningfully review some finite amount of code per day. The number is smaller than any of us admit, and it collapses further when the reviewer didn't write the surrounding context. Point an agentic workflow at that reviewer and you don't get a linear slowdown, you get &lt;strong&gt;rubber-stamping&lt;/strong&gt;, which is strictly worse than no review, because it manufactures false confidence and diffuses accountability across a team that now collectively believes someone looked.&lt;/p&gt;

&lt;p&gt;What helps, in order of effectiveness for me:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Vertical slices keep the review surface human-sized.&lt;/strong&gt; A 200-line slice with a clear spec is reviewable. A 2,000-line feature drop is theater.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review against the spec, not the diff.&lt;/strong&gt; The question is &lt;em&gt;"does this satisfy the stated invariants,"&lt;/em&gt; not &lt;em&gt;"would I have typed this."&lt;/em&gt; The second question is unanswerable at volume and mostly ego anyway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A reviewing agent with a different context than the implementing agent.&lt;/strong&gt; Same context reproduces the same blind spot. Different context, derived from the PRD rather than the implementation, catches things.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hard human gates where being wrong is expensive.&lt;/strong&gt; Non-negotiable, no exceptions for velocity.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  The answer is agentic validation, not more human eyes
&lt;/h3&gt;

&lt;p&gt;If generation is agentic and verification is manual, you have built a funnel that ends at a person.&lt;/p&gt;

&lt;p&gt;The only structurally sound response is to make verification agentic too: test generation derived independently from the spec, fault injection, contract conformance, invariant checking, spec-conformance review as a distinct automated pass.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The humans move up the stack. You review the specification and the invariants. Agents review conformance to them. That's a real change in what the job is, and I don't think it's a demotion.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  The backlog moved downstream, and that is the actual story
&lt;/h3&gt;

&lt;p&gt;Here's the thing nobody put in the launch demo.&lt;/p&gt;

&lt;p&gt;Theory of constraints has been telling us this for forty years: &lt;strong&gt;you do not eliminate a bottleneck, you relocate it.&lt;/strong&gt; Code was the constraint. Now it isn't. So go look at where it went.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BEFORE                          AFTER

 product   ▓▓                    product   ▓▓
 eng       ▓▓▓▓▓▓▓▓▓▓  ← here    eng       ▓
 QA        ▓▓▓                   QA        ▓▓▓▓▓▓
 security  ▓▓                    security  ▓▓▓▓▓
 compliance▓▓▓                   compliance▓▓▓▓▓▓▓▓▓▓▓  ← here now
 ops       ▓▓                    ops       ▓▓▓▓▓▓
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It went to &lt;strong&gt;QA&lt;/strong&gt;, handed features faster than it can write test plans. To &lt;strong&gt;compliance&lt;/strong&gt;, where PCI-DSS scope assessments, SOC 2 evidence collection, and data protection reviews run on a human cadence that did not just get 5x faster. To &lt;strong&gt;security review&lt;/strong&gt;. To &lt;strong&gt;SRE and ops&lt;/strong&gt;, absorbing a deployment frequency sized for a slower team. To &lt;strong&gt;change advisory boards&lt;/strong&gt;, release trains, and every governance process that assumed engineering was the slow part.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If you make engineering agentic and change nothing else, you have built a machine that generates work-in-progress inventory for teams downstream of you. WIP is not throughput. In some organizations it is a liability with a carrying cost.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The response is not to slow engineering down. It's to recognize that agentic practice is an organizational transformation wearing a developer-tooling costume:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Compliance evidence generated continuously as a &lt;strong&gt;build artifact&lt;/strong&gt;, not assembled in a panic before an audit&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Policy-as-code&lt;/strong&gt;, with control mapping evaluated in CI&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent-drafted change records&lt;/strong&gt; and risk assessments, human-approved&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compliance and QA present at the PRD stage&lt;/strong&gt;, requirements encoded up front rather than assessed as a veto at the end&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last bullet has the best return and the most organizational friction. Shifting compliance left is far easier when the spec is already a machine-readable artifact your agents produced. It's one of the genuinely underrated second-order benefits of spec-driven work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And a few costs I have not solved&lt;/strong&gt; (click to expand)&lt;/p&gt;



&lt;p&gt;&lt;strong&gt;Where does judgment come from now?&lt;/strong&gt; If agents do the mechanical work, junior engineers don't get the reps that used to produce senior engineers. I don't have an answer. I don't think the industry does either.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Spec drift.&lt;/strong&gt; A PRD that no longer describes the system is worse than no PRD, because agents trust it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost.&lt;/strong&gt; Four parallel agents burn tokens and CI minutes. Often worth it. Never free.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ownership.&lt;/strong&gt; Who is on call for code nobody wrote? The answer has to be a person, and the workflow has to make that person feel it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Nondeterminism.&lt;/strong&gt; Same prompt, different output. Your &lt;em&gt;process&lt;/em&gt; must be reproducible even when generation isn't, which is precisely why the spec, the slices, and the gates are the durable parts, and the generation is the fungible part.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I actually want you to take away
&lt;/h2&gt;

&lt;p&gt;The shift is not that a model can write your code. It's that &lt;strong&gt;your job moved from authoring the system to designing the system that authors the system.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The skills you're compensated for are increasingly: specification, decomposition, invariant definition, verification design, and knowing where a human must stand in the loop. That's architecture work. It was always the valuable part. It's now the &lt;em&gt;only&lt;/em&gt; part that doesn't commoditize.&lt;/p&gt;

&lt;p&gt;So, two things to go do this week:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Find your most repetitive class of work&lt;/strong&gt;, the feature shape you've built eleven times. Don't write a code-generation skill for it. Write the &lt;strong&gt;specification&lt;/strong&gt; skill first. Make it ask you clarifying questions. You'll be unsettled by how many you can't answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Go measure where your constraint actually is right now.&lt;/strong&gt; Not where it was when you last thought about it. If your engineers ship in two days and your compliance review takes three weeks, then every hour you spend making generation faster is an hour spent making the queue longer.&lt;/p&gt;

&lt;p&gt;Make it agentic end to end, or don't be surprised when the bottleneck simply moves in next door.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I write about distributed systems, payment infrastructure, and agentic engineering practice. If you're running a variation of this workflow, especially in a regulated environment, I want to hear what broke. That's the most useful thing anyone can share right now.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>ai</category>
      <category>programming</category>
      <category>automation</category>
    </item>
    <item>
      <title>What Payments Infrastructure Taught Me About Building Systems That Don't Break</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Fri, 31 Jul 2026 09:45:31 +0000</pubDate>
      <link>https://dev.to/mtsammy40/what-payments-infrastructure-taught-me-about-building-systems-that-dont-break-kjo</link>
      <guid>https://dev.to/mtsammy40/what-payments-infrastructure-taught-me-about-building-systems-that-dont-break-kjo</guid>
      <description>&lt;p&gt;&lt;strong&gt;Idempotency, vendor failure, monitoring that catches the invisible outages, and the tradeoffs nobody warns you about, lessons from scaling payments infrastructure.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most software fails quietly. A page renders slowly, a recommendation is a little off, a report is stale by an hour. Users shrug and move on.&lt;/p&gt;

&lt;p&gt;Payments doesn't work like that. When payments break, someone's money is in a place neither of you can account for, and the clock starts ticking on their patience. There's no graceful degradation. &lt;strong&gt;Either the money moved, or it didn't, and someone needs to know which.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I've spent a good chunk of my career building and scaling payments infrastructure, and it has quietly rewired how I think about engineering in general. Here's what stuck.&lt;/p&gt;




&lt;h2&gt;
  
  
  📋 The short version
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Lesson&lt;/th&gt;
&lt;th&gt;One-line summary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Idempotency&lt;/td&gt;
&lt;td&gt;You &lt;em&gt;will&lt;/em&gt; receive the same request twice. Design for it.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Vendor failure&lt;/td&gt;
&lt;td&gt;Gateways are vendors. Ask "when," not "if."&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Monitoring&lt;/td&gt;
&lt;td&gt;Never learn about an outage from a customer.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;The unglamorous stuff&lt;/td&gt;
&lt;td&gt;Ledgers, reconciliation, state machines, refunds.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Tradeoffs&lt;/td&gt;
&lt;td&gt;Every lesson above fights at least one other.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  1. 🔁 Idempotency isn't a feature. It's a foundation.
&lt;/h2&gt;

&lt;p&gt;The first hard lesson: you will receive the same request twice. Not "might." &lt;strong&gt;Will.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A client times out waiting for your response and retries. A user double-taps a button on a bad connection. A queue consumer crashes after processing but before acknowledging. A gateway sends the same webhook four times because it never got a &lt;code&gt;200&lt;/code&gt; back. None of these are exotic failure modes, they're Tuesday.&lt;/p&gt;

&lt;p&gt;If your system treats every incoming call as a new instruction, every one of those scenarios becomes a double charge. And a double charge isn't a bug you fix quietly in the next release. It's a support ticket, a refund, a reconciliation entry, and a customer who now checks their statement every time they use you.&lt;/p&gt;

&lt;p&gt;The fix is conceptually simple and operationally demanding: every operation that moves money must be uniquely identifiable and safely repeatable. The caller supplies an idempotency key. You store it &lt;em&gt;before&lt;/em&gt; you do anything else. If you see it again, you return the original result, not a new attempt, not an error, &lt;strong&gt;the same answer you gave the first time&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The naive version everyone writes first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# ❌ Race condition hiding in plain sight
&lt;/span&gt;&lt;span class="n"&gt;existing&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find_by_idempotency_key&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;existing&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;existing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;charge_customer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="c1"&gt;# two concurrent requests
&lt;/span&gt;&lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;save&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                  &lt;span class="c1"&gt;# both get here
&lt;/span&gt;&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two identical requests arriving 5ms apart both pass the &lt;code&gt;if&lt;/code&gt;. Let the database enforce it instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# ✅ The unique constraint does the real work
&lt;/span&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transaction&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="n"&gt;record&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;insert_idempotency_record&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;in_progress&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;UniqueConstraintViolation&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;await_or_return_existing&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;charge_customer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;record&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;complete&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The parts that actually take work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Idempotency has to be atomic with the operation itself.&lt;/strong&gt; If you record the key in one transaction and charge in another, you've just moved the race condition somewhere less obvious.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Concurrent duplicates are the hard case.&lt;/strong&gt; A naive "check then write" won't catch them. You need a unique constraint or a lock doing the real work, not application logic hoping for the best.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keys need a sensible lifetime.&lt;/strong&gt; Too short and legitimate retries slip through. Too long and you're storing a permanent record of every request you've ever seen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It has to extend past your API boundary.&lt;/strong&gt; Your internal queues, your workers, your webhook handlers, every hop is another opportunity to duplicate. Idempotency at the edge with at-least-once processing behind it just moves the problem inward.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;Idempotency is really just the discipline of making your system's behavior depend on &lt;strong&gt;intent&lt;/strong&gt; rather than on &lt;strong&gt;delivery&lt;/strong&gt;, and delivery is the thing you don't control.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  2. ⚠️ Every gateway you integrate is a vendor, and vendors fail
&lt;/h2&gt;

&lt;p&gt;When you integrate a payment gateway, you're not adding a library. You're taking on a dependency whose code you can't read, whose deploys you don't schedule, whose incidents you learn about from Twitter, and whose SLA is a document, not a guarantee.&lt;/p&gt;

&lt;p&gt;I stopped asking "what if this provider goes down" a long time ago. The right question is &lt;strong&gt;"what do we do when this provider goes down,"&lt;/strong&gt; and it's a question with a schedule attached, because it will happen this quarter.&lt;/p&gt;

&lt;p&gt;That shift in framing changes the architecture:&lt;/p&gt;

&lt;h3&gt;
  
  
  🧱 You isolate them
&lt;/h3&gt;

&lt;p&gt;Every provider sits behind your own interface, speaking your own domain language. Your core system should never know that Provider A calls it a &lt;code&gt;transaction_reference&lt;/code&gt; and Provider B calls it a &lt;code&gt;paymentId&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;         ┌──────────────────────┐
         │   Payment Service    │   ← speaks YOUR domain language
         └──────────┬───────────┘
                    │
         ┌──────────▼───────────┐
         │  Provider Interface  │   ← translation lives here
         └──┬────────┬────────┬─┘
            │        │        │
        ┌───▼──┐ ┌───▼──┐ ┌───▼──┐
        │  A   │ │  B   │ │  C   │   ← vendors. they will fail.
        └──────┘ └──────┘ └──────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That translation layer is annoying to build, and it is the thing that lets you swap a provider &lt;em&gt;under load&lt;/em&gt; instead of during a two-month migration.&lt;/p&gt;

&lt;h3&gt;
  
  
  ⏱️ You assume timeouts, not errors
&lt;/h3&gt;

&lt;p&gt;A clean error is a gift, it tells you what happened. The genuinely dangerous response is &lt;em&gt;no response&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Your request may have been processed. It may not have been. You don't know, and you can't just retry blindly. This is where reconciliation and status polling stop being nice-to-haves: when the network gives you ambiguity, you need a way to go ask &lt;strong&gt;"did this actually happen?"&lt;/strong&gt; and get an authoritative answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  🔌 You add circuit breakers
&lt;/h3&gt;

&lt;p&gt;When a provider starts failing, hammering it with retries makes their recovery slower and your queues longer. Fail fast, back off, route elsewhere, come back later.&lt;/p&gt;

&lt;h3&gt;
  
  
  🎲 You retry with jitter, and only where it's safe
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;delay&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;max_delay&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;delay&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uniform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;delay&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# ← the line people skip
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Exponential backoff &lt;em&gt;without&lt;/em&gt; jitter just means all your retries collide again in a synchronized wave. And "safe to retry" is only true because you did the work in lesson one.&lt;/p&gt;

&lt;h3&gt;
  
  
  🔀 You route around damage
&lt;/h3&gt;

&lt;p&gt;If you're on multiple providers, the point isn't just pricing leverage, it's that you can shift volume when one degrades.&lt;/p&gt;

&lt;p&gt;But automatic failover is only trustworthy if you can tell the difference between &lt;strong&gt;"provider is down"&lt;/strong&gt; and &lt;strong&gt;"provider is correctly declining these transactions."&lt;/strong&gt; Failing over on legitimate declines is a great way to turn a small problem into a fraud incident.&lt;/p&gt;

&lt;h3&gt;
  
  
  🧪 You never trust the sandbox
&lt;/h3&gt;

&lt;p&gt;Test environments are clean, fast, and always available. Production is none of those things. The behaviors that hurt you, partial failures, delayed settlement, out-of-order webhooks, undocumented status codes, almost never show up until real traffic does.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. 📡 You should never learn about an outage from a customer
&lt;/h2&gt;

&lt;p&gt;If a customer is telling you something is broken, you've already lost twice: once for the failure, and once for not knowing about it first.&lt;/p&gt;

&lt;p&gt;Monitoring in payments has to work on two levels, and most teams only build one.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;What you measure&lt;/th&gt;
&lt;th&gt;What it tells you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Technical&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Latency, error rates, queue depth, saturation, provider response times, retry counts&lt;/td&gt;
&lt;td&gt;The system is unhealthy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Business&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Success rate by provider / channel / card type / country, volume vs. last week, transactions stuck pending, settlement vs. ledger&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Money is not moving correctly&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The reason you need both is that the worst payments incidents don't look like outages.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every service is up. Latency is fine. Error rates are flat. And a specific bank's cards have been silently failing for forty minutes because a provider quietly changed something.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Infrastructure metrics will never catch that. A success-rate alert segmented by issuer catches it in five minutes.&lt;/p&gt;

&lt;p&gt;A few things I now consider non-negotiable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Alert on trends, not just thresholds.&lt;/strong&gt; A drop from 94% to 71% success is an emergency even though 71% isn't zero.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every alert needs a runbook.&lt;/strong&gt; An alert that fires at 3am with no attached "here's what to check and who to call" is just anxiety with a pager.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Traceability end to end.&lt;/strong&gt; When someone asks about one specific transaction, you should be able to reconstruct its entire life, every state change, every provider call, every retry, without SSH-ing into anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ops needs their own tools.&lt;/strong&gt; Your support and operations teams shouldn't be filing engineering tickets to answer routine customer questions. Give them read access to transaction state and safe, audited actions. This buys back an astonishing amount of engineering time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Availability is the other half of this. In most products, downtime costs you some engagement. In payments, downtime is transactions that didn't happen, revenue that evaporates and doesn't come back, plus something more expensive: &lt;strong&gt;trust that erodes a little each time and doesn't rebuild at the same rate.&lt;/strong&gt; People will forgive a slow app. They get quietly nervous about a payment system that failed on them once, and they stop reaching for it first.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. 🧾 The things nobody warns you about
&lt;/h2&gt;

&lt;p&gt;A few more lessons that cost me something to learn.&lt;/p&gt;

&lt;h3&gt;
  
  
  Your ledger is the source of truth, not your provider's dashboard
&lt;/h3&gt;

&lt;p&gt;Build a proper double-entry, append-only ledger early. Never update a balance in place. Never delete an entry. Corrections are new entries.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- ❌ Where did the money go? Nobody knows.&lt;/span&gt;
&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;accounts&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;balance&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;balance&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- ✅ Append-only. The history IS the balance.&lt;/span&gt;
&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;ledger_entries&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;account_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;amount_minor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;currency&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;direction&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;VALUES&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;50000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'KES'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'debit'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="s1"&gt;'txn_9f2a'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
       &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;17&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;50000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'KES'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'credit'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'txn_9f2a'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When something eventually goes wrong, and it will, the only thing that saves you is an immutable record of what your system believed at every point in time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reconcile continuously, not monthly
&lt;/h3&gt;

&lt;p&gt;Automated comparison between your ledger and each provider's settlement reports, running daily at minimum. Discrepancies compound. A mismatch found the next morning is a fix; the same mismatch found at quarter-end is an investigation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Never use floating point for money
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="mf"&gt;0.1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;0.2&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mf"&gt;0.3&lt;/span&gt;   &lt;span class="c1"&gt;// false. this is your revenue.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Store minor units as integers. Always attach a currency. Decide your rounding rules explicitly and write them down, because rounding disagreements between you and a provider will eventually surface as a real, unexplainable gap.&lt;/p&gt;

&lt;h3&gt;
  
  
  Model payment state as an explicit state machine
&lt;/h3&gt;

&lt;p&gt;Not a boolean. Not a status string that anyone can set to anything.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;initiated ──► pending ──┬─► succeeded ──► refunded
                        │
                        ├─► failed
                        │
                        └─► expired

# succeeded ──► pending is NOT a valid transition.
# Enforce that in ONE place, not in seven services.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Half the weird bugs in payments are illegal state transitions that nobody thought to prevent.&lt;/p&gt;

&lt;h3&gt;
  
  
  Webhooks arrive out of order, late, or twice
&lt;/h3&gt;

&lt;p&gt;Sometimes the "success" event lands before the "processing" event. Handle each one as a &lt;strong&gt;fact about a point in time&lt;/strong&gt;, not as an instruction to advance a status.&lt;/p&gt;

&lt;h3&gt;
  
  
  Design the unhappy paths first
&lt;/h3&gt;

&lt;p&gt;Refunds, partial refunds, reversals, chargebacks, disputes, expired authorizations. Teams build the happy path, ship it, and then discover that refunds don't fit the data model at all. The unhappy paths are where the actual complexity lives, and retrofitting them is far more expensive than designing for them.&lt;/p&gt;

&lt;h3&gt;
  
  
  Compliance is an architectural constraint, not a checklist
&lt;/h3&gt;

&lt;p&gt;What you're allowed to store, where you're allowed to store it, who can see it, how long you keep it. Discovering these requirements after you've built is a rewrite.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. ⚖️ Everything above is in tension with everything else
&lt;/h2&gt;

&lt;p&gt;Here's the part that took me longest to accept: none of these lessons is free, and several of them actively fight each other.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you gain&lt;/th&gt;
&lt;th&gt;What it costs you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Aggressive retries → higher success rates&lt;/td&gt;
&lt;td&gt;More duplicate risk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-provider redundancy → resilience&lt;/td&gt;
&lt;td&gt;Integration surface, reconciliation complexity, on-call load&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Comprehensive monitoring → catch real problems&lt;/td&gt;
&lt;td&gt;Alert fatigue that slows your real response&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Meticulous ledger → certainty&lt;/td&gt;
&lt;td&gt;Write throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;There's no configuration where you max out every dial. The engineering is in deciding, deliberately and with your eyes open, where you sit on each of these, and being honest that it's a choice with a cost, rather than pretending you got everything.&lt;/p&gt;




&lt;h2&gt;
  
  
  🎯 If you're starting on this today
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Build for the failure cases before you build for scale.&lt;/strong&gt; Scale problems announce themselves loudly and you get to fix them in daylight. Correctness problems hide, compound quietly, and surface at the worst possible time, usually as a number that doesn't add up and a customer who wants to know where their money went.&lt;/p&gt;

&lt;p&gt;Get the boring parts right:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;[ ] Idempotency keys, enforced at the database level&lt;/li&gt;
&lt;li&gt;[ ] An append-only, double-entry ledger&lt;/li&gt;
&lt;li&gt;[ ] Monitoring on business metrics, not just infrastructure&lt;/li&gt;
&lt;li&gt;[ ] An explicit state machine for payment status&lt;/li&gt;
&lt;li&gt;[ ] Automated daily reconciliation&lt;/li&gt;
&lt;li&gt;[ ] Refunds and chargebacks in the data model from day one&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else is much easier to build on top of that than it is to bolt on afterward.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Have you hit any of these the hard way? I'd like to hear which one bit you, the comments are open.&lt;/em&gt; 👇&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>devops</category>
      <category>performance</category>
    </item>
    <item>
      <title>Microservices or Micro-progress? The Perils of Pre-Scaling Too Soon</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Fri, 15 Aug 2025 21:28:09 +0000</pubDate>
      <link>https://dev.to/mtsammy40/microservices-or-micro-progress-the-perils-of-pre-scaling-too-soon-1049</link>
      <guid>https://dev.to/mtsammy40/microservices-or-micro-progress-the-perils-of-pre-scaling-too-soon-1049</guid>
      <description>&lt;h3&gt;
  
  
  &lt;strong&gt;The Grand Illusion of "Future-Proof" Code&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Ah, the siren song of scalability. It whispers sweet nothings into our ears:&amp;nbsp;&lt;em&gt;"Build it right the first time," "You’ll thank yourself later," "What if you get 10 million users tomorrow?"&lt;/em&gt;&amp;nbsp;And before you know it, you’ve spent three months architecting a dazzling microservices masterpiece, only to realize your "scalable" app has exactly one user:&amp;nbsp;&lt;strong&gt;you&lt;/strong&gt;, refreshing the page in incognito mode.&lt;/p&gt;

&lt;p&gt;This was me. I was&amp;nbsp;&lt;em&gt;that&lt;/em&gt;&amp;nbsp;developer. Convinced that my side project needed a Kubernetes cluster, a distributed notification system, and a payment service that could handle Stripe-level traffic, despite the fact that my MVP was still just a glorified to-do list.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Hard Truth: Scale Doesn’t Matter If Nobody Cares&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Here’s the uncomfortable reality:&amp;nbsp;&lt;strong&gt;You cannot optimize for problems you don’t have yet.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Are you drowning in user traffic?&lt;/strong&gt;&amp;nbsp;→ No? Then why are you building a CDN?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do you have 10,000 concurrent payment requests?&lt;/strong&gt;&amp;nbsp;→ No? Then why are you overengineering Stripe?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Are your notifications crashing under load?&lt;/strong&gt;&amp;nbsp;→ No? Then why did you build a pub-sub system when Novu’s free tier exists?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I had to face the music: I wasn’t building for scale, I was&amp;nbsp;&lt;strong&gt;procrasti-scaling&lt;/strong&gt;. Avoiding the real work (making something people actually wanted) by obsessing over infrastructure that&amp;nbsp;&lt;em&gt;might&lt;/em&gt;&amp;nbsp;matter someday.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Pivot: From Over-Engineered to "Good Enough"&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;So I did the unthinkable: I&amp;nbsp;&lt;strong&gt;deleted months of work&lt;/strong&gt;&amp;nbsp;and replaced it with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Supabase&lt;/strong&gt;&amp;nbsp;(Auth + DB) → Free tier. Works. Done.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Novu&lt;/strong&gt;&amp;nbsp;(Notifications) → Free tier. Works. Done.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stripe&lt;/strong&gt;&amp;nbsp;(Payments) → Basic integration. Works. Done.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Suddenly, I was making&amp;nbsp;&lt;strong&gt;real progress -&lt;/strong&gt; not on infrastructure, but on&amp;nbsp;&lt;strong&gt;the actual product&lt;/strong&gt;. And guess what? If (big&amp;nbsp;&lt;em&gt;if&lt;/em&gt;) my app ever outgrows these tools, I’ll&amp;nbsp;&lt;strong&gt;cross that bridge when I get there&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Lesson: Build Stupid First&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Your first version should be&amp;nbsp;&lt;strong&gt;embarrassingly simple&lt;/strong&gt;. Not because you’re lazy, but because:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;You don’t yet know what needs scaling.&lt;/strong&gt;&amp;nbsp;(Most things won’t.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Premature optimization is the root of all evil.&lt;/strong&gt;&amp;nbsp;(Or at least wasted time.)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Your biggest risk isn’t scale, it’s building something nobody wants.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So next time you catch yourself designing a Kafka pipeline for your cat blog, ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Is this solving a real problem, or just my fear of hypothetical ones?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Because in the end,&amp;nbsp;&lt;strong&gt;a working app with limits beats an unfinished "scalable" one every time.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Now go build something stupid. You can always make it smart later. 🚀&lt;/p&gt;

</description>
      <category>programming</category>
      <category>distributedsystems</category>
      <category>mvp</category>
    </item>
    <item>
      <title>Setting Up Keycloak for Passwordless Authentication</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Wed, 28 May 2025 09:57:59 +0000</pubDate>
      <link>https://dev.to/mtsammy40/setting-up-keycloak-for-passwordless-authentication-2fg1</link>
      <guid>https://dev.to/mtsammy40/setting-up-keycloak-for-passwordless-authentication-2fg1</guid>
      <description>&lt;p&gt;Passwordless authentication is becoming a must-have for modern applications, no more forgotten passwords, just seamless access via magic links, biometrics, or security keys. &lt;strong&gt;&lt;a href="https://www.keycloak.org/" rel="noopener noreferrer"&gt;Keycloak&lt;/a&gt;&lt;/strong&gt;, the popular open-source identity and access management solution, makes implementing passwordless auth surprisingly straightforward.  &lt;/p&gt;

&lt;p&gt;In this guide, we’ll walk through configuring Keycloak to support &lt;strong&gt;email-based magic links&lt;/strong&gt; (a common passwordless approach). Let’s dive in!  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Prerequisites&lt;/strong&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A running Keycloak instance (v20+)
&lt;/li&gt;
&lt;li&gt;SMTP server access (for sending magic links)
&lt;/li&gt;
&lt;li&gt;Basic familiarity with Keycloak admin console
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Step 1: Enable Email Verification&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Since passwordless auth relies on email links, we first need to ensure Keycloak can send emails.  &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Configure SMTP settings&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Go to &lt;strong&gt;Realm Settings → Email&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Fill in your SMTP server details (e.g., Gmail, SendGrid, Postmark)
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;   Host: smtp.example.com  
   Port: 587  
   From: no-reply@yourdomain.com  
   Enable SSL/TLS: Yes  
   Authentication: Enabled (provide credentials)  
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Test email delivery&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Click &lt;strong&gt;Test connection&lt;/strong&gt; to verify everything works.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Step 2: Set Up Passwordless Authentication Flow&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Keycloak uses &lt;strong&gt;authentication flows&lt;/strong&gt; to define login steps. We’ll customize the default flow.  &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Create a new authentication flow&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Navigate to &lt;strong&gt;Authentication → Flows&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Click &lt;strong&gt;New flow&lt;/strong&gt;, name it (e.g., "Passwordless Email")
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Add required steps&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Under your new flow, add these &lt;strong&gt;executions&lt;/strong&gt;:

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Username Form&lt;/strong&gt; (for email input)
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Send Email Verification Link&lt;/strong&gt; (replaces password check)
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conditional User Role&lt;/strong&gt; (optional, for additional security)
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Disable password requirement&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Go to &lt;strong&gt;Realm Settings → Login&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Disable &lt;strong&gt;"Password"&lt;/strong&gt; as a required credential
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Step 3: Customize the Magic Link Email&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Keycloak sends a verification email, let’s make it user-friendly.  &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Edit the email template&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Go to &lt;strong&gt;Realm Settings → Email → Templates&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Modify &lt;strong&gt;"Verify Email"&lt;/strong&gt; to include a clear call-to-action:
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;   &lt;span class="nt"&gt;&amp;lt;p&amp;gt;&lt;/span&gt;Click below to log in:&lt;span class="nt"&gt;&amp;lt;/p&amp;gt;&lt;/span&gt;  
   &lt;span class="nt"&gt;&amp;lt;a&lt;/span&gt; &lt;span class="na"&gt;href=&lt;/span&gt;&lt;span class="s"&gt;"${url}"&lt;/span&gt; &lt;span class="na"&gt;style=&lt;/span&gt;&lt;span class="s"&gt;"background: #2563eb; color: white; padding: 10px 20px; text-decoration: none; border-radius: 5px;"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;Sign In Instantly&lt;span class="nt"&gt;&amp;lt;/a&amp;gt;&lt;/span&gt;  
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Set link expiration&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Under &lt;strong&gt;Realm Settings → Tokens&lt;/strong&gt;, adjust &lt;strong&gt;"Email Verification Link Lifespan"&lt;/strong&gt; (e.g., 15 minutes).
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Step 4: Test the Flow&lt;/strong&gt;
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Try logging in as a test user.
&lt;/li&gt;
&lt;li&gt;Instead of a password field, you’ll see an email input.
&lt;/li&gt;
&lt;li&gt;After submitting, check your inbox for the magic link.
&lt;/li&gt;
&lt;li&gt;Clicking it should log you in directly!
&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Bonus: Adding WebAuthn (Biometric Auth)&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;For a more advanced passwordless experience, enable &lt;strong&gt;WebAuthn&lt;/strong&gt; (for security keys/biometrics):  &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Go to &lt;strong&gt;Authentication → Flows&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Add &lt;strong&gt;"WebAuthn Authenticator"&lt;/strong&gt; as an alternative.
&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Final Thoughts&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Keycloak makes passwordless auth surprisingly simple. With just a few tweaks, you can replace clunky passwords with secure, user-friendly magic links or biometric logins.  &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Need help?&lt;/strong&gt; Check out the &lt;a href="https://www.keycloak.org/docs/latest/server_admin/" rel="noopener noreferrer"&gt;official Keycloak docs&lt;/a&gt; or drop a question below!  &lt;/p&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
      <category>authentication</category>
      <category>keycloak</category>
    </item>
    <item>
      <title>AI in Corporate: The Coming Wave of Disastrous Blunders</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Wed, 14 May 2025 18:50:38 +0000</pubDate>
      <link>https://dev.to/mtsammy40/ai-in-corporate-the-coming-wave-of-disastrous-blunders-367h</link>
      <guid>https://dev.to/mtsammy40/ai-in-corporate-the-coming-wave-of-disastrous-blunders-367h</guid>
      <description>&lt;h3&gt;
  
  
  &lt;strong&gt;How AI-Powered Tools Will Accidentally Wreck Prod (And Maybe Some Companies)&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;We’re in the early days of AI integration in corporate environments, and I’m willing to bet &lt;strong&gt;we’re about to see an avalanche of catastrophic mistakes&lt;/strong&gt;.  &lt;/p&gt;

&lt;p&gt;Picture this:  &lt;/p&gt;

&lt;p&gt;A well-meaning infrastructure engineer, eager to speed up their workflow, starts using an AI-powered IDE like &lt;strong&gt;Cursor&lt;/strong&gt; with the &lt;em&gt;"auto-perform commands"&lt;/em&gt; feature enabled. They explain a problem—maybe not perfectly, maybe missing some key context—and the AI, confident as ever, starts running terminal commands &lt;strong&gt;the engineer has never even seen before&lt;/strong&gt;.  &lt;/p&gt;

&lt;p&gt;Before they know it:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Database tables are dropped.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Configs are overwritten.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prod is a smoking crater.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the worst part? &lt;strong&gt;No one fully understands what happened&lt;/strong&gt; because the AI made decisions based on incomplete or misinterpreted context.  &lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Why This Will Happen Over and Over&lt;/strong&gt;
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Over-Trust in AI’s “Understanding”&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI doesn’t &lt;em&gt;reason&lt;/em&gt;—it predicts. If your prompt is ambiguous, it will still generate &lt;em&gt;something&lt;/em&gt;, and that something might be &lt;code&gt;rm -rf&lt;/code&gt; in the wrong directory.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;The Illusion of Control&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tools that auto-execute commands (like GitHub Copilot with shell gen, Cursor’s AI agent, etc.) &lt;strong&gt;remove the human review step&lt;/strong&gt;. Engineers might assume the AI "gets it" until it very clearly doesn’t.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Silent Failures&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Unlike a human who might say &lt;em&gt;"Wait, this looks dangerous,"&lt;/em&gt; AI will happily run destructive commands &lt;strong&gt;with confidence&lt;/strong&gt;. By the time logs are checked, it’s too late.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Fallout: AI Will Kill Some Companies&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;We’ve already seen AI blunders:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Legal briefs citing fake cases&lt;/strong&gt; (because the AI hallucinated precedents).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Customer service bots going rogue&lt;/strong&gt; (see: Air Canada’s chatbot inventing refund policies).
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now imagine:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A fintech AI misinterprets a deployment script and wipes transaction records.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A cloud AI "optimizes" costs by deleting "unused" resources… like prod databases.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Some companies &lt;strong&gt;won’t recover&lt;/strong&gt; from these mistakes—especially if they happen at scale.  &lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;How to Survive the AI Wild West&lt;/strong&gt;
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Treat AI Like a Junior Dev Who Lies Sometimes&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Always review generated code/commands before execution.
&lt;/li&gt;
&lt;li&gt;Sandbox &lt;strong&gt;everything&lt;/strong&gt; before it touches prod.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Disable Auto-Execute (For Now)&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tools that let AI run commands directly are &lt;strong&gt;dangerous&lt;/strong&gt;. Keep it in "suggestion mode" until guardrails improve.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Assume It Will Fail&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Log every AI-generated action.
&lt;/li&gt;
&lt;li&gt;Build rollback plans for when (not if) AI breaks something.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Bottom Line&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;AI is powerful, but &lt;strong&gt;we’re in the "move fast and break things" phase&lt;/strong&gt;—except now, the "things" might be entire companies.  &lt;/p&gt;

&lt;p&gt;Brace for the chaos. &lt;strong&gt;The first wave of AI-induced disasters is coming.&lt;/strong&gt;  &lt;/p&gt;




&lt;p&gt;&lt;strong&gt;What’s the worst AI blunder you’ve seen (or caused)?&lt;/strong&gt; Drop your horror stories below. 👇&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>data</category>
      <category>webdev</category>
    </item>
    <item>
      <title>AI Won’t Replace Developers Anytime Soon—Here’s Why</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Wed, 07 May 2025 18:15:53 +0000</pubDate>
      <link>https://dev.to/mtsammy40/ai-wont-replace-developers-anytime-soon-heres-why-2166</link>
      <guid>https://dev.to/mtsammy40/ai-wont-replace-developers-anytime-soon-heres-why-2166</guid>
      <description>&lt;p&gt;(Critiques of the distant future, please consider the date of publication 😅) &lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Promise (and Pitfalls) of AI-Generated Code&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;AI coding assistants like &lt;a href="https://cursor.sh" rel="noopener noreferrer"&gt;Cursor&lt;/a&gt; and &lt;a href="https://claude.ai" rel="noopener noreferrer"&gt;Claude&lt;/a&gt; have been game-changers for developer productivity. They can scaffold entire applications, debug tricky issues, and even explain complex concepts in seconds. But can they &lt;em&gt;fully&lt;/em&gt; replace software engineers?  &lt;/p&gt;

&lt;p&gt;Not yet—and here’s a funny (and frustrating) example of why.  &lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Case of the Broken Countdown Timer&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;I recently asked Cursor to generate a Next.js landing page with a countdown timer to my product launch. It did &lt;strong&gt;most&lt;/strong&gt; of the work well—the UI looked great, the logic seemed sound, but when I tested it… the timer was &lt;strong&gt;stuck&lt;/strong&gt;.  &lt;/p&gt;

&lt;p&gt;I alerted the AI, but instead of fixing the issue, it just:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Repeated the same code
&lt;/li&gt;
&lt;li&gt;Gave me a generic checklist (e.g., "Check if the date is correct")
&lt;/li&gt;
&lt;li&gt;Missed the glaring problem
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Next, I pasted the code into Claude. It &lt;strong&gt;thought&lt;/strong&gt; it was a hydration issue (a reasonable guess in Next.js) and tweaked the code—but the timer &lt;strong&gt;still didn’t work&lt;/strong&gt;.  &lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Obvious Bug AI Missed&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;After some manual debugging (thankfully, I know JavaScript), I spotted the issue:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;getTimeLeft&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;launchDate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// 🚨 Problem: This resets EVERY render!&lt;/span&gt;
  &lt;span class="nx"&gt;launchDate&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setDate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;launchDate&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getDate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;35&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// 5 weeks from today&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;now&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;diff&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;launchDate&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getTime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;now&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getTime&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="c1"&gt;// ... rest of the logic&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The bug? &lt;strong&gt;&lt;code&gt;launchDate&lt;/code&gt; was being recalculated on every render&lt;/strong&gt;, meaning the countdown always showed &lt;code&gt;0&lt;/code&gt; (since &lt;code&gt;now&lt;/code&gt; and &lt;code&gt;launchDate&lt;/code&gt; were effectively the same time).  &lt;/p&gt;

&lt;p&gt;The fix? **Make &lt;code&gt;launchDate&lt;/code&gt; a static field (From what future date are we counting down?):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;LAUNCH_DATE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;2025-06-01&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;getTimeLeft&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;now&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;diff&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;LAUNCH_DATE&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getTime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;now&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getTime&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="c1"&gt;// ... rest of the logic&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  &lt;strong&gt;Why AI Still Needs Human Oversight&lt;/strong&gt;
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;AI Lacks Deep Context&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It didn’t realize &lt;code&gt;launchDate&lt;/code&gt; should be static.
&lt;/li&gt;
&lt;li&gt;It followed patterns but didn’t &lt;em&gt;understand&lt;/em&gt; the intent.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Debugging Requires Reasoning, Not Just Repetition&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Both assistants gave &lt;em&gt;plausible&lt;/em&gt; suggestions but didn’t &lt;em&gt;diagnose&lt;/em&gt; the root cause.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Trivial Mistakes Are Hard for AI to Spot&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Humans recognize "obvious" errors faster because we think in terms of goals, not just syntax.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Verdict: AI is a Powerful Assistant, Not a Replacement&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;AI has made &lt;strong&gt;incredible progress&lt;/strong&gt;, but it still:&lt;br&gt;&lt;br&gt;
✔ Struggles with nuanced logic&lt;br&gt;&lt;br&gt;
✔ Misses simple but critical bugs&lt;br&gt;&lt;br&gt;
✔ Needs human guidance for real-world scenarios  &lt;/p&gt;

&lt;p&gt;So, developers, rest easy, your job is safe (for now). AI is a &lt;strong&gt;tool&lt;/strong&gt;, not a replacement. And honestly? That’s a good thing.  &lt;/p&gt;

</description>
      <category>vibecoding</category>
      <category>programming</category>
      <category>ai</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Mistakes Are Part of the Dev Environment: It's How You Handle Them That Counts</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Wed, 07 May 2025 08:26:35 +0000</pubDate>
      <link>https://dev.to/mtsammy40/mistakes-are-part-of-the-dev-environment-its-how-you-handle-them-that-counts-2m1p</link>
      <guid>https://dev.to/mtsammy40/mistakes-are-part-of-the-dev-environment-its-how-you-handle-them-that-counts-2m1p</guid>
      <description>&lt;p&gt;“To err is human, to debug divine.”&lt;/p&gt;

&lt;p&gt;If you’ve been writing code for any stretch of time, you’ve probably stared at your screen wondering how something so small could break so much. Whether it’s a misplaced semicolon, an off-by-one error, or deploying to production with test credentials (guilty 😅), every developer—junior or senior—has stories of mistakes they’ve made.&lt;/p&gt;

&lt;p&gt;Let’s face it: mistakes are not just part of the development process—they are the development process.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Illusion of Perfection&lt;/strong&gt;&lt;br&gt;
There's a common myth, especially among beginners, that good developers don't make mistakes. That once you’re “senior,” you write perfect code the first time.&lt;/p&gt;

&lt;p&gt;Spoiler: Nobody does.&lt;/p&gt;

&lt;p&gt;In fact, senior developers often make bigger mistakes—but they catch and fix them faster because they’ve learned how to deal with them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why Mistakes Matter&lt;/strong&gt;&lt;br&gt;
Mistakes aren’t just inevitable—they’re useful. They reveal assumptions, force you to dig deeper into how systems work, and often lead to better solutions than you would have considered otherwise. Here’s what mistakes teach us:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Clarity: They expose where our understanding is shallow.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Resilience: They teach us how to handle failure calmly.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Humility: They remind us that there’s always more to learn.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;A Culture of Blame vs. a Culture of Learning&lt;/strong&gt;&lt;br&gt;
One of the worst things a dev team can do is foster a culture where mistakes are punished. This leads to fear, cover-ups, and stagnation.&lt;/p&gt;

&lt;p&gt;Great teams treat mistakes as data points. When something breaks, the question isn't “Who messed up?”—it’s “How can we make this less likely to happen again?”&lt;/p&gt;

&lt;p&gt;That might mean:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Improving test coverage.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Setting up guardrails (like staging environments or automated linting).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Writing clearer documentation.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Doing post-mortems without finger-pointing.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Your Mistake Recovery Toolbox&lt;/strong&gt;&lt;br&gt;
Here are a few practical things every developer should keep in their mental toolkit when mistakes happen:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Stay Calm&lt;br&gt;
Panicking never helps. Take a breath, grab a coffee, and remember: you are not your bug.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Reproduce It&lt;br&gt;
If you can reproduce the bug, you can fix the bug. Step-by-step isolation is often the fastest way forward.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Use Version Control&lt;br&gt;
If you’re not using Git (or similar), start now. Version control is the ultimate undo button.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Write Tests Around It&lt;br&gt;
Write a test that fails because of the bug. Then fix it. This not only confirms the fix—it prevents regressions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Learn and Share&lt;br&gt;
Write about it, tweet it, blog it, bring it up in retros. Your mistake could prevent someone else’s.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Some Famous Dev Mistakes&lt;/strong&gt;&lt;br&gt;
GitHub once deleted part of their production database.&lt;/p&gt;

&lt;p&gt;AWS once took down a major part of the internet with a typo.&lt;/p&gt;

&lt;p&gt;Google once lost an entire day of email due to a bad config.&lt;/p&gt;

&lt;p&gt;These are massive companies with brilliant engineers. Mistakes still happen. What sets them apart is how they respond.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Final Thoughts&lt;/strong&gt;&lt;br&gt;
Programming is a constant process of learning, breaking, fixing, and improving. Mistakes are not bugs in you—they’re features of the process.&lt;/p&gt;

&lt;p&gt;So next time you bork the deployment or introduce a hard-to-find bug, don’t beat yourself up. Reflect, fix, learn—and maybe even laugh about it later.&lt;/p&gt;

&lt;p&gt;Because at the end of the day, real growth in software engineering happens not in the absence of mistakes, but in how we respond to them.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
      <category>beginners</category>
      <category>career</category>
    </item>
    <item>
      <title>Quantum Computing is About to Pwn Your Encryption – Time to Wake Up!</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Mon, 05 May 2025 11:37:27 +0000</pubDate>
      <link>https://dev.to/mtsammy40/quantum-computing-is-about-to-pwn-your-encryption-time-to-wake-up-4c5i</link>
      <guid>https://dev.to/mtsammy40/quantum-computing-is-about-to-pwn-your-encryption-time-to-wake-up-4c5i</guid>
      <description>&lt;h2&gt;
  
  
  &lt;strong&gt;"Harvest Now, Decrypt Later" – The World’s Slowest Heist&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Imagine a burglar breaking into your house, but instead of stealing your TV, they take a &lt;strong&gt;photocopy of your safe’s lock&lt;/strong&gt; and say:&lt;br&gt;&lt;br&gt;
&lt;em&gt;"I’ll crack this later when I invent lock-picking lasers."&lt;/em&gt;  &lt;/p&gt;

&lt;p&gt;That’s essentially what hackers are doing right now with &lt;strong&gt;"Harvest Now, Decrypt Later" (HNDL) attacks&lt;/strong&gt;. They’re hoarding encrypted data (your emails, bank details, even those embarrassing selfies) and waiting for quantum computers to crack them open like a cheap piñata.  &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You:&lt;/strong&gt; &lt;em&gt;"But quantum computing isn’t ready yet!"&lt;/em&gt;&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Hackers:&lt;/strong&gt; &lt;em&gt;"We can wait. Your data isn’t going anywhere."&lt;/em&gt;  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Shor’s Algorithm: The Math Bully That Eats RSA for Breakfast&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Current encryption relies on math problems so hard that even supercomputers cry trying to solve them. &lt;strong&gt;But quantum computers? They cheat.&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;RSA Encryption:&lt;/strong&gt; &lt;em&gt;"It’ll take a billion years to factor this large prime!"&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shor’s Algorithm:&lt;/strong&gt; &lt;em&gt;"Hold my qubit."&lt;/em&gt; 💥
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What’s at risk?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
✔ Your &lt;strong&gt;HTTPS connections&lt;/strong&gt; (bye-bye, secure banking).&lt;br&gt;&lt;br&gt;
✔ Your &lt;strong&gt;SSH keys&lt;/strong&gt; (hope you like unexpected server guests).&lt;br&gt;&lt;br&gt;
✔ &lt;strong&gt;Bitcoin &amp;amp; blockchain&lt;/strong&gt; (unless they upgrade fast, quantum miners will be the new crypto whales).  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Grover’s Algorithm: The Unwanted Gym Bro of Encryption&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;While Shor’s algorithm &lt;strong&gt;destroys&lt;/strong&gt; RSA &amp;amp; ECC, Grover’s algorithm is more of a &lt;strong&gt;persistent annoyance&lt;/strong&gt; to symmetric encryption like AES:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AES-256?&lt;/strong&gt; Still strong, but now with &lt;strong&gt;only AES-128-level security&lt;/strong&gt;.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SHA-256?&lt;/strong&gt; Collision attacks just got &lt;strong&gt;way easier&lt;/strong&gt;.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Translation:&lt;/strong&gt; Your encryption just lost half its gains. Time to hit the &lt;strong&gt;quantum-resistant crypto gym&lt;/strong&gt;. 💪  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Post-Quantum Cryptography: The Superhero We Need (But Don’t Deserve)&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;NIST has been working on &lt;strong&gt;quantum-proof algorithms&lt;/strong&gt;, because apparently, we can’t just &lt;strong&gt;unplug the quantum computers&lt;/strong&gt; and call it a day. Here’s the new lineup:  &lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;1. CRYSTALS-Kyber – The New RSA (But Fancier)&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Good for:&lt;/strong&gt; Key exchanges (so quantum hackers can’t eavesdrop).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bad for:&lt;/strong&gt; People who miss the good ol’ days of RSA (which, let’s be honest, were never that good).
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;2. CRYSTALS-Dilithium – Like ECDSA, But Won’t Die in 5 Years&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Replaces:&lt;/strong&gt; Digital signatures (so your GitHub commits stay legit).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bonus:&lt;/strong&gt; Sounds like a &lt;strong&gt;Power Rangers weapon&lt;/strong&gt;. &lt;em&gt;"Go go Dilithium Signatures!"&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;3. SPHINCS+ – The Backup That Nobody Wants to Use&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;How it works:&lt;/strong&gt; Hash-based, so even if quantum breaks everything else, this still stands.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Downside:&lt;/strong&gt; Bigger, slower, like that one relative who still uses a &lt;strong&gt;flip phone&lt;/strong&gt;.
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;What You Should Do Before Quantum Hackers Ruin Your Day&lt;/strong&gt;
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Stop pretending this isn’t happening.&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;"Quantum computing is decades away!"&lt;/em&gt; – People who will be hacked in 5 years.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Check if your crypto is already obsolete.&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Still using &lt;strong&gt;RSA-2048?&lt;/strong&gt; Start planning your migration &lt;strong&gt;yesterday&lt;/strong&gt;.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Demand quantum-safe encryption in your tools.&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ask your cloud provider: &lt;em&gt;"Hey, when are you adding Kyber support?"&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;If they say &lt;em&gt;"What’s Kyber?"&lt;/em&gt; – &lt;strong&gt;panic&lt;/strong&gt;.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Prepare for the inevitable "Oh crap" moment.&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Because someday, a headline will say: &lt;em&gt;"Quantum computer just broke Bitcoin"&lt;/em&gt;, and you don’t want to be scrambling then.
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Final Thought: Don’t Be the Last One Using Broken Crypto&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Quantum computing is coming, and it &lt;strong&gt;doesn’t care&lt;/strong&gt; if your security team is ready. The good news? &lt;strong&gt;We have solutions now.&lt;/strong&gt; The bad news? &lt;strong&gt;Most people won’t act until it’s too late.&lt;/strong&gt;  &lt;/p&gt;

&lt;p&gt;Will you be the &lt;strong&gt;early adopter&lt;/strong&gt; sipping coffee while others panic? Or the one &lt;strong&gt;rewriting your entire auth system at 3 AM&lt;/strong&gt; after the quantum apocalypse hits?  &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The choice is yours.&lt;/strong&gt;  &lt;/p&gt;




&lt;p&gt;&lt;em&gt;"But I don’t even understand quantum mechanics!"&lt;/em&gt; – Don’t worry, neither do most quantum physicists. Just start learning &lt;strong&gt;post-quantum crypto&lt;/strong&gt; today.* 😉  &lt;/p&gt;

</description>
    </item>
    <item>
      <title>Passwords Are a Ticking Timebomb—And These Breaches Prove It</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Mon, 05 May 2025 11:27:16 +0000</pubDate>
      <link>https://dev.to/mtsammy40/passwords-are-a-ticking-timebomb-and-these-breaches-prove-it-25h1</link>
      <guid>https://dev.to/mtsammy40/passwords-are-a-ticking-timebomb-and-these-breaches-prove-it-25h1</guid>
      <description>&lt;p&gt;Passwords have been the default authentication method for decades, but their flaws are more dangerous than ever. High-profile breaches and cyberattacks consistently expose how fragile password-based security really is. Below, we’ll examine two major case studies that highlight why passwords are failing us—and what we should use instead.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Case Study 1: The LinkedIn Breach (2012, 2016, and Beyond)&lt;/strong&gt;
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;What Happened?&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;In &lt;strong&gt;2012&lt;/strong&gt;, LinkedIn suffered a breach that exposed &lt;strong&gt;164 million email and password combinations&lt;/strong&gt;.
&lt;/li&gt;
&lt;li&gt;Hackers didn’t just steal passwords—they cracked weak hashes (SHA-1 without salting), revealing plaintext credentials.
&lt;/li&gt;
&lt;li&gt;In &lt;strong&gt;2016&lt;/strong&gt;, another batch of &lt;strong&gt;117 million passwords&lt;/strong&gt; from the same breach resurfaced on the dark web.
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Why Passwords Failed&lt;/strong&gt;
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Weak Hashing:&lt;/strong&gt; LinkedIn stored passwords with weak encryption, making them easy to crack.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Password Reuse:&lt;/strong&gt; Many users reused the same passwords across multiple sites, leading to &lt;strong&gt;credential stuffing attacks&lt;/strong&gt; on other platforms.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delayed Impact:&lt;/strong&gt; Even years later, these passwords were still being used in attacks.
&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Aftermath&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;LinkedIn forced password resets, but the damage was done.
&lt;/li&gt;
&lt;li&gt;Many users who reused passwords saw their other accounts (email, banking, social media) compromised.
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;How It Could Have Been Prevented&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;✅ &lt;strong&gt;Passwordless Auth:&lt;/strong&gt; Passkeys or biometric logins would have made stolen credentials useless.&lt;br&gt;&lt;br&gt;
✅ &lt;strong&gt;Better Hashing:&lt;/strong&gt; Modern algorithms (bcrypt, Argon2) could have slowed down cracking.&lt;br&gt;&lt;br&gt;
✅ &lt;strong&gt;MFA Enforcement:&lt;/strong&gt; Even with leaked passwords, MFA would have blocked unauthorized access.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Case Study 2: The Colonial Pipeline Ransomware Attack (2021)&lt;/strong&gt;
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;What Happened?&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Hackers breached &lt;strong&gt;Colonial Pipeline&lt;/strong&gt;, a major U.S. fuel supplier, causing a &lt;strong&gt;six-day shutdown&lt;/strong&gt; and fuel shortages across the East Coast.
&lt;/li&gt;
&lt;li&gt;The attack started with &lt;strong&gt;a single compromised password&lt;/strong&gt; to an old VPN account that &lt;strong&gt;lacked multi-factor authentication (MFA).&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Why Passwords Failed&lt;/strong&gt;
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;No MFA:&lt;/strong&gt; A single weak password was all hackers needed to infiltrate the network.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Legacy Account Exposure:&lt;/strong&gt; The VPN account was no longer in use but wasn’t deactivated.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Password Reuse:&lt;/strong&gt; The password may have been reused or easily guessed (though exact details weren’t disclosed).
&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Aftermath&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Colonial Pipeline paid &lt;strong&gt;$4.4 million in Bitcoin&lt;/strong&gt; as ransom.
&lt;/li&gt;
&lt;li&gt;The U.S. government recovered some funds, but the incident highlighted how &lt;strong&gt;passwords alone are insufficient for critical infrastructure.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;How It Could Have Been Prevented&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;✅ &lt;strong&gt;Passwordless VPN Access:&lt;/strong&gt; A hardware security key (YubiKey) or certificate-based auth would have prevented the breach.&lt;br&gt;&lt;br&gt;
✅ &lt;strong&gt;Strict MFA Policies:&lt;/strong&gt; Even a simple TOTP (Google Authenticator) check would have stopped the attack.&lt;br&gt;&lt;br&gt;
✅ &lt;strong&gt;Automated Account Deactivation:&lt;/strong&gt; Unused accounts should be disabled automatically.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;The Way Forward: Killing the Password for Good&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;These cases prove that &lt;strong&gt;passwords alone are a security liability.&lt;/strong&gt; Here’s what we should adopt instead:  &lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;1. Passkeys (FIDO2 / WebAuthn)&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;No passwords, just biometrics or hardware keys.
&lt;/li&gt;
&lt;li&gt;Immune to phishing &amp;amp; credential stuffing.
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;2. Universal MFA Adoption&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mandate MFA everywhere&lt;/strong&gt;, especially for remote access and admin accounts.
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;3. Better Credential Management&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Password managers&lt;/strong&gt; for generating and storing strong passwords (if still needed).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Regular audits&lt;/strong&gt; to deactivate unused accounts.
&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  &lt;strong&gt;Final Thoughts&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Passwords are outdated, insecure, and costly. The LinkedIn and Colonial Pipeline breaches show just how dangerous reliance on passwords can be. The sooner we move to &lt;strong&gt;passwordless authentication&lt;/strong&gt;, the safer we’ll all be.  &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are you still using passwords, or have you switched to passkeys/MFA? Share your experience below!&lt;/strong&gt; 🔐  &lt;/p&gt;

</description>
      <category>security</category>
      <category>webdev</category>
      <category>cybersecurity</category>
      <category>devops</category>
    </item>
    <item>
      <title>Stop Using Keycloak Like a Basic Auth Server—Try These Features</title>
      <dc:creator>Samuel Mutemi</dc:creator>
      <pubDate>Tue, 29 Apr 2025 14:20:02 +0000</pubDate>
      <link>https://dev.to/mtsammy40/stop-using-keycloak-like-a-basic-auth-server-try-these-features-5dkm</link>
      <guid>https://dev.to/mtsammy40/stop-using-keycloak-like-a-basic-auth-server-try-these-features-5dkm</guid>
      <description>&lt;p&gt;Keycloak is widely recognized as a powerful open-source Identity and Access Management (IAM) solution, offering SSO, OAuth, and user federation out of the box. But beyond the basics, Keycloak has several underrated features that can significantly enhance security, customization, and usability.  &lt;/p&gt;

&lt;p&gt;In this post, we’ll explore some of these &lt;strong&gt;lesser-known Keycloak features&lt;/strong&gt; that can save you time, improve security, and unlock new capabilities.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;1. Fine-Grained Admin Permissions (Client Policies &amp;amp; Admin Fine-Grained AuthZ)&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Most Keycloak admins know about realm roles, but did you know you can &lt;strong&gt;delegate admin permissions with surgical precision&lt;/strong&gt;?  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Client Policies&lt;/strong&gt;: Define rules for client registrations (e.g., enforce HTTPS redirect URIs).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fine-Grained Admin Permissions&lt;/strong&gt;: Restrict admin console access (e.g., allow a user to manage only specific clients).
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔹 &lt;strong&gt;Use Case&lt;/strong&gt;: Securely delegate administration without giving full realm access.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;2. Dynamic Client Registration (DCR) &amp;amp; Initial Access Tokens&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Instead of manually configuring every client, Keycloak supports &lt;strong&gt;dynamic client registration&lt;/strong&gt; via REST API.  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Initial Access Tokens&lt;/strong&gt;: Generate short-lived tokens to allow self-service client registration.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Client Registration Policies&lt;/strong&gt;: Enforce constraints (e.g., allowed redirect URIs).
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔹 &lt;strong&gt;Use Case&lt;/strong&gt;: SaaS platforms where tenants need to onboard their own apps securely.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;3. WebAuthn &amp;amp; Passwordless Authentication&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;While OTP is common, Keycloak supports &lt;strong&gt;FIDO2/WebAuthn&lt;/strong&gt; for phishing-resistant logins.  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Biometric &amp;amp; Security Key Auth&lt;/strong&gt;: Users can log in via Face ID, Touch ID, or YubiKey.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conditional Policies&lt;/strong&gt;: Require WebAuthn only for high-risk actions.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔹 &lt;strong&gt;Use Case&lt;/strong&gt;: Financial apps or internal systems needing stronger MFA.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;4. Token Exchange (Impersonation &amp;amp; Delegation)&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Keycloak allows &lt;strong&gt;token exchange&lt;/strong&gt;, letting one token be swapped for another under controlled conditions.  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Impersonation&lt;/strong&gt;: Admins can act on behalf of users (for debugging).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delegation&lt;/strong&gt;: A service can obtain a token for another service securely.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔹 &lt;strong&gt;Use Case&lt;/strong&gt;: Microservices architectures where services need to call each other.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;5. Custom User Attributes &amp;amp; Declarative User Profiles&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Keycloak supports &lt;strong&gt;custom user metadata&lt;/strong&gt;, but with &lt;strong&gt;Declarative User Profiles (DUP)&lt;/strong&gt;, you can enforce validation.  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Define Required Fields&lt;/strong&gt;: Mandate certain attributes at registration.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validation Rules&lt;/strong&gt;: Enforce email formats, phone numbers, etc.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔹 &lt;strong&gt;Use Case&lt;/strong&gt;: Compliance-heavy industries needing structured user data.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;6. Event Listeners &amp;amp; Webhooks&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Keycloak emits &lt;strong&gt;real-time events&lt;/strong&gt; (logins, token exchanges, failures). You can:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Forward Events to Kafka/RabbitMQ&lt;/strong&gt;: For SIEM integration.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trigger Webhooks&lt;/strong&gt;: Notify external systems on user actions.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔹 &lt;strong&gt;Use Case&lt;/strong&gt;: Fraud detection, audit logging, or real-time analytics.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;7. Script-Based Authentication Flows&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Instead of hardcoding auth logic, Keycloak supports &lt;strong&gt;JavaScript/Python scripts&lt;/strong&gt; in auth flows.  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Custom Validation&lt;/strong&gt;: Check user attributes before login.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conditional Steps&lt;/strong&gt;: Skip MFA for trusted IPs.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔹 &lt;strong&gt;Use Case&lt;/strong&gt;: Adaptive authentication without custom code.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;8. Lightweight Directory Services (LDAP) with Write-Back&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Keycloak can sync with LDAP/Active Directory &lt;strong&gt;bidirectionally&lt;/strong&gt;.  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;User Write-Back&lt;/strong&gt;: Changes in Keycloak (e.g., password updates) sync back to LDAP.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On-Demand Sync&lt;/strong&gt;: Avoid full syncs with lazy loading.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔹 &lt;strong&gt;Use Case&lt;/strong&gt;: Enterprises migrating from legacy LDAP systems.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;9. Built-In Token Revocation &amp;amp; Offline Sessions&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Revoking tokens is usually manual, but Keycloak offers:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not-Before Policy&lt;/strong&gt;: Invalidate all tokens issued before a certain time.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offline Session Limits&lt;/strong&gt;: Control how long refresh tokens last.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔹 &lt;strong&gt;Use Case&lt;/strong&gt;: Responding to breaches or employee offboarding.  &lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;10. Themeable Emails &amp;amp; Localization&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Most Keycloak emails look generic, but you can:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Customize Templates&lt;/strong&gt;: Branded password reset emails.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-Language Support&lt;/strong&gt;: Auto-send emails in the user’s language.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔹 &lt;strong&gt;Use Case&lt;/strong&gt;: Improving UX for global user bases.  &lt;/p&gt;




&lt;h3&gt;
  
  
  &lt;strong&gt;Final Thoughts&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Keycloak is far more than just an OAuth server—it’s a &lt;strong&gt;swiss-army knife for IAM&lt;/strong&gt;. By leveraging these lesser-known features, you can:&lt;br&gt;&lt;br&gt;
✅ &lt;strong&gt;Enhance security&lt;/strong&gt; (WebAuthn, token exchange)&lt;br&gt;&lt;br&gt;
✅ &lt;strong&gt;Automate workflows&lt;/strong&gt; (dynamic client registration)&lt;br&gt;&lt;br&gt;
✅ &lt;strong&gt;Improve UX&lt;/strong&gt; (custom emails, declarative profiles)  &lt;/p&gt;

&lt;p&gt;Have you used any of these features? Any hidden gems I missed? Let me know in the comments! 🚀  &lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Further Reading&lt;/strong&gt;:  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.keycloak.org/documentation" rel="noopener noreferrer"&gt;Keycloak Docs&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/keycloak/keycloak" rel="noopener noreferrer"&gt;Keycloak GitHub&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Would you like a deep dive into any of these features? Let me know!&lt;/p&gt;

</description>
      <category>security</category>
      <category>webdev</category>
      <category>microservices</category>
      <category>oauth</category>
    </item>
  </channel>
</rss>
