<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Parsa Mohammadi</title>
    <description>The latest articles on DEV Community by Parsa Mohammadi (@parsa_m).</description>
    <link>https://dev.to/parsa_m</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4123367%2F2199da8a-7ce0-4062-b7d9-0a75f65bf20d.png</url>
      <title>DEV Community: Parsa Mohammadi</title>
      <link>https://dev.to/parsa_m</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/parsa_m"/>
    <language>en</language>
    <item>
      <title>How to Assess the Blast Radius of a Code Change</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Mon, 05 Oct 2026 20:00:00 +0000</pubDate>
      <link>https://dev.to/parsa_m/how-to-assess-the-blast-radius-of-a-code-change-342l</link>
      <guid>https://dev.to/parsa_m/how-to-assess-the-blast-radius-of-a-code-change-342l</guid>
      <description>&lt;p&gt;A code change can be small in lines of code and large in production impact.&lt;/p&gt;

&lt;p&gt;That is the basic idea behind blast radius.&lt;/p&gt;

&lt;p&gt;The blast radius of a change is the set of systems, components, users, or business functions that could be affected if the change behaves unexpectedly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the changed component
&lt;/h2&gt;

&lt;p&gt;Identify exactly what the pull request modifies.&lt;/p&gt;

&lt;p&gt;Then ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is it shared?&lt;/li&gt;
&lt;li&gt;Is it customer facing?&lt;/li&gt;
&lt;li&gt;Is it part of a critical workflow?&lt;/li&gt;
&lt;li&gt;Does it interact with a database?&lt;/li&gt;
&lt;li&gt;Does it handle authentication?&lt;/li&gt;
&lt;li&gt;Does it control infrastructure?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Map direct dependencies
&lt;/h2&gt;

&lt;p&gt;Find the components that directly depend on the changed code.&lt;/p&gt;

&lt;p&gt;Then look one level further.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Changed component
      ↓
Direct dependencies
      ↓
Dependent services
      ↓
Customer workflows
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is to understand propagation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Look at indirect impact
&lt;/h2&gt;

&lt;p&gt;Some changes affect components that don't appear in the pull request.&lt;/p&gt;

&lt;p&gt;For example, changing a shared API contract can affect multiple services even when those services are untouched.&lt;/p&gt;

&lt;p&gt;That is why dependency mapping matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check production usage
&lt;/h2&gt;

&lt;p&gt;Not every dependency has equal importance.&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How heavily is it used?&lt;/li&gt;
&lt;li&gt;Which customers depend on it?&lt;/li&gt;
&lt;li&gt;Which workflows rely on it?&lt;/li&gt;
&lt;li&gt;Is it part of a critical path?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A shared component with low production usage may have a smaller practical blast radius than a heavily used component.&lt;/p&gt;

&lt;h2&gt;
  
  
  Look at history
&lt;/h2&gt;

&lt;p&gt;Previous incidents can reveal where changes have caused problems before.&lt;/p&gt;

&lt;p&gt;Also look at recent changes and regressions in the same area.&lt;/p&gt;

&lt;h2&gt;
  
  
  Consider the deployment
&lt;/h2&gt;

&lt;p&gt;Blast radius also depends on how a change is deployed.&lt;/p&gt;

&lt;p&gt;A gradual rollout, feature flag, or quick rollback can limit the impact of a failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Diff size is not blast radius
&lt;/h2&gt;

&lt;p&gt;A useful rule is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Lines changed ≠ production impact&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A five line change to a shared authentication library can matter more than a 500 line change to an isolated internal tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where PRI fits
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;Production Reliability Index (PRI)&lt;/strong&gt; can use change context, dependencies, production signals, and other reliability evidence to help surface changes that deserve additional attention.&lt;/p&gt;

&lt;p&gt;Run a PRI assessment:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://tomosu.ai/start" rel="noopener noreferrer"&gt;https://tomosu.ai/start&lt;/a&gt;&lt;/p&gt;

</description>
      <category>programming</category>
      <category>devops</category>
      <category>softwareengineering</category>
      <category>sre</category>
    </item>
    <item>
      <title>Hacktoberfest tip before you submit: break one function on purpose (return []) and run your tests. If the suite still passes, those tests check nothing. Fix them or delete them, a fake green is worse than no test. #hacktoberfest</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Mon, 05 Oct 2026 12:17:44 +0000</pubDate>
      <link>https://dev.to/parsa_m/hacktoberfest-tip-before-you-submit-break-one-function-on-purpose-return-and-run-your-tests-8j3</link>
      <guid>https://dev.to/parsa_m/hacktoberfest-tip-before-you-submit-break-one-function-on-purpose-return-and-run-your-tests-8j3</guid>
      <description></description>
      <category>hacktoberfest</category>
      <category>opensource</category>
      <category>programming</category>
      <category>testing</category>
    </item>
    <item>
      <title>Agent wrote the code and the tests? Break one function on purpose. If tests pass, they're fake.</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:23:06 +0000</pubDate>
      <link>https://dev.to/parsa_m/-3b1p</link>
      <guid>https://dev.to/parsa_m/-3b1p</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/parsa_m/your-hacktoberfest-tests-are-green-that-doesnt-mean-they-ran-4jn0" class="crayons-story__hidden-navigation-link"&gt;Your Hacktoberfest tests are green. That doesn't mean they ran.&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
      &lt;a href="https://dev.to/parsa_m/your-hacktoberfest-tests-are-green-that-doesnt-mean-they-ran-4jn0" class="crayons-article__context-note crayons-article__context-note__feed"&gt;&lt;p&gt;Hacktoberfest: Contribution Chronicles&lt;/p&gt;

&lt;/a&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/parsa_m" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4123367%2F2199da8a-7ce0-4062-b7d9-0a75f65bf20d.png" alt="parsa_m profile" class="crayons-avatar__image" width="800" height="800"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/parsa_m" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Parsa Mohammadi
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Parsa Mohammadi
                
                
              
              &lt;div id="story-author-preview-content-4800730" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/parsa_m" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4123367%2F2199da8a-7ce0-4062-b7d9-0a75f65bf20d.png" class="crayons-avatar__image" alt="" width="800" height="800"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Parsa Mohammadi&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/parsa_m/your-hacktoberfest-tests-are-green-that-doesnt-mean-they-ran-4jn0" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Oct 5&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/parsa_m/your-hacktoberfest-tests-are-green-that-doesnt-mean-they-ran-4jn0" id="article-link-4800730"&gt;
          Your Hacktoberfest tests are green. That doesn't mean they ran.
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag crayons-tag--filled  " href="/t/discuss"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;discuss&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/hacktoberfest"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;hacktoberfest&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/testing"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;testing&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/parsa_m/your-hacktoberfest-tests-are-green-that-doesnt-mean-they-ran-4jn0" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;1&lt;span class="hidden s:inline"&gt;&amp;nbsp;reaction&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/parsa_m/your-hacktoberfest-tests-are-green-that-doesnt-mean-they-ran-4jn0#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            3 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Your Hacktoberfest tests are green. That doesn't mean they ran.</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Mon, 05 Oct 2026 09:35:59 +0000</pubDate>
      <link>https://dev.to/parsa_m/your-hacktoberfest-tests-are-green-that-doesnt-mean-they-ran-4jn0</link>
      <guid>https://dev.to/parsa_m/your-hacktoberfest-tests-are-green-that-doesnt-mean-they-ran-4jn0</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; If an agent wrote your code and your tests, a green build means the agent is satisfied. Whether the project works is a separate question. Break one function on purpose before you submit. If the suite still passes, those tests aren't checking anything. Checklist and a free scan at the end.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;More and more Hacktoberfest projects are built with coding agents. That's fine. The risky moment is the last hour, when the agent says "all tests passing ✅" and you believe it because you're tired and the deadline is tonight.&lt;/p&gt;

&lt;p&gt;Coding agents are generally optimized to satisfy the checks they're given. If the tests are weak, an agent can produce a green build without proving much about the implementation. Sometimes it gets there by making the code work. Sometimes it gets there by making the tests stop checking anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a fake green looks like
&lt;/h2&gt;

&lt;p&gt;None of these fail. All of them look fine at a glance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The assertion that can't fail:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_parse_transactions&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parse_transactions&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SAMPLE_SMS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;  &lt;span class="c1"&gt;# it never returns None, so this passes forever
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The test that quietly left:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@pytest.mark.skip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;flaky, fix later&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_handles_empty_input&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="bp"&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The mock that answers its own question:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@patch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;app.model.classify&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;return_value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;groceries&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_classifies_groceries&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mock_classify&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;categorize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;grocery order #4412&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;groceries&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="c1"&gt;# the real model never ran
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The try/except that eats the failure:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_model_loads&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;load_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemma-3-4b-q4.gguf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;pass&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The one check that catches most of it
&lt;/h2&gt;

&lt;p&gt;Break your own code on purpose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;parse_transactions&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;  &lt;span class="c1"&gt;# temporary sabotage
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the suite. If it still passes, those tests weren't testing that function. Revert, then fix the tests or delete them. A deleted fake test is more honest than a green one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rest of the pre-submit list
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Put agent work through pull requests&lt;/strong&gt;, even solo. You get a diff to read, and your commit history shows the project was built inside the challenge window, which is a rule.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Search for secrets&lt;/strong&gt; before the repo goes public: &lt;code&gt;sk-&lt;/code&gt;, &lt;code&gt;api_key&lt;/code&gt;, &lt;code&gt;token&lt;/code&gt;, and every service name you used. If a key was ever committed, rotate it. Deleting the line doesn't remove it from git history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clone fresh and follow your own README.&lt;/strong&gt; Judges will. The usual misses are a model file nothing downloads, an undocumented env var, and a globally installed dependency.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write down the model&lt;/strong&gt;: name, size, quantization, and the hardware you tested on. The prompt asks why open-source AI matters for your project, and "runs on a laptop" is a claim readers will check.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Note post-deadline commits in your README.&lt;/strong&gt; The challenge FAQ says skipping this can get an entry disqualified.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Credit borrowed code.&lt;/strong&gt; If your agent pulled a chunk of someone else's project in, find it and credit it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A second check before you submit, free
&lt;/h2&gt;

&lt;p&gt;If you'd rather not do all of this by eye at 11pm, &lt;a href="https://tomosu.ai" rel="noopener noreferrer"&gt;Tomosu AI&lt;/a&gt; can scan the project for you. There are two free ways in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Small repo (under 50 files): scan it online.&lt;/strong&gt; Paste your GitHub repo URL at &lt;a href="https://tomosu.ai/start" rel="noopener noreferrer"&gt;tomosu.ai/start&lt;/a&gt;. Public repos scan without a login, and the scan is read-only. Most one-week hackathon projects fit under the limit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bigger repo: use the free editor extension.&lt;/strong&gt; Install it in VS Code, Cursor, or Antigravity and scan your workspace from the editor:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://marketplace.visualstudio.com/items?itemName=tomosuai.prod-free-tomosu-ai-vscode-plugin" rel="noopener noreferrer"&gt;VS Code Marketplace&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://open-vsx.org/extension/tomosuai/prod-free-tomosu-ai-vscode-plugin" rel="noopener noreferrer"&gt;Open VSX&lt;/a&gt; (for Cursor, Antigravity, and other VS Code forks)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://tomosu.ai/docs/vscode-installation" rel="noopener noreferrer"&gt;Installation guide&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Either way you get a Production Reliability Index (PRI) score for the project. Run it once before your final commit, fix what it flags, and run it again.&lt;/p&gt;

&lt;p&gt;A screenshot of the first scan, the second scan, and what you changed in between makes a good "show your work" section. Judges weight the write-up most heavily, and a real before/after tells them more about your process than a feature list.&lt;/p&gt;

&lt;h2&gt;
  
  
  Over to you
&lt;/h2&gt;

&lt;p&gt;What's the sneakiest thing an agent has done to get your tests green? Drop it in the comments. The best ones go into a follow-up post, with credit.&lt;/p&gt;

</description>
      <category>discuss</category>
      <category>hacktoberfest</category>
      <category>ai</category>
      <category>testing</category>
    </item>
    <item>
      <title>Pre Merge Reliability Analysis</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Sun, 04 Oct 2026 20:00:00 +0000</pubDate>
      <link>https://dev.to/parsa_m/pre-merge-reliability-analysis-3df</link>
      <guid>https://dev.to/parsa_m/pre-merge-reliability-analysis-3df</guid>
      <description>&lt;p&gt;Most code review happens before merge.&lt;/p&gt;

&lt;p&gt;That makes the pull request a useful place to assess more than implementation quality.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;pre merge reliability analysis&lt;/strong&gt; looks at the proposed change together with the production context around it.&lt;/p&gt;

&lt;p&gt;The goal is simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Understand the potential production impact before the change becomes part of the next release.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Start with change scope
&lt;/h2&gt;

&lt;p&gt;Identify what the pull request changes.&lt;/p&gt;

&lt;p&gt;Look at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Files&lt;/li&gt;
&lt;li&gt;Services&lt;/li&gt;
&lt;li&gt;APIs&lt;/li&gt;
&lt;li&gt;Databases&lt;/li&gt;
&lt;li&gt;Infrastructure&lt;/li&gt;
&lt;li&gt;Shared components&lt;/li&gt;
&lt;li&gt;Customer facing behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then ask what is indirectly affected.&lt;/p&gt;

&lt;h2&gt;
  
  
  Map dependencies
&lt;/h2&gt;

&lt;p&gt;A change can propagate through dependencies.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PR
 ↓
Changed component
 ↓
Dependencies
 ↓
Production services
 ↓
Users and workflows
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is often more useful than looking at the number of changed lines.&lt;/p&gt;

&lt;h2&gt;
  
  
  Review testing evidence
&lt;/h2&gt;

&lt;p&gt;Ask whether the important behavior is actually verified.&lt;/p&gt;

&lt;p&gt;Consider:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Unit tests&lt;/li&gt;
&lt;li&gt;Integration tests&lt;/li&gt;
&lt;li&gt;End to end tests&lt;/li&gt;
&lt;li&gt;Regression tests&lt;/li&gt;
&lt;li&gt;Failure scenarios&lt;/li&gt;
&lt;li&gt;Dependency behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A green test suite is evidence.&lt;/p&gt;

&lt;p&gt;It is not proof that every production condition has been covered.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add production context
&lt;/h2&gt;

&lt;p&gt;Look at how the affected components behave in production.&lt;/p&gt;

&lt;p&gt;Consider traffic, customer usage, runtime anomalies, and operational constraints.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check history
&lt;/h2&gt;

&lt;p&gt;Previous incidents, rollbacks, and repeated regressions can provide useful context.&lt;/p&gt;

&lt;p&gt;The goal is not to automatically reject changes in areas with a history of problems.&lt;/p&gt;

&lt;p&gt;The goal is to make that history visible during review.&lt;/p&gt;

&lt;h2&gt;
  
  
  Consider deployment conditions
&lt;/h2&gt;

&lt;p&gt;Ask whether the change can be safely deployed and reversed.&lt;/p&gt;

&lt;p&gt;A change with a large potential impact may have manageable operational risk if it can be deployed gradually and rolled back quickly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should block a merge?
&lt;/h2&gt;

&lt;p&gt;There is no universal threshold.&lt;/p&gt;

&lt;p&gt;Different teams have different systems, deployment processes, and tolerance for uncertainty.&lt;/p&gt;

&lt;p&gt;The useful question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Does the available evidence justify merging this change under the current deployment conditions?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Where PRI fits
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;Production Reliability Index (PRI)&lt;/strong&gt; can provide a summary signal around the change while preserving the underlying findings that explain it.&lt;/p&gt;

&lt;p&gt;That makes it useful as another input into the merge decision rather than a replacement for review.&lt;/p&gt;

&lt;p&gt;Run a PRI assessment:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://tomosu.ai/start" rel="noopener noreferrer"&gt;https://tomosu.ai/start&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How to Measure Software Change Reliability</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Thu, 01 Oct 2026 20:00:00 +0000</pubDate>
      <link>https://dev.to/parsa_m/how-to-measure-software-change-reliability-227h</link>
      <guid>https://dev.to/parsa_m/how-to-measure-software-change-reliability-227h</guid>
      <description>&lt;p&gt;Teams have plenty of metrics for services.&lt;/p&gt;

&lt;p&gt;They measure availability, latency, error rates, deployment frequency, and incident recovery.&lt;/p&gt;

&lt;p&gt;But there is another useful question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How reliable is this particular software change?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a change level question rather than a service level question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Service reliability vs change reliability
&lt;/h2&gt;

&lt;p&gt;Service level metrics describe how a system behaves over time.&lt;/p&gt;

&lt;p&gt;Change reliability asks how a specific modification is likely to affect that system.&lt;/p&gt;

&lt;p&gt;For example, a service may have excellent availability while a particular change still has a large potential blast radius.&lt;/p&gt;

&lt;p&gt;Both views are useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  What signals can you measure?
&lt;/h2&gt;

&lt;p&gt;There is no universally accepted formula for software change reliability.&lt;/p&gt;

&lt;p&gt;Instead, start with useful evidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Change scope
&lt;/h3&gt;

&lt;p&gt;What components, files, APIs, databases, and infrastructure are affected?&lt;/p&gt;

&lt;h3&gt;
  
  
  Dependencies
&lt;/h3&gt;

&lt;p&gt;How many systems depend on the changed component?&lt;/p&gt;

&lt;h3&gt;
  
  
  Testing
&lt;/h3&gt;

&lt;p&gt;What evidence exists that the important behavior works?&lt;/p&gt;

&lt;h3&gt;
  
  
  Production behavior
&lt;/h3&gt;

&lt;p&gt;What do runtime signals tell you about the affected area?&lt;/p&gt;

&lt;h3&gt;
  
  
  Code volatility
&lt;/h3&gt;

&lt;p&gt;How frequently does this part of the system change?&lt;/p&gt;

&lt;h3&gt;
  
  
  Incident history
&lt;/h3&gt;

&lt;p&gt;Have similar changes caused problems before?&lt;/p&gt;

&lt;h3&gt;
  
  
  Deployment outcomes
&lt;/h3&gt;

&lt;p&gt;What happened to similar changes after deployment?&lt;/p&gt;

&lt;h2&gt;
  
  
  Why one metric is difficult
&lt;/h2&gt;

&lt;p&gt;A single number can hide important context.&lt;/p&gt;

&lt;p&gt;Consider two changes:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Change A&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Small diff&lt;/li&gt;
&lt;li&gt;Shared library&lt;/li&gt;
&lt;li&gt;Limited integration testing&lt;/li&gt;
&lt;li&gt;Recent incident history&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Change B&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Large diff&lt;/li&gt;
&lt;li&gt;Isolated internal service&lt;/li&gt;
&lt;li&gt;Strong integration coverage&lt;/li&gt;
&lt;li&gt;Easy rollback&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The second diff is larger.&lt;/p&gt;

&lt;p&gt;That does not automatically make it less reliable.&lt;/p&gt;

&lt;p&gt;The point of measurement is to bring the relevant evidence together.&lt;/p&gt;

&lt;h2&gt;
  
  
  What makes a useful reliability measure?
&lt;/h2&gt;

&lt;p&gt;A useful measurement should be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Relevant to the change&lt;/li&gt;
&lt;li&gt;Traceable to evidence&lt;/li&gt;
&lt;li&gt;Actionable&lt;/li&gt;
&lt;li&gt;Comparable over time&lt;/li&gt;
&lt;li&gt;Clear about uncertainty&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A score is useful when engineers can understand why it looks the way it does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where PRI fits
&lt;/h2&gt;

&lt;p&gt;Tomosu's &lt;strong&gt;Production Reliability Index (PRI)&lt;/strong&gt; provides a summary signal across multiple areas of reliability.&lt;/p&gt;

&lt;p&gt;The signal can help teams identify which changes deserve closer investigation while keeping the underlying evidence visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main idea
&lt;/h2&gt;

&lt;p&gt;Measuring software change reliability is not about finding one perfect formula.&lt;/p&gt;

&lt;p&gt;It is about consistently combining the signals that matter:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Change + dependencies + testing + production behavior + history + deployment outcomes&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That gives teams a way to discuss reliability at the level where changes actually happen.&lt;/p&gt;

&lt;p&gt;Run a PRI assessment:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://tomosu.ai/start" rel="noopener noreferrer"&gt;https://tomosu.ai/start&lt;/a&gt;&lt;/p&gt;

</description>
      <category>softwareengineering</category>
      <category>sre</category>
      <category>devops</category>
      <category>codereview</category>
    </item>
    <item>
      <title>How to Assess Production Reliability Before Deployment</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Wed, 30 Sep 2026 20:00:00 +0000</pubDate>
      <link>https://dev.to/parsa_m/how-to-assess-production-reliability-before-deployment-4g3h</link>
      <guid>https://dev.to/parsa_m/how-to-assess-production-reliability-before-deployment-4g3h</guid>
      <description>&lt;p&gt;Most teams assess a change using code review and tests.&lt;/p&gt;

&lt;p&gt;That catches a lot. But it doesn't answer everything.&lt;/p&gt;

&lt;p&gt;A change can pass the test suite and still create problems in production because of its dependencies, how the affected code is used, recent incidents, or the size of its potential blast radius.&lt;/p&gt;

&lt;p&gt;So the question before deployment is not just:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Does this code work?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How is this change likely to behave in production?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Start with the change itself
&lt;/h2&gt;

&lt;p&gt;First, understand what the change actually touches.&lt;/p&gt;

&lt;p&gt;Look at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Files and components changed&lt;/li&gt;
&lt;li&gt;APIs modified&lt;/li&gt;
&lt;li&gt;Database changes&lt;/li&gt;
&lt;li&gt;Configuration changes&lt;/li&gt;
&lt;li&gt;Authentication or authorization logic&lt;/li&gt;
&lt;li&gt;Shared libraries&lt;/li&gt;
&lt;li&gt;Customer facing functionality&lt;/li&gt;
&lt;li&gt;Infrastructure changes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A small diff isn't necessarily a small production change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Look at dependencies
&lt;/h2&gt;

&lt;p&gt;Next, ask what depends on the code being changed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Change
  ↓
Dependencies
  ↓
Affected Components
  ↓
Production Usage
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The further a change can propagate, the more context you need before deploying it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check the testing evidence
&lt;/h2&gt;

&lt;p&gt;Tests provide evidence about a change. They don't prove that every production scenario will behave correctly.&lt;/p&gt;

&lt;p&gt;Look at unit, integration, end to end, regression, and failure tests.&lt;/p&gt;

&lt;p&gt;Then ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What important production behavior is still unverified?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Look at production behavior
&lt;/h2&gt;

&lt;p&gt;Production context can change how a change should be evaluated.&lt;/p&gt;

&lt;p&gt;Consider:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Real customer traffic&lt;/li&gt;
&lt;li&gt;Large datasets&lt;/li&gt;
&lt;li&gt;Concurrent requests&lt;/li&gt;
&lt;li&gt;Production configuration&lt;/li&gt;
&lt;li&gt;External service failures&lt;/li&gt;
&lt;li&gt;Unusual inputs&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Check incident and change history
&lt;/h2&gt;

&lt;p&gt;Look at previous incidents, rollbacks, related fixes, change frequency, and repeated failures.&lt;/p&gt;

&lt;p&gt;History doesn't decide whether a change is safe.&lt;/p&gt;

&lt;p&gt;It provides context.&lt;/p&gt;

&lt;h2&gt;
  
  
  Consider code volatility
&lt;/h2&gt;

&lt;p&gt;A component that changes constantly can require different attention from one that has remained stable for months.&lt;/p&gt;

&lt;p&gt;Look at how often it changes and whether previous changes caused problems.&lt;/p&gt;

&lt;h2&gt;
  
  
  Assess deployment and rollback conditions
&lt;/h2&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Can the change be rolled back quickly?&lt;/li&gt;
&lt;li&gt;Can it be deployed gradually?&lt;/li&gt;
&lt;li&gt;Is a feature flag available?&lt;/li&gt;
&lt;li&gt;Are database changes reversible?&lt;/li&gt;
&lt;li&gt;Can the affected functionality be disabled?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A simple framework
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Change
  ↓
Evidence
  ↓
Uncertainty
  ↓
Action
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Understand the change.&lt;/p&gt;

&lt;p&gt;Collect evidence.&lt;/p&gt;

&lt;p&gt;Identify what remains uncertain.&lt;/p&gt;

&lt;p&gt;Then decide what additional action makes sense.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pre deployment checklist
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[ ] What changed?&lt;/li&gt;
&lt;li&gt;[ ] What components are affected?&lt;/li&gt;
&lt;li&gt;[ ] What dependencies are involved?&lt;/li&gt;
&lt;li&gt;[ ] Is the important behavior tested?&lt;/li&gt;
&lt;li&gt;[ ] Has this area caused incidents before?&lt;/li&gt;
&lt;li&gt;[ ] How large could the blast radius be?&lt;/li&gt;
&lt;li&gt;[ ] Can the change be rolled back?&lt;/li&gt;
&lt;li&gt;[ ] What monitoring will be available after deployment?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where PRI fits
&lt;/h2&gt;

&lt;p&gt;Tomosu's &lt;strong&gt;Production Reliability Index (PRI)&lt;/strong&gt; brings multiple reliability signals together to give engineers a change level signal before deployment.&lt;/p&gt;

&lt;p&gt;The goal is not to replace engineering judgment.&lt;/p&gt;

&lt;p&gt;It is to surface useful evidence earlier.&lt;/p&gt;

&lt;p&gt;Run a PRI assessment:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://tomosu.ai/start" rel="noopener noreferrer"&gt;https://tomosu.ai/start&lt;/a&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>sre</category>
      <category>softwareengineering</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Production Reliability vs Observability</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Tue, 29 Sep 2026 20:00:00 +0000</pubDate>
      <link>https://dev.to/parsa_m/production-reliability-vs-code-review-4f4b</link>
      <guid>https://dev.to/parsa_m/production-reliability-vs-code-review-4f4b</guid>
      <description>&lt;p&gt;Observability tells you what is happening inside a running system.&lt;/p&gt;

&lt;p&gt;Production reliability asks a different question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What does this change mean for production?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The two are closely related, but they operate at different points in the engineering workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  What observability gives you
&lt;/h2&gt;

&lt;p&gt;Observability helps engineers understand system behavior through signals such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Logs&lt;/li&gt;
&lt;li&gt;Metrics&lt;/li&gt;
&lt;li&gt;Traces&lt;/li&gt;
&lt;li&gt;Errors&lt;/li&gt;
&lt;li&gt;Latency&lt;/li&gt;
&lt;li&gt;Resource usage&lt;/li&gt;
&lt;li&gt;Request behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For AI systems, that can extend to prompts, responses, tool calls, retrieved context, token usage, and evaluation results.&lt;/p&gt;

&lt;p&gt;Observability is extremely useful when you need to understand what is happening.&lt;/p&gt;

&lt;h2&gt;
  
  
  What observability does not tell you before a change ships
&lt;/h2&gt;

&lt;p&gt;Imagine a pull request changes a shared database library.&lt;/p&gt;

&lt;p&gt;Your observability system can tell you what happens after deployment.&lt;/p&gt;

&lt;p&gt;But before deployment, you also want to know:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which services use the library?&lt;/li&gt;
&lt;li&gt;How important are those services?&lt;/li&gt;
&lt;li&gt;What tests cover the change?&lt;/li&gt;
&lt;li&gt;Has the component caused incidents before?&lt;/li&gt;
&lt;li&gt;How large is the potential blast radius?&lt;/li&gt;
&lt;li&gt;Can the deployment be rolled back?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those questions require change context.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production reliability uses observability as evidence
&lt;/h2&gt;

&lt;p&gt;A useful model is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Code change
    +
Production context
    +
Observability signals
    ↓
Reliability assessment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Observability is therefore an input.&lt;/p&gt;

&lt;p&gt;It does not replace the broader assessment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Example
&lt;/h2&gt;

&lt;p&gt;Suppose a change modifies a payment service.&lt;/p&gt;

&lt;p&gt;Observability might show that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Error rates are stable&lt;/li&gt;
&lt;li&gt;Latency is normal&lt;/li&gt;
&lt;li&gt;No dependency is currently failing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is useful.&lt;/p&gt;

&lt;p&gt;But you may still need to know whether the change touches a high volume path, whether the relevant integration tests exist, and whether similar changes caused incidents before.&lt;/p&gt;

&lt;p&gt;Healthy runtime behavior does not automatically mean every new change is reliable.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI systems make this distinction clearer
&lt;/h2&gt;

&lt;p&gt;An AI application can have healthy infrastructure while producing bad results.&lt;/p&gt;

&lt;p&gt;AI observability can help you inspect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt and response behavior&lt;/li&gt;
&lt;li&gt;Tool calls&lt;/li&gt;
&lt;li&gt;Retrieved context&lt;/li&gt;
&lt;li&gt;Latency&lt;/li&gt;
&lt;li&gt;Cost&lt;/li&gt;
&lt;li&gt;Evaluation results&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Production reliability asks what a change to that system could mean before it reaches users.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main difference
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Observability:&lt;/strong&gt; What is happening?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Production reliability:&lt;/strong&gt; What does this change mean for production?&lt;/p&gt;

&lt;p&gt;Both are useful.&lt;/p&gt;

&lt;p&gt;Neither needs to replace the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where PRI fits
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;Production Reliability Index (PRI)&lt;/strong&gt; combines different signals around a software change, including runtime information, code volatility, dependencies, deployment context, and other reliability indicators.&lt;/p&gt;

&lt;p&gt;The purpose is to make change level reliability easier to assess before deployment.&lt;/p&gt;

&lt;p&gt;Run a PRI assessment:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://tomosu.ai/start" rel="noopener noreferrer"&gt;https://tomosu.ai/start&lt;/a&gt;&lt;/p&gt;

&lt;p&gt; &lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>observability</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>Production Reliability vs Static Analysis</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Tue, 29 Sep 2026 20:00:00 +0000</pubDate>
      <link>https://dev.to/parsa_m/production-reliability-vs-static-analysis-4l9f</link>
      <guid>https://dev.to/parsa_m/production-reliability-vs-static-analysis-4l9f</guid>
      <description>&lt;p&gt;Static analysis examines code without running it.&lt;/p&gt;

&lt;p&gt;Production reliability looks at what a change could mean when it reaches the real system.&lt;/p&gt;

&lt;p&gt;Both are useful, but they answer different questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  What static analysis does
&lt;/h2&gt;

&lt;p&gt;Static analysis can identify problems such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Bugs and suspicious patterns&lt;/li&gt;
&lt;li&gt;Type errors&lt;/li&gt;
&lt;li&gt;Security issues&lt;/li&gt;
&lt;li&gt;Code smells&lt;/li&gt;
&lt;li&gt;Dependency problems&lt;/li&gt;
&lt;li&gt;Violations of coding rules&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It works directly on the code and its structure.&lt;/p&gt;

&lt;p&gt;That makes it valuable early in the development process.&lt;/p&gt;

&lt;h2&gt;
  
  
  What static analysis cannot see by itself
&lt;/h2&gt;

&lt;p&gt;A static analyzer does not necessarily know:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which components are most important in production&lt;/li&gt;
&lt;li&gt;How heavily a service is used&lt;/li&gt;
&lt;li&gt;What incidents happened recently&lt;/li&gt;
&lt;li&gt;Which deployment conditions apply&lt;/li&gt;
&lt;li&gt;How a change interacts with current runtime behavior&lt;/li&gt;
&lt;li&gt;How large the practical blast radius could be&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is where production context becomes useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Example: a shared authentication library
&lt;/h2&gt;

&lt;p&gt;Suppose a pull request changes a shared authentication library.&lt;/p&gt;

&lt;p&gt;Static analysis may find no obvious code problems.&lt;/p&gt;

&lt;p&gt;The tests may pass.&lt;/p&gt;

&lt;p&gt;But the library is used across many production services and the affected authentication path has recently been involved in incidents.&lt;/p&gt;

&lt;p&gt;Nothing about the implementation necessarily looks wrong.&lt;/p&gt;

&lt;p&gt;The production context still matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Static analysis is one signal
&lt;/h2&gt;

&lt;p&gt;A broader model looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Static analysis
       +
Testing
       +
Dependencies
       +
Production behavior
       +
Incident history
       +
Deployment conditions
       ↓
Reliability assessment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is not to make one tool responsible for everything.&lt;/p&gt;

&lt;p&gt;It is to combine useful evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI generated code
&lt;/h2&gt;

&lt;p&gt;AI assisted development increases the volume of code changes.&lt;/p&gt;

&lt;p&gt;Static analysis remains useful for checking the implementation.&lt;/p&gt;

&lt;p&gt;But generated code can still introduce changes with unexpected dependencies or production impact.&lt;/p&gt;

&lt;p&gt;That makes context important alongside code level analysis.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main difference
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Static analysis:&lt;/strong&gt; What problems can we identify in the code?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Production reliability:&lt;/strong&gt; What could this change mean for the production system?&lt;/p&gt;

&lt;p&gt;Those questions overlap, but they are not interchangeable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where PRI fits
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;Production Reliability Index (PRI)&lt;/strong&gt; is designed to bring multiple signals around a change together so engineers can identify changes that deserve additional attention.&lt;/p&gt;

&lt;p&gt;Run a PRI assessment:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://tomosu.ai/start" rel="noopener noreferrer"&gt;https://tomosu.ai/start&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>softwareengineering</category>
      <category>codequality</category>
      <category>staticanalysis</category>
    </item>
    <item>
      <title>Production Reliability vs Code Review</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Mon, 28 Sep 2026 20:00:00 +0000</pubDate>
      <link>https://dev.to/parsa_m/production-reliability-vs-code-review-2off</link>
      <guid>https://dev.to/parsa_m/production-reliability-vs-code-review-2off</guid>
      <description>&lt;p&gt;Code review is one of the most established ways engineers evaluate software changes.&lt;/p&gt;

&lt;p&gt;But code review and production reliability are not the same thing.&lt;/p&gt;

&lt;p&gt;A reviewer can determine that an implementation is correct while still not knowing exactly how the change will behave across production dependencies, traffic, configuration, and existing system behavior.&lt;/p&gt;

&lt;p&gt;The two questions are different:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Code review:&lt;/strong&gt; Is the implementation correct?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Production reliability:&lt;/strong&gt; What does this change mean for the running system?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2fep56udl75tc0d992yd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2fep56udl75tc0d992yd.png" alt=" " width="800" height="359"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What code review is good at
&lt;/h2&gt;

&lt;p&gt;Code review can catch problems such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Incorrect logic&lt;/li&gt;
&lt;li&gt;Bad API usage&lt;/li&gt;
&lt;li&gt;Security issues&lt;/li&gt;
&lt;li&gt;Poor error handling&lt;/li&gt;
&lt;li&gt;Maintainability problems&lt;/li&gt;
&lt;li&gt;Unclear implementation choices&lt;/li&gt;
&lt;li&gt;Violations of project conventions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It is primarily focused on the implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What production reliability adds
&lt;/h2&gt;

&lt;p&gt;A production reliability assessment adds context around the change.&lt;/p&gt;

&lt;p&gt;It asks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What components are affected?&lt;/li&gt;
&lt;li&gt;Which dependencies are involved?&lt;/li&gt;
&lt;li&gt;How is this code used in production?&lt;/li&gt;
&lt;li&gt;What testing covers the change?&lt;/li&gt;
&lt;li&gt;Has this area caused incidents before?&lt;/li&gt;
&lt;li&gt;How large could the blast radius be?&lt;/li&gt;
&lt;li&gt;How easy is the change to roll back?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A reviewer may understand the code perfectly and still lack some of this context.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsqzzxfndiphh82rljfn1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsqzzxfndiphh82rljfn1.png" alt=" " width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Example: a shared database library
&lt;/h2&gt;

&lt;p&gt;Imagine a small change to a shared database library.&lt;/p&gt;

&lt;p&gt;The diff is straightforward.&lt;/p&gt;

&lt;p&gt;The code review looks good.&lt;/p&gt;

&lt;p&gt;But the library is used by 40 services, several of those services handle customer transactions, and the affected query path has caused performance problems before.&lt;/p&gt;

&lt;p&gt;The implementation may still be correct.&lt;/p&gt;

&lt;p&gt;The production context simply makes the change more important to investigate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Blast radius matters
&lt;/h2&gt;

&lt;p&gt;The number of changed lines is not the same thing as production impact.&lt;/p&gt;

&lt;p&gt;A small change in a shared component can have a large blast radius.&lt;/p&gt;

&lt;p&gt;A large change in an isolated internal service may have a much smaller one.&lt;/p&gt;

&lt;p&gt;That is why dependency and production usage are important parts of change assessment.&lt;/p&gt;

&lt;h2&gt;
  
  
  What about AI generated code?
&lt;/h2&gt;

&lt;p&gt;AI assisted development makes this distinction more important.&lt;/p&gt;

&lt;p&gt;AI can increase the amount of code produced and the number of changes entering review.&lt;/p&gt;

&lt;p&gt;Code review still matters.&lt;/p&gt;

&lt;p&gt;But the reviewer may need more context to understand what a generated change could affect.&lt;/p&gt;

&lt;p&gt;Production reliability can provide another layer of evidence around the change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code review and observability
&lt;/h2&gt;

&lt;p&gt;Observability answers another question: what is happening after the system is running?&lt;/p&gt;

&lt;p&gt;Production reliability connects change context with production evidence before and around deployment.&lt;/p&gt;

&lt;p&gt;They work together rather than replacing each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical workflow
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Code change
    ↓
Code review
    ↓
Production context
    ↓
Dependencies + testing + history
    ↓
Deployment decision
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is not to replace code review.&lt;/p&gt;

&lt;p&gt;It is to make the decision around the change more informed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where PRI fits
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;Production Reliability Index (PRI)&lt;/strong&gt; is designed to summarize multiple signals around a software change so engineers can identify where additional investigation may be useful.&lt;/p&gt;

&lt;p&gt;The score is not a substitute for reviewing the underlying findings.&lt;/p&gt;

&lt;p&gt;It is a way to bring the signals into one place.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main idea
&lt;/h2&gt;

&lt;p&gt;Code review asks whether the implementation makes sense.&lt;/p&gt;

&lt;p&gt;Production reliability asks what the implementation could mean in production.&lt;/p&gt;

&lt;p&gt;You need both perspectives when the goal is to understand a change before it reaches users.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://tomosu.ai/start" rel="noopener noreferrer"&gt;https://tomosu.ai/start&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>softwareengineering</category>
      <category>devops</category>
      <category>sre</category>
    </item>
    <item>
      <title>What Is Production Reliability?</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Mon, 28 Sep 2026 12:23:10 +0000</pubDate>
      <link>https://dev.to/parsa_m/what-is-production-reliability-33bn</link>
      <guid>https://dev.to/parsa_m/what-is-production-reliability-33bn</guid>
      <description>&lt;p&gt;A change can pass tests, pass code review, and still cause problems in production.&lt;/p&gt;

&lt;p&gt;That is because production reliability is about more than whether code works in isolation.&lt;/p&gt;

&lt;p&gt;It is about how a change behaves when it meets the real system: real traffic, real dependencies, real data, real configuration, and real users.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Production reliability is the ability of a software change to behave correctly and consistently under real production conditions.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F25y8w6hpr9ko2wr6cwkv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F25y8w6hpr9ko2wr6cwkv.png" alt=" " width="800" height="409"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why production reliability is different from test results
&lt;/h2&gt;

&lt;p&gt;Tests tell you what happened under the conditions you tested.&lt;/p&gt;

&lt;p&gt;Production introduces conditions that are difficult to reproduce completely:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Different traffic patterns&lt;/li&gt;
&lt;li&gt;Larger datasets&lt;/li&gt;
&lt;li&gt;External dependency failures&lt;/li&gt;
&lt;li&gt;Configuration differences&lt;/li&gt;
&lt;li&gt;Concurrent requests&lt;/li&gt;
&lt;li&gt;Unexpected inputs&lt;/li&gt;
&lt;li&gt;Existing production behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A change can therefore have strong test evidence while still having meaningful uncertainty around production behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should you look at?
&lt;/h2&gt;

&lt;p&gt;A useful assessment starts with the change itself and then adds context.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Change scope
&lt;/h3&gt;

&lt;p&gt;What files, services, APIs, databases, or infrastructure are affected?&lt;/p&gt;

&lt;p&gt;The size of a diff is useful context, but it does not tell you the full impact.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Dependencies
&lt;/h3&gt;

&lt;p&gt;What depends on the changed component?&lt;/p&gt;

&lt;p&gt;A small change to a shared library can have a much larger production impact than a large change inside an isolated service.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Testing and verification
&lt;/h3&gt;

&lt;p&gt;What evidence exists that the important behavior works?&lt;/p&gt;

&lt;p&gt;Look at unit tests, integration tests, end to end tests, regression tests, and failure cases.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Production behavior
&lt;/h3&gt;

&lt;p&gt;How is the affected component actually used?&lt;/p&gt;

&lt;p&gt;Traffic, data volume, runtime behavior, and external integrations can all change the practical impact of a change.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Change and incident history
&lt;/h3&gt;

&lt;p&gt;Has this area caused incidents, rollbacks, or repeated regressions before?&lt;/p&gt;

&lt;p&gt;History does not determine whether a change is safe, but it is useful context.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Deployment and rollback
&lt;/h3&gt;

&lt;p&gt;What happens if the change behaves badly?&lt;/p&gt;

&lt;p&gt;Can it be rolled back quickly? Can it be deployed gradually? Is a feature flag available?&lt;/p&gt;

&lt;h2&gt;
  
  
  Production reliability vs code review
&lt;/h2&gt;

&lt;p&gt;Code review asks whether the implementation makes sense.&lt;/p&gt;

&lt;p&gt;Production reliability asks what the implementation means for the running system.&lt;/p&gt;

&lt;p&gt;You generally want both.&lt;/p&gt;

&lt;p&gt;Read more: &lt;a href="https://tomosu.ai/blogs/production-reliability-vs-code-review.html" rel="noopener noreferrer"&gt;https://tomosu.ai/blogs/production-reliability-vs-code-review.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Production reliability vs observability
&lt;/h2&gt;

&lt;p&gt;Observability tells you what is happening in a running system.&lt;/p&gt;

&lt;p&gt;Production reliability adds change context: what changed, what it affects, and what could happen when it reaches production.&lt;/p&gt;

&lt;p&gt;Observability is therefore one useful input into reliability assessment, not a replacement for it.&lt;/p&gt;

&lt;p&gt;Read more: &lt;a href="https://tomosu.ai/blogs/production-reliability-vs-observability.html" rel="noopener noreferrer"&gt;https://tomosu.ai/blogs/production-reliability-vs-observability.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical model
&lt;/h2&gt;

&lt;p&gt;A useful way to think about the assessment is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Change
  ↓
Evidence
  ↓
Uncertainty
  ↓
Action
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is not to eliminate uncertainty before every deployment.&lt;/p&gt;

&lt;p&gt;The goal is to find changes where the potential production impact and remaining uncertainty justify additional attention.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where PRI fits
&lt;/h2&gt;

&lt;p&gt;Tomosu's &lt;strong&gt;Production Reliability Index (PRI)&lt;/strong&gt; brings multiple reliability signals together, including:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fragility&lt;/li&gt;
&lt;li&gt;Drift&lt;/li&gt;
&lt;li&gt;Governance Compliance&lt;/li&gt;
&lt;li&gt;Runtime Signals&lt;/li&gt;
&lt;li&gt;Code Volatility&lt;/li&gt;
&lt;li&gt;Deployment Velocity&lt;/li&gt;
&lt;li&gt;Escalation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The number is useful as a summary, but the underlying evidence matters more than the number by itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main idea
&lt;/h2&gt;

&lt;p&gt;Production reliability is not a property of a diff alone.&lt;/p&gt;

&lt;p&gt;It is a property of a change in context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Change + dependencies + testing + production behavior + history + deployment conditions&lt;/strong&gt; give engineers a much better picture of what may happen after a change ships.&lt;/p&gt;

&lt;p&gt;Run a Production Reliability Index assessment:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://tomosu.ai/start" rel="noopener noreferrer"&gt;https://tomosu.ai/start&lt;/a&gt;&lt;/p&gt;

</description>
      <category>softwareengineering</category>
      <category>sre</category>
      <category>devops</category>
      <category>productivity</category>
    </item>
    <item>
      <title>What Is AI Observability? A Definition for Engineers</title>
      <dc:creator>Parsa Mohammadi</dc:creator>
      <pubDate>Wed, 23 Sep 2026 14:42:36 +0000</pubDate>
      <link>https://dev.to/parsa_m/what-is-ai-observability-a-definition-for-engineers-57j8</link>
      <guid>https://dev.to/parsa_m/what-is-ai-observability-a-definition-for-engineers-57j8</guid>
      <description>&lt;p&gt;AI observability is the practice of tracking what an AI system did and why.&lt;/p&gt;

&lt;p&gt;That means looking beyond uptime and request latency and capturing things such as prompts, model responses, retrieved context, tool calls, agent decisions, latency, token usage, and evaluation results.&lt;/p&gt;

&lt;p&gt;The distinction matters because an AI system can return a successful HTTP response and still produce the wrong result.&lt;/p&gt;

&lt;p&gt;A normal monitoring system might tell you that the request returned 200. AI observability helps answer what happened inside that request.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where standard monitoring stops
&lt;/h2&gt;

&lt;p&gt;Traditional monitoring is built around questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is the service up?&lt;/li&gt;
&lt;li&gt;Is latency increasing?&lt;/li&gt;
&lt;li&gt;Are requests failing?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those signals are still useful.&lt;/p&gt;

&lt;p&gt;But an AI application can fail while everything looks healthy from an infrastructure perspective.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Return an incorrect answer&lt;/li&gt;
&lt;li&gt;Retrieve the wrong context&lt;/li&gt;
&lt;li&gt;Call the wrong tool&lt;/li&gt;
&lt;li&gt;Produce an unsafe action&lt;/li&gt;
&lt;li&gt;Follow the wrong branch in an agent workflow&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The infrastructure worked. The behavior didn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an AI observability tool needs to capture
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Prompt and response pairs
&lt;/h3&gt;

&lt;p&gt;You need to know what went into the model and what came out.&lt;/p&gt;

&lt;h3&gt;
  
  
  Latency and cost
&lt;/h3&gt;

&lt;p&gt;Track model usage at the call and workflow level.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tool calls
&lt;/h3&gt;

&lt;p&gt;For agents, the sequence matters. You want to know what the agent called, in what order, what it received, and whether it retried.&lt;/p&gt;

&lt;h3&gt;
  
  
  Retrieved context
&lt;/h3&gt;

&lt;p&gt;For RAG systems, a bad response may come from bad retrieval rather than the model itself. Without the retrieved context, it's difficult to separate those failure modes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Evaluation results
&lt;/h3&gt;

&lt;p&gt;A system may behave differently over time even when the underlying infrastructure hasn't changed. Evaluation and observability together can show whether output quality is drifting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Monitoring vs AI observability
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Standard monitoring&lt;/th&gt;
&lt;th&gt;AI observability&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Uptime&lt;/td&gt;
&lt;td&gt;Model behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;Prompt and response&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Error rate&lt;/td&gt;
&lt;td&gt;Tool calls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Request failures&lt;/td&gt;
&lt;td&gt;Agent decision sequence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Infrastructure state&lt;/td&gt;
&lt;td&gt;Retrieved context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Is it running?&lt;/td&gt;
&lt;td&gt;What did it do?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What changes when an agent writes code?
&lt;/h2&gt;

&lt;p&gt;The problem becomes larger when the AI system isn't simply answering a user.&lt;/p&gt;

&lt;p&gt;A coding agent can plan a change, read files, run commands, modify code, run tests, interpret results, modify the code again, and produce a final diff.&lt;/p&gt;

&lt;p&gt;The final diff doesn't contain the whole decision path. If something breaks, you may need to know which tool call or intermediate decision caused the problem.&lt;/p&gt;

&lt;p&gt;That's why agent observability needs to capture the sequence, not just the final output.&lt;/p&gt;

&lt;h2&gt;
  
  
  What AI observability doesn't do
&lt;/h2&gt;

&lt;p&gt;Observability tells you what happened. It doesn't automatically stop a bad change from shipping.&lt;/p&gt;

&lt;p&gt;You can know that an agent hallucinated a configuration value after the fact and still have no mechanism that prevented the change from merging.&lt;/p&gt;

&lt;p&gt;That's a different problem.&lt;/p&gt;

&lt;p&gt;You need a separate verification or governance step that asks: Should this change actually ship?&lt;/p&gt;

&lt;p&gt;Observability helps explain failures. Production risk analysis is concerned with identifying risky changes before they become failures.&lt;/p&gt;

&lt;p&gt;Full article: &lt;a href="https://tomosu.ai/blogs/what-is-ai-observability.html" rel="noopener noreferrer"&gt;https://tomosu.ai/blogs/what-is-ai-observability.html&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>devops</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
