<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: AI Works</title>
    <description>The latest articles on DEV Community by AI Works (@_797a7c3a31b7c8547037).</description>
    <link>https://dev.to/_797a7c3a31b7c8547037</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4156871%2Fc57722e1-85a0-4911-a3ff-23f8c2b219b5.png</url>
      <title>DEV Community: AI Works</title>
      <link>https://dev.to/_797a7c3a31b7c8547037</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/_797a7c3a31b7c8547037"/>
    <language>en</language>
    <item>
      <title>“It Worked” Is Not the Same as “It Can Run in Production”</title>
      <dc:creator>AI Works</dc:creator>
      <pubDate>Wed, 07 Oct 2026 16:02:58 +0000</pubDate>
      <link>https://dev.to/_797a7c3a31b7c8547037/it-worked-is-not-the-same-as-it-can-run-in-production-198j</link>
      <guid>https://dev.to/_797a7c3a31b7c8547037/it-worked-is-not-the-same-as-it-can-run-in-production-198j</guid>
      <description>&lt;h1&gt;
  
  
  “It Worked” Is Not the Same as “It Can Run in Production”
&lt;/h1&gt;

&lt;p&gt;You are probably the kind of reader I have in mind if:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;you already use ChatGPT, Claude, or another LLM;&lt;/li&gt;
&lt;li&gt;you have used GitHub or built a simple automation;&lt;/li&gt;
&lt;li&gt;AI agents and automation interest you, but governance or approval design is not your specialty;&lt;/li&gt;
&lt;li&gt;you are starting to ask, &lt;strong&gt;"How much of this should I let AI do automatically?"&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last question becomes important when you move from &lt;strong&gt;AI that suggests something&lt;/strong&gt; to &lt;strong&gt;AI or automation that can actually do something next&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For example, suppose an AI agent drafts a technical article.&lt;/p&gt;

&lt;p&gt;The Markdown is valid. Required fields are present. Automated checks return &lt;code&gt;PASS&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Should the system publish it automatically?&lt;/p&gt;

&lt;p&gt;Not necessarily.&lt;/p&gt;

&lt;p&gt;A check can tell you that the output has the right format. It may also give you enough evidence to verify what was produced.&lt;/p&gt;

&lt;p&gt;But that still does not answer two different questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Has someone accepted the content?&lt;/li&gt;
&lt;li&gt;Is the system actually allowed to publish it?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That difference is the core of this article.&lt;/p&gt;

&lt;p&gt;A simplified version looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI drafts content
    ↓
Validation PASS
    ↓
Evidence available
    ↓
Human acceptance PENDING
    ↓
Publish DENIED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same idea appears in CI/CD:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Tests PASS
    ↓
Release evidence available
    ↓
Approval PENDING
    ↓
Production deploy BLOCKED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In both cases, one step succeeded. But the next, higher-impact action is still not allowed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this article is for
&lt;/h2&gt;

&lt;p&gt;This article is for people who are moving beyond "ask an AI, read the answer, decide manually" and starting to connect AI or automation to actions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;changing a repository;&lt;/li&gt;
&lt;li&gt;merging code;&lt;/li&gt;
&lt;li&gt;deploying software;&lt;/li&gt;
&lt;li&gt;sending a message;&lt;/li&gt;
&lt;li&gt;publishing content;&lt;/li&gt;
&lt;li&gt;updating another system.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are only using ChatGPT, Claude, or another LLM as a conversational assistant and you always decide the next step yourself, the model in this article is probably more than you need.&lt;/p&gt;

&lt;p&gt;It becomes useful when &lt;strong&gt;a successful check can trigger another action automatically&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Assumptions
&lt;/h2&gt;

&lt;p&gt;This model is useful when your workflow has some combination of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;multiple stages rather than one isolated check;&lt;/li&gt;
&lt;li&gt;automated checks that return a local success state such as &lt;code&gt;PASS&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;logs, test records, screenshots, or other evidence;&lt;/li&gt;
&lt;li&gt;human or policy review;&lt;/li&gt;
&lt;li&gt;a later action with a larger consequence than the check before it, such as merging, deploying, sending, publishing, or updating another system.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the workflow is a single, side-effect-free check with no separate approval or execution decision, this model is probably unnecessary.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you'll take away
&lt;/h2&gt;

&lt;p&gt;The practical idea is to keep four questions separate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Result
  ↓
Evidence
  ↓
Acceptance
  ↓
Production Execution
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Result:&lt;/strong&gt; did the thing work?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evidence:&lt;/strong&gt; can we show that it worked?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Acceptance:&lt;/strong&gt; has the result been accepted?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production Execution:&lt;/strong&gt; is the next production-impacting action authorized?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is not a universal four-gate architecture. It is a thinking aid for making the meaning of &lt;code&gt;PASS&lt;/code&gt; explicit.&lt;/p&gt;

&lt;p&gt;The rule I want to preserve is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A PASS at one stage should not automatically become a PASS at the next stage.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Four questions that look similar but are not
&lt;/h2&gt;

&lt;p&gt;The four states answer different questions.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Result — did the thing work?
&lt;/h3&gt;

&lt;p&gt;This is the most immediate question.&lt;/p&gt;

&lt;p&gt;Did the implementation run?&lt;br&gt;
Did the expected output appear?&lt;br&gt;
Did the test complete?&lt;/p&gt;

&lt;p&gt;A positive result is important, but it only tells us that something happened as expected under the observed conditions.&lt;/p&gt;

&lt;p&gt;It does not yet answer whether the result is sufficiently evidenced, formally accepted, or authorized for production.&lt;/p&gt;
&lt;h3&gt;
  
  
  2. Evidence — can we show that it worked?
&lt;/h3&gt;

&lt;p&gt;A result can exist without durable evidence.&lt;/p&gt;

&lt;p&gt;For example, a run may succeed, but the supporting logs, checks, screenshots, test records, or other evidence may be missing, incomplete, or not tied to the exact version that ran.&lt;/p&gt;

&lt;p&gt;Evidence asks a different question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can another reviewer verify the claim we are making about the result?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a stronger state than "I saw it work."&lt;/p&gt;
&lt;h3&gt;
  
  
  3. Acceptance — has the result been formally accepted?
&lt;/h3&gt;

&lt;p&gt;Even good evidence does not automatically create formal acceptance.&lt;/p&gt;

&lt;p&gt;Acceptance may depend on whatever review rules apply to the system: quality criteria, security checks, policy requirements, scope limits, rollback readiness, or another explicit decision.&lt;/p&gt;

&lt;p&gt;The concrete gate varies by environment.&lt;/p&gt;

&lt;p&gt;The reusable point is simply that &lt;strong&gt;evidence and acceptance are different states&lt;/strong&gt;.&lt;/p&gt;
&lt;h3&gt;
  
  
  4. Production Execution — is it authorized to run in production?
&lt;/h3&gt;

&lt;p&gt;This is another separate decision.&lt;/p&gt;

&lt;p&gt;A system may be implemented, evidenced, and even accepted as a valid artifact without being authorized for production execution yet.&lt;/p&gt;

&lt;p&gt;That distinction is not unusual in deployment tooling.&lt;/p&gt;

&lt;p&gt;For example, GitHub Actions environments can require reviewers or other deployment protection rules before a job targeting an environment is allowed to proceed. Google Cloud Deploy can require approval on deployment targets before promotion. AWS CodePipeline can stop a pipeline at a manual approval action until an authorized approver allows it to continue.&lt;/p&gt;

&lt;p&gt;Those platforms do not prove that every system should use the four states above. They are useful examples of a broader idea: &lt;strong&gt;upstream technical success and downstream execution authorization can be modeled separately&lt;/strong&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I observed across more than one context
&lt;/h2&gt;

&lt;p&gt;The pattern became more useful when it appeared in more than one workflow.&lt;/p&gt;

&lt;p&gt;In one context, a technical or artifact-level step could pass while the overall promotion decision remained unresolved.&lt;/p&gt;

&lt;p&gt;In another, an initial phase could pass while the broader artifact still required targeted rework, a later phase remained on hold, and final approval had not been granted.&lt;/p&gt;

&lt;p&gt;The exact states were system-specific.&lt;/p&gt;

&lt;p&gt;The reusable lesson was not their names. It was that they were &lt;strong&gt;not collapsed into one boolean&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That made it possible to say:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;this part worked;&lt;/li&gt;
&lt;li&gt;the evidence for this part is sufficient;&lt;/li&gt;
&lt;li&gt;the overall artifact is not yet accepted;&lt;/li&gt;
&lt;li&gt;production execution is still not authorized.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is a more precise description than a single "PASS."&lt;/p&gt;
&lt;h2&gt;
  
  
  Why state separation helps
&lt;/h2&gt;

&lt;p&gt;Separating the decisions can prevent several kinds of error.&lt;/p&gt;
&lt;h3&gt;
  
  
  Premature promotion
&lt;/h3&gt;

&lt;p&gt;A successful implementation can look finished before evidence, review, or runtime conditions are ready.&lt;/p&gt;
&lt;h3&gt;
  
  
  Ambiguous status
&lt;/h3&gt;

&lt;p&gt;If one field called &lt;code&gt;PASS&lt;/code&gt; is used for several different meanings, people and automation may interpret it differently.&lt;/p&gt;
&lt;h3&gt;
  
  
  Unsafe automation
&lt;/h3&gt;

&lt;p&gt;Automation is especially sensitive to ambiguous state.&lt;/p&gt;

&lt;p&gt;If a downstream action sees "PASS" and cannot tell whether that means "test passed" or "production authorized," the system can move farther than intended.&lt;/p&gt;
&lt;h3&gt;
  
  
  Weak auditability
&lt;/h3&gt;

&lt;p&gt;Separate states make it easier to reconstruct which decision was made, by which process, and what remained pending.&lt;/p&gt;
&lt;h2&gt;
  
  
  The four-gate model is a thinking aid, not a universal architecture
&lt;/h2&gt;

&lt;p&gt;I would not turn the diagram into a rule that every system must implement exactly four gates.&lt;/p&gt;

&lt;p&gt;Some systems may combine stages safely.&lt;br&gt;
Others may need more states.&lt;br&gt;
Some may use automated policy checks instead of human acceptance.&lt;br&gt;
Some may not have a production environment at all.&lt;/p&gt;

&lt;p&gt;The useful design question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Which decisions in this workflow are genuinely different, and which ones are we accidentally allowing to inherit each other's PASS state?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For my use case, separating Result, Evidence, Acceptance, and Production Execution made the state model clearer.&lt;/p&gt;
&lt;h2&gt;
  
  
  A practical rule I plan to reuse
&lt;/h2&gt;

&lt;p&gt;When designing a workflow that can change production state, I now try to avoid rules like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;if result == PASS:
    promote()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I prefer something conceptually closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;result_passed
evidence_passed
acceptance_passed
production_execution_authorized
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact implementation can vary.&lt;/p&gt;

&lt;p&gt;What matters is that one variable does not silently stand in for all four decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;These references are included as general context on deployment approval and protection mechanisms. They do not establish that the four-state model in this article is a universal best practice.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GitHub Docs — Deployments and environments:
&lt;a href="https://docs.github.com/en/actions/reference/workflows-and-actions/deployments-and-environments" rel="noopener noreferrer"&gt;https://docs.github.com/en/actions/reference/workflows-and-actions/deployments-and-environments&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;GitHub Docs — Controlling deployments with environments and protection rules:
&lt;a href="https://docs.github.com/en/actions/how-tos/deploy/configure-and-manage-deployments/control-deployments" rel="noopener noreferrer"&gt;https://docs.github.com/en/actions/how-tos/deploy/configure-and-manage-deployments/control-deployments&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Google Cloud Deploy — Promote releases and manage approvals:
&lt;a href="https://docs.cloud.google.com/deploy/docs/promote-release" rel="noopener noreferrer"&gt;https://docs.cloud.google.com/deploy/docs/promote-release&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;AWS CodePipeline — Manual approval actions:
&lt;a href="https://docs.aws.amazon.com/codepipeline/latest/userguide/approvals.html" rel="noopener noreferrer"&gt;https://docs.aws.amazon.com/codepipeline/latest/userguide/approvals.html&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;Scope note: the observations behind this article come from more than one context, but they do not establish a universal release architecture. Treat the model as a reusable design principle to test against your own workflow.&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>architecture</category>
      <category>automation</category>
    </item>
  </channel>
</rss>
