“It Worked” Is Not the Same as “It Can Run in Production”
You are probably the kind of reader I have in mind if:
- you already use ChatGPT, Claude, or another LLM;
- you have used GitHub or built a simple automation;
- AI agents and automation interest you, but governance or approval design is not your specialty;
- you are starting to ask, "How much of this should I let AI do automatically?"
That last question becomes important when you move from AI that suggests something to AI or automation that can actually do something next.
For example, suppose an AI agent drafts a technical article.
The Markdown is valid. Required fields are present. Automated checks return PASS.
Should the system publish it automatically?
Not necessarily.
A check can tell you that the output has the right format. It may also give you enough evidence to verify what was produced.
But that still does not answer two different questions:
- Has someone accepted the content?
- Is the system actually allowed to publish it?
That difference is the core of this article.
A simplified version looks like this:
AI drafts content
↓
Validation PASS
↓
Evidence available
↓
Human acceptance PENDING
↓
Publish DENIED
The same idea appears in CI/CD:
Tests PASS
↓
Release evidence available
↓
Approval PENDING
↓
Production deploy BLOCKED
In both cases, one step succeeded. But the next, higher-impact action is still not allowed.
Who this article is for
This article is for people who are moving beyond "ask an AI, read the answer, decide manually" and starting to connect AI or automation to actions such as:
- changing a repository;
- merging code;
- deploying software;
- sending a message;
- publishing content;
- updating another system.
If you are only using ChatGPT, Claude, or another LLM as a conversational assistant and you always decide the next step yourself, the model in this article is probably more than you need.
It becomes useful when a successful check can trigger another action automatically.
Assumptions
This model is useful when your workflow has some combination of:
- multiple stages rather than one isolated check;
- automated checks that return a local success state such as
PASS; - logs, test records, screenshots, or other evidence;
- human or policy review;
- a later action with a larger consequence than the check before it, such as merging, deploying, sending, publishing, or updating another system.
If the workflow is a single, side-effect-free check with no separate approval or execution decision, this model is probably unnecessary.
What you'll take away
The practical idea is to keep four questions separate:
Result
↓
Evidence
↓
Acceptance
↓
Production Execution
- Result: did the thing work?
- Evidence: can we show that it worked?
- Acceptance: has the result been accepted?
- Production Execution: is the next production-impacting action authorized?
This is not a universal four-gate architecture. It is a thinking aid for making the meaning of PASS explicit.
The rule I want to preserve is:
A PASS at one stage should not automatically become a PASS at the next stage.
Four questions that look similar but are not
The four states answer different questions.
1. Result — did the thing work?
This is the most immediate question.
Did the implementation run?
Did the expected output appear?
Did the test complete?
A positive result is important, but it only tells us that something happened as expected under the observed conditions.
It does not yet answer whether the result is sufficiently evidenced, formally accepted, or authorized for production.
2. Evidence — can we show that it worked?
A result can exist without durable evidence.
For example, a run may succeed, but the supporting logs, checks, screenshots, test records, or other evidence may be missing, incomplete, or not tied to the exact version that ran.
Evidence asks a different question:
Can another reviewer verify the claim we are making about the result?
That is a stronger state than "I saw it work."
3. Acceptance — has the result been formally accepted?
Even good evidence does not automatically create formal acceptance.
Acceptance may depend on whatever review rules apply to the system: quality criteria, security checks, policy requirements, scope limits, rollback readiness, or another explicit decision.
The concrete gate varies by environment.
The reusable point is simply that evidence and acceptance are different states.
4. Production Execution — is it authorized to run in production?
This is another separate decision.
A system may be implemented, evidenced, and even accepted as a valid artifact without being authorized for production execution yet.
That distinction is not unusual in deployment tooling.
For example, GitHub Actions environments can require reviewers or other deployment protection rules before a job targeting an environment is allowed to proceed. Google Cloud Deploy can require approval on deployment targets before promotion. AWS CodePipeline can stop a pipeline at a manual approval action until an authorized approver allows it to continue.
Those platforms do not prove that every system should use the four states above. They are useful examples of a broader idea: upstream technical success and downstream execution authorization can be modeled separately.
What I observed across more than one context
The pattern became more useful when it appeared in more than one workflow.
In one context, a technical or artifact-level step could pass while the overall promotion decision remained unresolved.
In another, an initial phase could pass while the broader artifact still required targeted rework, a later phase remained on hold, and final approval had not been granted.
The exact states were system-specific.
The reusable lesson was not their names. It was that they were not collapsed into one boolean.
That made it possible to say:
- this part worked;
- the evidence for this part is sufficient;
- the overall artifact is not yet accepted;
- production execution is still not authorized.
That is a more precise description than a single "PASS."
Why state separation helps
Separating the decisions can prevent several kinds of error.
Premature promotion
A successful implementation can look finished before evidence, review, or runtime conditions are ready.
Ambiguous status
If one field called PASS is used for several different meanings, people and automation may interpret it differently.
Unsafe automation
Automation is especially sensitive to ambiguous state.
If a downstream action sees "PASS" and cannot tell whether that means "test passed" or "production authorized," the system can move farther than intended.
Weak auditability
Separate states make it easier to reconstruct which decision was made, by which process, and what remained pending.
The four-gate model is a thinking aid, not a universal architecture
I would not turn the diagram into a rule that every system must implement exactly four gates.
Some systems may combine stages safely.
Others may need more states.
Some may use automated policy checks instead of human acceptance.
Some may not have a production environment at all.
The useful design question is:
Which decisions in this workflow are genuinely different, and which ones are we accidentally allowing to inherit each other's PASS state?
For my use case, separating Result, Evidence, Acceptance, and Production Execution made the state model clearer.
A practical rule I plan to reuse
When designing a workflow that can change production state, I now try to avoid rules like:
if result == PASS:
promote()
I prefer something conceptually closer to:
result_passed
evidence_passed
acceptance_passed
production_execution_authorized
The exact implementation can vary.
What matters is that one variable does not silently stand in for all four decisions.
References
These references are included as general context on deployment approval and protection mechanisms. They do not establish that the four-state model in this article is a universal best practice.
- GitHub Docs — Deployments and environments: https://docs.github.com/en/actions/reference/workflows-and-actions/deployments-and-environments
- GitHub Docs — Controlling deployments with environments and protection rules: https://docs.github.com/en/actions/how-tos/deploy/configure-and-manage-deployments/control-deployments
- Google Cloud Deploy — Promote releases and manage approvals: https://docs.cloud.google.com/deploy/docs/promote-release
- AWS CodePipeline — Manual approval actions: https://docs.aws.amazon.com/codepipeline/latest/userguide/approvals.html
Scope note: the observations behind this article come from more than one context, but they do not establish a universal release architecture. Treat the model as a reusable design principle to test against your own workflow.
Top comments (2)
This maps almost one-to-one onto how we gate an autonomous pipeline, and the sentence that earns the whole post is "a PASS at one stage should not automatically become a PASS at the next." Silent promotion between stages is precisely how a cheated eval reaches production — the result passed, so nobody re-asks whether the evidence actually supports the claim.
Your Evidence stage line about logs "tied to the exact version that ran" is the one I'd underline hardest. We learned the hard way that evidence not bound to the exact artifact is almost worse than no evidence, because it looks authoritative. A run succeeds, the screenshot is from a slightly different build, and now the Acceptance reviewer is approving a claim about code that never ran. Binding the evidence to a content hash of what executed closes that gap cheaply.
One question: who owns the Acceptance → Production transition in your model when it's an agent doing the work? That's the gate we keep wanting a human on, and the one most teams quietly automate away.
We run a site where an agent ships a post a day, and this is the exact line we had to draw: the agent publishes content unattended, but a named list of decisions — UI deploys, anything that sends mail, anything that moves money — stays a hard human gate no threshold negotiates. The interesting question your framing exposes: format-PASS is a probability, not a verdict. Did you land on per-decision thresholds, or a flat approval list?