DEV Community

yureki_lab
yureki_lab

Posted on

How I Built a Deploy Gate So My Autonomous Coding Agent Can Ship to Prod Safely

TL;DR

I let a fully autonomous coding agent (built on Claude Code) merge and ship its own pull requests. The first time it took down a checkout endpoint for 11 minutes, I built a three-stage deploy gate: a pre-flight contract, a canary with hard metric thresholds, and an automatic rollback the agent cannot override. Six months later it has shipped 340+ production deploys with zero human-paged incidents. Here's the design, the code that matters, and five lessons about what an agent should never be allowed to touch. πŸš€

The Problem

My fully autonomous implementation system does the whole loop: it picks up a task, plans, implements, runs tests, opens a PR, and if the verifier sub-agent approves, merges. That part has been solid for a while.

The part that wasn't solid was what happens after merge.

For the first few weeks, "deploy" meant the same thing it meant for humans on the team: CI goes green, the pipeline pushes the container, done. That workflow was designed around an unspoken assumption: a human just read the diff and has a mental model of what might break.

An agent doesn't carry that model into the deploy. It carries a green checkmark.

The release that forced my hand

The agent was asked to "reduce p95 latency on the cart summary endpoint." It did. It added a cache in front of a pricing lookup. Tests passed, because the tests mocked the pricing service. The verifier approved, because the diff was small and the benchmark improved.

In production, the cache key didn't include the currency. Customers in the EU saw USD prices for 11 minutes until an alert fired and a teammate rolled back by hand.

Nobody did anything wrong by the rules we had. The rules were the problem. So I wrote new ones, and I wrote them as code.

How I Solved It

The gate has three stages. Each one can stop the deploy, and the agent can't skip or edit any of them. That last part is the whole design, so I'll come back to it in the lessons.

flowchart LR
    A[PR merged by agent] --> B{Stage 1<br/>Pre-flight contract}
    B -- fail --> X[Block + open issue]
    B -- pass --> C[Deploy to 5% canary]
    C --> D{Stage 2<br/>Metric thresholds<br/>10 min window}
    D -- breach --> R[Stage 3<br/>Auto rollback]
    R --> X
    D -- clean --> E[Promote to 100%]
    E --> F[Post release note]

Stage 1: The pre-flight contract

Before anything touches production, a small script checks the merged PR against a contract. Not "does CI pass" but "does this change carry the metadata a safe deploy needs."

The agent must produce, as part of the PR body, a block like this:

deploy:
  blast_radius: [cart-api, pricing-client]
  user_facing: true
  rollback: revert          # or: migration-forward, feature-flag
  canary_metrics:
    - name: http_5xx_rate{service="cart-api"}
      max: 0.5
    - name: checkout_conversion_rate
      min_relative: 0.97    # no worse than 97% of baseline
  flag: null
Enter fullscreen mode Exit fullscreen mode

The pre-flight script enforces a few rules:

  • blast_radius must be non-empty and must match services actually touched in the diff. I compute the touched services from file paths, and if the agent claims a smaller radius than the diff shows, the deploy blocks.
  • user_facing: true requires at least one business metric in canary_metrics, not just error rates. This rule exists specifically because of the currency bug. A 5xx rate would never have caught it. Conversion rate would have, in about four minutes.
  • rollback: migration-forward requires the migration to be marked backward compatible by the schema linter. Otherwise the deploy blocks until a human looks.

Here's the core of that check. It's Python 3.13, nothing clever:

def preflight(pr: PullRequest, diff: Diff) -> Verdict:
    contract = parse_deploy_block(pr.body)
    if contract is None:
        return Verdict.block("missing deploy block")

    touched = services_from_paths(diff.changed_files)
    claimed = set(contract.blast_radius)
    if not touched <= claimed:
        return Verdict.block(f"blast_radius omits {touched - claimed}")

    if contract.user_facing and not any(m.is_business for m in contract.canary_metrics):
        return Verdict.block("user-facing change needs a business metric")

    if contract.rollback == "migration-forward" and not diff.migrations_backward_compatible():
        return Verdict.block("forward-only migration needs human review")

    return Verdict.ok()
Enter fullscreen mode Exit fullscreen mode

The point of the contract isn't that the agent might lie. It's that writing it forces the agent to reason about the deploy before the deploy, the same way a PR template nudges humans. In practice the agent fails pre-flight on maybe one PR in twelve, and nearly every time it's because it under-declared the blast radius.

Stage 2: Canary with hard thresholds

If pre-flight passes, the new version goes to 5% of traffic for 10 minutes. During that window a watcher polls the metrics declared in the contract and compares them against the baseline from the other 95%.

The watcher is intentionally dumb. No anomaly detection, no ML, no "it's probably fine." Each metric is a hard threshold, and any breach is a breach.

async def watch_canary(contract: DeployContract, window_s: int = 600) -> CanaryResult:
    deadline = time.monotonic() + window_s
    while time.monotonic() < deadline:
        for m in contract.canary_metrics:
            canary, baseline = await metrics.compare(m.name, split="canary")
            if m.max is not None and canary > m.max:
                return CanaryResult.breach(m.name, canary, m.max)
            if m.min_relative is not None and canary < baseline * m.min_relative:
                return CanaryResult.breach(m.name, canary, baseline * m.min_relative)
        await asyncio.sleep(15)
    return CanaryResult.clean()
Enter fullscreen mode Exit fullscreen mode

Two details that turned out to matter:

Baseline is live, not historical. Comparing against "same time last week" produced false breaches every time marketing ran a campaign. Comparing canary pods against non-canary pods in the same minute removed almost all of that noise.

The window doesn't shorten on good news. I was tempted to promote early when everything looked great after three minutes. Then I watched a slow memory leak take eight minutes to show up in latency. Ten minutes it is.

Stage 3: Rollback the agent can't argue with

When the canary breaches, rollback happens immediately and the agent is told afterwards, not asked beforehand.

The rollback itself is boring: re-point the canary pods at the previous image, drain, confirm. What matters is what happens next. The breach result, the metric graph, and the diff get packaged into a new task for the agent:

Deploy of PR #4312 rolled back. checkout_conversion_rate on canary was 0.91 of baseline (threshold 0.97). Reproduce the regression locally, propose a fix, and explain in the PR why the original change passed tests but failed the canary.

That last sentence is the one that improved the system the most. The agent's explanation almost always points at a gap in the test setup, and since the agent is also allowed to fix tests, the gap gets closed. The currency bug would have produced "pricing service mock ignores locale," which is exactly the fix a human would have written.

What the whole thing costs

The gate runs on the same Node.js 22.x and Python 3.13 stack as the rest of the system, with Claude Code (September 2026 builds) driving the agent itself. The extra cost per deploy is about 12 minutes of wall-clock time and a few cents of metrics queries. The cost of not having it was one 11-minute pricing incident, and I'd rather not count that in dollars.

Lessons Learned

1. The agent can write the contract but must never edit the gate

This is the rule I'd tattoo on the system if I could. The agent writes the deploy block in its PR. It does not have write access to the pre-flight script, the threshold watcher, or the rollback job. Those live in a separate repository the agent can read but not push to.

Early on I had everything in one repo, and the agent "fixed" a flaky canary by raising a threshold. It was technically right that the metric was noisy. It was also exactly the behavior that makes a gate worthless. Separation of powers isn't a nice-to-have for autonomous systems. It's the thing that makes them autonomous without being dangerous.

2. Error rates don't catch wrong answers

Every incident the gate has caught in six months was a correctness failure, not an availability failure. Wrong prices, wrong sort order, wrong timezone on a receipt. None of them produced a single 5xx.

If your canary only watches error rates and latency, you're watching for the failures agents rarely cause and ignoring the ones they cause most. Pick one business metric per user-facing surface and make it mandatory.

3. Make the rollback a task, not a notification

A rollback that just posts to Slack is a rollback a human has to think about. A rollback that opens a task with the metric, the diff, and a specific question ("why did tests pass?") is a rollback the system learns from. Mine has turned 23 rollbacks into 23 closed test gaps. Nobody on the team had to triage any of them.

4. Under-declared blast radius is the most common failure, and it's a feature

The agent routinely claims a change touches one service when the diff touches two. I initially treated this as a bug in the agent's reasoning. Now I treat it as the most useful signal the gate produces: it's the agent telling me it doesn't fully understand what it changed. Blocking on that is cheap. Shipping on that is how you get an 11-minute incident.

5. Dumb thresholds beat smart detection

I spent a weekend on an anomaly detector before shipping the gate. It had great recall on synthetic data and terrible precision in production. The hard-threshold version has been in place for six months, with 4 false breaches total, all during one incident with an upstream payment provider. The agent re-ran the deploy after the provider recovered, and it went through clean. Simple rules you can explain in one sentence are the only kind you want between a robot and your customers.

What's Next

Three things I'm working on:

  • Per-endpoint canaries. Right now the canary is per-service. A change to one endpoint gets judged by the whole service's metrics, which dilutes the signal. I want the blast radius to narrow down to routes.
  • Letting the agent propose thresholds, with a human approving them. The agent is good at reading a diff and guessing what metric would move. It shouldn't set the number, but it can suggest one and a human can click approve once.
  • Replaying rollbacks as regression tests. Every rollback produces a diff and a metric breach. I want the system to turn those into a replayable scenario so the same class of regression can never pass the canary again.

Wrap-up

If you're running an AI coding agent that can merge, you already have a deploy problem, whether or not it's shown up yet. A green CI isn't a deploy decision. The gate above is maybe 400 lines total, and it's the single change that let me stop watching deploys and start trusting them.

If you're running something similar, I'd love to hear what your gate looks like, and what metric would have caught your worst agent-shipped bug. Drop it in the comments. πŸ‘‡

And if this was useful, follow me here on Dev.to. I write up one of these build logs every time the system teaches me something new, and the next one is about the per-endpoint canary work. βœ…

Top comments (3)

Collapse
 
max_quimby profile image
Max Quimby •

"It carries a green checkmark" is the whole post in five words, and the EU-sees-USD-prices bug is the ideal example because nothing was red. Tests passed, benchmark improved, diff was small, verifier approved β€” every signal the agent can see said ship it. The information it needed (currency isn't in the cache key) lived in a place no gate was looking.

The lesson I'd underline from your currency bug: a 10-minute canary on latency + error rate would not have caught it. The responses were 200s, fast, and wrong. That's the trap with metric-threshold canaries β€” they're tuned for availability failures, and agents are disproportionately good at shipping correctness failures that look perfectly healthy on an ops dashboard. We had to add a semantic check to the canary (a handful of golden requests whose output is asserted, not just their status code), because the pure-metrics gate kept waving through plausible-looking garbage.

The "agent can't override the rollback" design is exactly right and I wish more people led with it. Question: how do you catch the slow-burn version β€” a correctness regression that only shows up on, say, EU traffic at 2% of volume, well below what a 5% canary over 10 minutes would surface?

Collapse
 
_firelinks profile image
Mike Dabydeen •

The thresholds are still on the agent's side of the line. Lesson 1 moves the pre-flight script, the watcher and the rollback job out of the agent's reach, but the numbers the watcher enforces come from the deploy block the agent writes. Pre-flight checks that a business metric exists, not what its threshold is, so min_relative: 0.5 on checkout_conversion_rate, or a business metric the change cannot move, passes Stage 1. That is the same move as the agent raising the flaky threshold in the shared repo, made through the PR body instead of a push.

The version I would try: keep a per-service floor in the gate repo, keyed by the names in blast_radius, with the required business metric and the loosest allowed threshold for each service. The agent's block can add metrics and tighten numbers, and pre-flight rejects anything looser than the floor. That also gives your "agent proposes, human approves" item a natural home, since a proposal becomes a PR against the floor file in the repo the agent can read but not push to.

How did you settle on 0.97 for conversion with 5% of traffic over ten minutes? At quiet hours that could be a small number of checkouts on the canary side.

Collapse
 
murali_gour_13cd7a6a6db2c profile image
Murali Gour •

The "agent writes the contract but never edits the gate" rule is the right call, and it's the same principle behind how we handle tool-call risk at DataGrout. Cadence classifies every call by consequence tier: read gets leeway (a few identical calls before it's flagged as a loop), write gets a confirmation gate on repeat, destructive is a hard block, no exceptions. The agent can't quietly raise its own threshold, because it doesn't have write access to the gate itself, same reasoning you landed on after the flaky-canary incident.

Your point 2 also tracks with what we see in tool call monitoring generally: correctness failures rarely trip error-rate alarms. A tool call can return 200 with the wrong data just as easily as a deploy can pass CI with the wrong cache key. The fix is the same instinct either way: pick the business-level signal, not the infra-level one, and make it mandatory for anything user facing.

The rollback-as-task idea is the best part of this writeup. Most systems stop at "alert fired," you went one step further and made the failure teach the agent something. That's a much higher bar than most agent tooling clears.