DEV Community

Oleksandr Kuryzhev
Oleksandr Kuryzhev

Posted on Originally published at kuryzhev.cloud

CI/CD Quality Gate Mistakes: Approvals and Rollback Pitfalls

Originally published on kuryzhev.cloud


A team adds a CI/CD quality gate after a bad release, bolts on a manual approval step for production, and writes "rollback: redeploy the previous version" in the runbook. Everyone feels safer. The pipeline is green, the approval button exists, and the rollback sentence is documented. This retrospective looks at how those three controls typically fail anyway. It draws on documented platform behavior and common design patterns rather than any specific incident.

Context: three controls, three different jobs

The three controls answer different questions, and mixing them up is where trouble starts. A quality gate answers "is this artifact technically acceptable?" It is automated, repeatable and should give the same verdict twice. A manual approval answers "is it acceptable to release this now?" That is a human judgment about timing, risk and business context. A rollback answers "how do we return to a known-good state when the first two were wrong?"

Each control is cheap to add as a checkbox and expensive to make meaningful. A gate that never fails, an approval nobody reads and a rollback nobody has run all look like working controls until the day they are needed. A common pattern is that the control exists in configuration but does not constrain anything.

The sections below cover the three most common ways these controls degrade, followed by a pattern that holds up better. The examples use GitHub Actions and Kubernetes because their behavior is well documented. The same ideas apply to GitLab CI, Jenkins or Argo CD.

Common failure 1: the gate that can be skipped, retried or ignored

The most frequent weakness is a CI/CD quality gate that is not actually required. A test job runs and fails, but the deploy job still starts. The cause is usually continue-on-error: true, a misconfigured needs, or a deploy workflow triggered separately from the test workflow. The dashboard shows red, but nothing stopped the release.

A second variant is the gate that measures the wrong thing. Coverage thresholds computed over the whole repository barely move when a risky change lands, so the gate passes. A better approach is to gate on new or changed code, plus a small set of meaningful checks: unit tests, a vulnerability scan with an explicit severity threshold, and a policy check on infrastructure changes. That is more honest than one large aggregate number.

A third variant is flakiness. When a test fails one run in ten, people rerun until green, and the gate trains the team to treat red as noise. Quarantine flaky tests visibly and track them, rather than letting reruns become the workflow.

Watch out for path filters on required checks. In GitHub, a workflow skipped because of path or branch filtering leaves its required check in a "Pending" state, which blocks the merge. Teams often "fix" this by removing the requirement. Check the current behavior in the GitHub Docs on required status checks, and design a lightweight always-run job instead.

Also verify that the gate runs against the exact commit being deployed. A gate on the pull request branch says little about the merge result on main if main has moved.

Common failure 2: the manual approval that approves nothing

Manual approval is usually added as a reaction to risk, and then it drifts into ceremony. The approver gets a notification saying "Deployment to production is waiting." It has no diff, no test summary, no link to the change and no indication of what differs from what is running now. Clicking "approve" becomes the path of least resistance, and approval fatigue sets in over time.

The deeper problem is approving a pipeline rather than an artifact. The approval may happen first with the build running afterward, or the deploy job may rebuild from source. Either way, the thing that was reviewed is not the thing that ships. A moving base image tag or an unpinned dependency can change the result between approval and deploy.

Watch out for self-approval and long waits. If the person who triggered the deployment can also approve it, the control is decorative. GitHub environments offer a "prevent self-review" option, and required reviewers are configured per environment. Waiting approvals also time out, and the exact limit is platform-dependent. Verify it in the docs before relying on a deploy that waits over a weekend.

Approvals also tend to be applied uniformly. Gating a documentation change and a database schema change identically teaches people that the gate carries no information. Reserve human approval for changes that are genuinely hard to reverse or time-sensitive. Let low-risk changes flow through the automated gate alone.

Common failure 3: rollback that was never exercised

"Roll back" often means "redeploy the previous commit." That is a rebuild, not a rollback. It depends on the build system working, registries being reachable and dependencies resolving the same way. It also depends on the quality gate (or an emergency bypass of it) running under pressure. A true rollback redeploys the previous immutable artifact, ideally by digest, without rebuilding anything.

Even then, rollback only reverts what the tool tracks. A Kubernetes Deployment rollback restores the previous pod template from the revision history. It does not undo changes to ConfigMaps, Secrets, CRDs or external resources. The retained history is bounded by revisionHistoryLimit, so check the value your manifests actually set. The Kubernetes Deployment documentation describes what a rollback does and does not revert.

The hardest case is state. Suppose release N+1 ran a destructive migration, such as dropping a column or rewriting data. Rolling the application back to N may then leave it running against a schema it does not understand. This is often discovered only during an attempted rollback.

Watch out for rollback paths that need the thing that just broke. The rollback may require the same pipeline, the same approver and the same registry. If so, a single outage can block both the release and the recovery. Rehearse the rollback in a non-production environment on a schedule, so the first real attempt is not the first attempt.

Safer operating pattern: build once, gate the artifact, verify after deploy

A more reliable CI/CD quality gate design follows a few principles:

  • Gate the exact commit that produces the artifact, and optionally scan the built image by digest.
  • Build the artifact once and identify it by digest.
  • Attach the approval to the environment, not to a workflow step.
  • Verify health after deploy, and undo automatically when verification fails.

The sketch below shows the shape in GitHub Actions; adapt names and secrets to your setup.

This workflow gates the commit, builds one image from that tree, deploys that exact digest to a protected environment, and reverts if rollout verification fails:

name: release
on:
  push:
    branches: [main]

jobs:
  build-and-gate:
    runs-on: ubuntu-24.04
    outputs:
      digest: ${{ steps.push.outputs.digest }}   # immutable reference reused downstream
    steps:
      - uses: actions/checkout@v5                # pin to a commit SHA in real pipelines
      - name: Unit tests and policy checks
        run: make test lint                      # no continue-on-error: failure stops the run
      - name: Set up Buildx
        uses: docker/setup-buildx-action@v3
      - name: Log in to registry
        uses: docker/login-action@v3
        with:
          registry: registry.example.com
          username: ${{ secrets.REGISTRY_USERNAME }}   # placeholder secret names
          password: ${{ secrets.REGISTRY_TOKEN }}
      - name: Build and push image
        id: push
        uses: docker/build-push-action@v6
        with:
          context: .                             # build the tree that was just tested
          push: true
          tags: registry.example.com/app:${{ github.sha }}

  deploy-prod:
    needs: build-and-gate
    runs-on: ubuntu-24.04
    environment: production                      # required reviewers + prevent self-review live here
    steps:
      - name: Configure cluster access
        env:
          KUBECONFIG_DATA: ${{ secrets.KUBECONFIG_PROD }}  # placeholder; prefer OIDC short-lived creds
        run: |
          mkdir -p ~/.kube
          printf '%s' "$KUBECONFIG_DATA" > ~/.kube/config
      - name: Deploy by digest
        run: |
          kubectl set image deployment/app \
            app=registry.example.com/app@${{ needs.build-and-gate.outputs.digest }}
      - name: Verify rollout
        id: verify
        run: kubectl rollout status deployment/app --timeout=180s
      - name: Roll back on failed verification
        if: failure() && steps.verify.outcome == 'failure'   # only undo a rollout that actually started
        run: kubectl rollout undo deployment/app

Environment protection rules, including required reviewers, are configured in repository settings rather than in the YAML. See the GitHub Docs on deployment environments for current options and plan limitations.

The rollback condition is deliberately narrow. If the deploy step itself fails, no new revision exists, and an unconditional rollout undo would revert to an even older release. Note also that rollout status only proves pods became ready, so add a real smoke test or metric check for anything user-facing.

The following checklist captures the decisions that matter before calling the pipeline production-ready:

GATE
[ ] Required check on the protected branch (not just "runs")
[ ] Always-run job so path filters cannot leave it pending
[ ] Thresholds apply to changed code; flaky tests quarantined, not retried

APPROVAL
[ ] Attached to the environment, shown with diff + test summary
[ ] Approver is not the deployer; timeout behavior known
[ ] Approves a digest, never a rebuild

ROLLBACK
[ ] Previous artifact still in registry (retention checked)
[ ] Migrations follow expand/contract; N-1 code works with N schema
[ ] Rollback rehearsed in staging; works without bypassing the gate
[ ] Post-deploy verification triggers it automatically

The migration line deserves emphasis. Expand/contract means three steps:

  1. Add the new schema first.
  2. Ship code that tolerates both shapes.
  3. Remove the old shape in a later release.

That keeps the previous application version valid. It makes rollback a routine operation instead of a data-recovery project.

More pipeline design notes are available on the DevOps_DayS home page. None of this removes the need for judgment. It only ensures that the gate, the approval and the rollback each do the one job they were added for.

Related

Top comments (0)