DEV Community

Cover image for Solving the deploy frequency wall: from weekly releases to multiple daily deploys without new incidents
binadit
binadit

Posted on Originally published at binadit.com

Solving the deploy frequency wall: from weekly releases to multiple daily deploys without new incidents

Why your team keeps reverting to weekly deploys (and how to actually fix it)

Here's a pattern that shows up constantly in teams of 15-40 engineers: leadership pushes for faster releases, the team bumps deploy frequency up, and within a few weeks someone quietly walks it back to weekly. No mandate, no meeting. Just two or three rough releases in a row, and the team self-polices back to "safe" cadence.

It looks like a discipline problem. It isn't. It's an infrastructure problem, and it's fixable.

The real issue: your pipeline was built for a different era

When you deploy weekly, you can afford sloppiness because a human has time to babysit each release. Watch the error rate for 20 minutes, glance at a dashboard, ship it. Try that five times a day and the manual check either disappears or becomes theater nobody trusts.

Under the hood, this usually comes down to four coupled problems:

  • Deploys aren't atomic. Code, migrations, and config all ship in one shot. If the migration is slow, the code can't go out without it.
  • Rollback means redeploying. Reverting takes the same pipeline, the same time, the same risk as deploying forward.
  • Monitoring lags deploy speed. Alert thresholds and log delays were tuned for hourly checks, not minute-by-minute validation.
  • Blast radius is 100%. Every deploy hits the whole fleet at once, so any bad release is automatically a full incident.

None of this is about code quality or test coverage. It's about the deploy path, routing layer, and monitoring stack not being built for the frequency you're asking of them.

The fix: decouple, automate rollback, shrink the blast radius

1. Split schema changes from code deploys

Migrations are usually the biggest source of deploy anxiety. Use the expand/contract pattern so schema changes never ship in the same step as code:

-- Step 1: expand (safe, backward compatible)
ALTER TABLE orders ADD COLUMN shipping_method_v2 VARCHAR(50) NULL;

-- Step 2: dual-write in app code, backfill
UPDATE orders SET shipping_method_v2 = shipping_method WHERE shipping_method_v2 IS NULL;

-- Step 3: contract (separate deploy, days later)
ALTER TABLE orders DROP COLUMN shipping_method;
Enter fullscreen mode Exit fullscreen mode

Code deploys stop waiting on migrations, and migrations stop blocking rollbacks.

2. Make rollback a routing flip, not a redeploy

If reverting means rerunning the pipeline backward, people will hesitate to ship under pressure. Rollback should be a traffic switch, done in seconds:

# Nginx upstream weight shift for canary rollback
upstream backend {
    server app-v124-1:8080 weight=0;   # new version, was live
    server app-v123-1:8080 weight=100; # previous stable, restored
}
# reload takes effect in under 1 second
nginx -s reload
Enter fullscreen mode Exit fullscreen mode

Or on Kubernetes:

kubectl rollout undo deployment/checkout-api
kubectl rollout status deployment/checkout-api --timeout=30s
Enter fullscreen mode Exit fullscreen mode

Ship to a small slice of instances first. Healthy? Shift more weight over. Not healthy? Shift back to zero. No twelve-minute rebuild required to recover.

3. Automate the canary check

A human staring at Grafana for 20 minutes works at weekly cadence. At five deploys a day, it's neither sustainable nor reliable, people get worse at spotting anomalies the more times they repeat the check.

canary:
  metrics:
    - name: error_rate
      threshold: baseline + 0.5%
      window: 5m
    - name: p95_latency
      threshold: baseline + 15%
      window: 5m
    - name: 5xx_count
      threshold: baseline + 10
      window: 5m
  action_on_fail: auto_rollback
  promotion_steps: [5%, 25%, 50%, 100%]
  step_duration: 5m
Enter fullscreen mode Exit fullscreen mode

This is probably the single highest-leverage change for teams moving to daily deploys. It turns a subjective judgment call into a repeatable gate.

4. Shrink the blast radius with progressive rollout

Deploying to 100% of instances at once makes every release a full-fleet bet. Instead:

  1. Ship to 5% of instances or one availability zone
  2. Hold through a full traffic cycle, including cron jobs and batch work
  3. Auto-promote on healthy metrics, auto-rollback on bad ones
  4. Only reach 100% after clearing every stage

Same principle as safe database migrations: keep each irreversible step small.

How to know it actually worked

Don't just track "deploy success." Watch these specifically:

  • Change failure rate: percentage of deploys triggering a rollback or hotfix. Aim under 15% (DORA elite benchmark), trending down as frequency goes up.
  • MTTR: with automated rollback via traffic weight shift, this should drop under 5 minutes.
  • Rollback execution time: measure from "threshold breached" to "traffic fully reverted." Should be seconds to low minutes, not tied to build time.
  • Canary false negative rate: how often something bad slips past the gate and gets caught by users instead. If it happens more than once a quarter, retune thresholds, don't remove the gate.
  • Deploy frequency vs. incident count: these two lines should decouple. If incidents still climb with frequency, your canary metrics are probably missing a business signal like checkout completion or cart abandonment.

Give this at least four weeks at the new cadence before declaring victory. One good week proves nothing.

Keeping it from backsliding

  • Alert on deploy frequency itself. A quiet drop without an explicit decision usually means the pipeline is causing avoidance.
  • Treat canary thresholds as living config, not a one-time setup. Traffic patterns shift, thresholds need to shift with them.

Full technical breakdown, including the zero-downtime migration mechanics: read the original article

Originally published on binadit.com

Top comments (0)