Why your team keeps reverting to weekly deploys (and how to actually fix it)
Here's a pattern that shows up constantly in teams of 15-40 engineers: leadership pushes for faster releases, the team bumps deploy frequency up, and within a few weeks someone quietly walks it back to weekly. No mandate, no meeting. Just two or three rough releases in a row, and the team self-polices back to "safe" cadence.
It looks like a discipline problem. It isn't. It's an infrastructure problem, and it's fixable.
The real issue: your pipeline was built for a different era
When you deploy weekly, you can afford sloppiness because a human has time to babysit each release. Watch the error rate for 20 minutes, glance at a dashboard, ship it. Try that five times a day and the manual check either disappears or becomes theater nobody trusts.
Under the hood, this usually comes down to four coupled problems:
- Deploys aren't atomic. Code, migrations, and config all ship in one shot. If the migration is slow, the code can't go out without it.
- Rollback means redeploying. Reverting takes the same pipeline, the same time, the same risk as deploying forward.
- Monitoring lags deploy speed. Alert thresholds and log delays were tuned for hourly checks, not minute-by-minute validation.
- Blast radius is 100%. Every deploy hits the whole fleet at once, so any bad release is automatically a full incident.
None of this is about code quality or test coverage. It's about the deploy path, routing layer, and monitoring stack not being built for the frequency you're asking of them.
The fix: decouple, automate rollback, shrink the blast radius
1. Split schema changes from code deploys
Migrations are usually the biggest source of deploy anxiety. Use the expand/contract pattern so schema changes never ship in the same step as code:
-- Step 1: expand (safe, backward compatible)
ALTER TABLE orders ADD COLUMN shipping_method_v2 VARCHAR(50) NULL;
-- Step 2: dual-write in app code, backfill
UPDATE orders SET shipping_method_v2 = shipping_method WHERE shipping_method_v2 IS NULL;
-- Step 3: contract (separate deploy, days later)
ALTER TABLE orders DROP COLUMN shipping_method;
Code deploys stop waiting on migrations, and migrations stop blocking rollbacks.
2. Make rollback a routing flip, not a redeploy
If reverting means rerunning the pipeline backward, people will hesitate to ship under pressure. Rollback should be a traffic switch, done in seconds:
# Nginx upstream weight shift for canary rollback
upstream backend {
server app-v124-1:8080 weight=0; # new version, was live
server app-v123-1:8080 weight=100; # previous stable, restored
}
# reload takes effect in under 1 second
nginx -s reload
Or on Kubernetes:
kubectl rollout undo deployment/checkout-api
kubectl rollout status deployment/checkout-api --timeout=30s
Ship to a small slice of instances first. Healthy? Shift more weight over. Not healthy? Shift back to zero. No twelve-minute rebuild required to recover.
3. Automate the canary check
A human staring at Grafana for 20 minutes works at weekly cadence. At five deploys a day, it's neither sustainable nor reliable, people get worse at spotting anomalies the more times they repeat the check.
canary:
metrics:
- name: error_rate
threshold: baseline + 0.5%
window: 5m
- name: p95_latency
threshold: baseline + 15%
window: 5m
- name: 5xx_count
threshold: baseline + 10
window: 5m
action_on_fail: auto_rollback
promotion_steps: [5%, 25%, 50%, 100%]
step_duration: 5m
This is probably the single highest-leverage change for teams moving to daily deploys. It turns a subjective judgment call into a repeatable gate.
4. Shrink the blast radius with progressive rollout
Deploying to 100% of instances at once makes every release a full-fleet bet. Instead:
- Ship to 5% of instances or one availability zone
- Hold through a full traffic cycle, including cron jobs and batch work
- Auto-promote on healthy metrics, auto-rollback on bad ones
- Only reach 100% after clearing every stage
Same principle as safe database migrations: keep each irreversible step small.
How to know it actually worked
Don't just track "deploy success." Watch these specifically:
- Change failure rate: percentage of deploys triggering a rollback or hotfix. Aim under 15% (DORA elite benchmark), trending down as frequency goes up.
- MTTR: with automated rollback via traffic weight shift, this should drop under 5 minutes.
- Rollback execution time: measure from "threshold breached" to "traffic fully reverted." Should be seconds to low minutes, not tied to build time.
- Canary false negative rate: how often something bad slips past the gate and gets caught by users instead. If it happens more than once a quarter, retune thresholds, don't remove the gate.
- Deploy frequency vs. incident count: these two lines should decouple. If incidents still climb with frequency, your canary metrics are probably missing a business signal like checkout completion or cart abandonment.
Give this at least four weeks at the new cadence before declaring victory. One good week proves nothing.
Keeping it from backsliding
- Alert on deploy frequency itself. A quiet drop without an explicit decision usually means the pipeline is causing avoidance.
- Treat canary thresholds as living config, not a one-time setup. Traffic patterns shift, thresholds need to shift with them.
Full technical breakdown, including the zero-downtime migration mechanics: read the original article
Originally published on binadit.com
Top comments (0)