DEV Community

Clara Decker
Clara Decker

Posted on

Canary Deployments on Error Budgets: A Fixed Error-Rate Threshold Is a Coin Flip

The standard canary: shift 5% of traffic to the new version, watch for five minutes, abort if the error rate exceeds 1%.
Consider what that actually measures. 5% of 200 requests/minute over five minutes is 50 requests. One error is a 2% error rate — abort. Zero errors “proves” nothing; a version with a genuine 1.5% regression has roughly a 47% chance of producing zero errors in 50 requests.
You are flipping a coin and calling it a deployment gate. It fails both ways: false aborts on healthy deploys, and false confidence on broken ones. Teams respond by loosening the threshold, which only removes the false aborts and keeps the false confidence.
Two things fix it: compare against the baseline rather than an absolute number, and require enough samples for the comparison to mean anything.
Sample size first
from future import annotations

import math
from dataclasses import dataclass

def required_samples(
baseline_rate: float,
min_detectable_effect: float,
power: float = 0.8,
alpha: float = 0.05,
) -> int:
"""Samples per arm needed to detect a rate change of min_detectable_effect.

Two-proportion z-test. The output is usually much larger than teams
expect, and that is the actual finding: at low traffic you cannot canary
on error rate at all, and pretending otherwise is theatre."""
p1 = max(baseline_rate, 1e-6)
p2 = p1 + min_detectable_effect
p_bar = (p1 + p2) / 2

z_alpha = 1.959963985 if alpha == 0.05 else abs(_z(1 - alpha / 2))
z_beta = 0.841621234 if power == 0.8 else abs(_z(power))

num = (
    z_alpha * math.sqrt(2 * p_bar * (1 - p_bar))
    + z_beta * math.sqrt(p1 * (1 - p1) + p2 * (1 - p2))
) ** 2
return math.ceil(num / (min_detectable_effect ** 2))
Enter fullscreen mode Exit fullscreen mode

def _z(p: float) -> float:
"""Inverse normal CDF, Acklam's rational approximation."""
a = [-3.969683028665376e+01, 2.209460984245205e+02, -2.759285104469687e+02,
1.383577518672690e+02, -3.066479806614716e+01, 2.506628277459239e+00]
b = [-5.447609879822406e+01, 1.615858368580409e+02, -1.556989798598866e+02,
6.680131188771972e+01, -1.328068155288572e+01]
c = [-7.784894002430293e-03, -3.223964580411365e-01, -2.400758277161838e+00,
-2.549732539343734e+00, 4.374664141464968e+00, 2.938163982698783e+00]
d = [7.784695709041462e-03, 3.224671290700398e-01, 2.445134137142996e+00,
3.754408661907416e+00]
plow, phigh = 0.02425, 1 - 0.02425
if p < plow:
q = math.sqrt(-2 * math.log(p))
return (((((c[0]*q+c[1])*q+c[2])*q+c[3])*q+c[4])*q+c[5]) / \
((((d[0]*q+d[1])*q+d[2])*q+d[3])*q+1)
if p > phigh:
q = math.sqrt(-2 * math.log(1 - p))
return -(((((c[0]*q+c[1])*q+c[2])*q+c[3])*q+c[4])*q+c[5]) / \
((((d[0]*q+d[1])*q+d[2])*q+d[3])*q+1)
q = p - 0.5
r = q * q
return (((((a[0]*r+a[1])*r+a[2])*r+a[3])*r+a[4])*r+a[5])*q / \
(((((b[0]*r+b[1])*r+b[2])*r+b[3])*r+b[4])*r+1)

A 0.5% baseline and wanting to catch a 0.5pp regression needs

4,673 requests per arm. At 5% canary traffic and 200 rps that is

~8 min. At 20 rps it is ~1.3 hours - and that is the

honest answer, not 5 minutes.

The gate itself
@dataclass
class ArmStats:
requests: int
errors: int
latency_p99_ms: float

@property
def rate(self) -> float:
    return self.errors / self.requests if self.requests else 0.0
Enter fullscreen mode Exit fullscreen mode

@dataclass
class Verdict:
action: str # "promote" | "hold" | "abort"
reason: str
confidence: float = 0.0

def two_proportion_z(a: ArmStats, b: ArmStats) -> float:
"""Positive z means the canary (b) is worse."""
n1, n2 = a.requests, b.requests
if n1 == 0 or n2 == 0:
return 0.0
p_pool = (a.errors + b.errors) / (n1 + n2)
se = math.sqrt(p_pool * (1 - p_pool) * (1 / n1 + 1 / n2))
return 0.0 if se == 0 else (b.rate - a.rate) / se

def evaluate(
baseline: ArmStats,
canary: ArmStats,
budget_remaining: float, # fraction of the SLO error budget left
min_samples: int,
hard_abort_rate: float = 0.05,
) -> Verdict:
# Hard circuit breaker, no statistics required. A canary returning 5%
# errors is broken; do not wait for significance.
if canary.requests >= 50 and canary.rate >= hard_abort_rate:
return Verdict("abort", f"canary error rate {canary.rate:.1%} exceeds hard limit")

if canary.requests < min_samples:
    return Verdict("hold", f"{canary.requests}/{min_samples} samples")

z = two_proportion_z(baseline, canary)
if z > 1.96:
    return Verdict("abort", f"canary significantly worse (z={z:.2f})", 0.95)

# Error budget governs how much risk you may take, not the canary itself.
# Nearly exhausted budget means small steps and long bakes.
if budget_remaining < 0.1:
    return Verdict("hold", "error budget nearly exhausted; deploys frozen")

if canary.latency_p99_ms > baseline.latency_p99_ms * 1.25:
    return Verdict("abort", "p99 latency regression >25%")

return Verdict("promote", f"no significant regression (z={z:.2f})", 0.95)
Enter fullscreen mode Exit fullscreen mode

Why baseline comparison, not an absolute threshold
Your error rate is not constant. It varies with time of day, with a flaky upstream, with a bot crawl. An absolute 1% threshold aborts a perfectly good canary during a period when the baseline is at 1.2%, and it passes a bad canary during a quiet period when the baseline is 0.05%.
Comparing arms at the same moment cancels all of that. Both versions see the same upstream, the same traffic mix, the same hour. That is the entire reason to run a canary rather than deploying to staging.
Traffic splitting has to be sticky
A user who hits the canary on request one and the baseline on request two experiences a version-mismatch bug you will never reproduce. Split on a stable hash of a session or user identifier, not per request:
import hashlib

def in_canary(session_id: str, percentage: float) -> bool:
"""Stable per session. Also means increasing the percentage only ever
ADDS sessions to the canary - nobody gets moved back."""
h = hashlib.blake2b(session_id.encode(), digest_size=8).digest()
return (int.from_bytes(h, "big") % 10_000) < percentage * 100
The monotonicity property matters. With a naive random split, raising the canary from 5% to 10% reshuffles everyone; with a stable hash, the original 5% stay put and 5% more join. That makes the comparison valid across steps and makes rollback clean.
What the error budget is actually for
The budget does not decide whether this canary is healthy — the statistics do. It decides how much risk you are permitted to take right now.
Full budget: 20% steps, short bakes, automatic promotion. Half consumed: 5% steps, longer bakes, a human approves promotion. Under 10%: no deploys except fixes. That policy converts the SLO from a dashboard into something that changes behaviour, which is the only reason to have one.
If your traffic is too low
Accept it. At 20 requests per second you cannot statistically canary a 0.5pp regression in a useful timeframe, and no amount of dashboard sophistication changes that.
What works instead: shadow traffic (mirror production requests to the new version and compare responses without serving them), heavier pre-production testing, and feature flags that let you turn off behaviour without redeploying. Those are the tools for low-volume services, and reaching for a canary there produces a gate that looks rigorous and decides nothing.

We build delivery pipelines and platform tooling at SoluLab — more on our software consulting work.

Top comments (0)