<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Clara Decker</title>
    <description>The latest articles on DEV Community by Clara Decker (@clara_decker_99c61ab5fe0e).</description>
    <link>https://dev.to/clara_decker_99c61ab5fe0e</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4163012%2F0e051f96-bfd5-4961-8d4f-65301cdc7c9c.png</url>
      <title>DEV Community: Clara Decker</title>
      <link>https://dev.to/clara_decker_99c61ab5fe0e</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/clara_decker_99c61ab5fe0e"/>
    <language>en</language>
    <item>
      <title>Postgres Indexes Under Write Load: Every Index Is a Tax on Every Write</title>
      <dc:creator>Clara Decker</dc:creator>
      <pubDate>Mon, 05 Oct 2026 06:53:45 +0000</pubDate>
      <link>https://dev.to/clara_decker_99c61ab5fe0e/postgres-indexes-under-write-load-every-index-is-a-tax-on-every-write-3pmm</link>
      <guid>https://dev.to/clara_decker_99c61ab5fe0e/postgres-indexes-under-write-load-every-index-is-a-tax-on-every-write-3pmm</guid>
      <description>&lt;p&gt;An index speeds up reads and slows down writes. That trade is well known and almost never quantified, so tables accumulate indexes nobody removes.&lt;br&gt;
The cost is larger than it looks. Every INSERT updates every index on the table. Every update that changes an indexed column updates that index — and critically, an update that touches any indexed column may lose the HOT (heap-only tuple) optimization, which turns a cheap in-page update into one that writes to every index on the table. One badly chosen index can measurably slow writes that do not even reference it.&lt;br&gt;
Three techniques give you most of the read benefit at a fraction of the write cost.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Partial indexes
If your query always filters on a predicate, put the predicate in the index:
-- Full index: every row, updated on every insert.
CREATE INDEX idx_orders_status ON orders (status);&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;-- Partial: only rows in a state you actually query, and one that is a&lt;br&gt;
-- small fraction of the table. Rows that never match are never indexed,&lt;br&gt;
-- so inserting a completed order touches this index not at all.&lt;br&gt;
CREATE INDEX idx_orders_pending ON orders (created_at)&lt;br&gt;
  WHERE status IN ('pending', 'processing');&lt;br&gt;
On a table where 99% of rows are completed, the partial index is roughly 1% of the size, and — because completed orders never enter it — inserting and updating them skips it entirely. This is the single highest-leverage index pattern in Postgres and it is underused.&lt;br&gt;
The catch: the planner will only use it if it can prove your query’s WHERE clause implies the index predicate. WHERE status = 'pending' works. WHERE status = ANY($1) with a parameter array does not, because the value is not known at plan time. Check with EXPLAIN using realistic parameters, not literals.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Covering indexes with INCLUDE
An index-only scan avoids touching the heap at all — but only if every column the query needs is in the index. INCLUDE adds payload columns without making them part of the search key:
CREATE INDEX idx_orders_customer_covering
ON orders (customer_id, created_at DESC)
INCLUDE (total_cents, status);
INCLUDE columns live only in leaf pages, so they do not bloat internal nodes or affect the sort order. The index is larger than a plain one but supports index-only scans.
Index-only scans depend on the visibility map, which is maintained by VACUUM. On a heavily-updated table with lazy autovacuum, the visibility map is stale and the “index-only” scan falls back to heap fetches anyway. Check Heap Fetches in EXPLAIN (ANALYZE, BUFFERS) — if it is high, tune autovacuum before adding more covering indexes.&lt;/li&gt;
&lt;li&gt;Find what you are paying for
UNUSED_INDEXES = """
SELECT
s.schemaname, s.relname AS table_name, s.indexrelname AS index_name,
pg_size_pretty(pg_relation_size(s.indexrelid)) AS size,
pg_relation_size(s.indexrelid) AS size_bytes,
s.idx_scan
FROM pg_stat_user_indexes s
JOIN pg_index i ON i.indexrelid = s.indexrelid
WHERE s.idx_scan &amp;lt; %(min_scans)s
AND NOT i.indisunique          -- unique indexes enforce constraints
AND NOT i.indisprimary
AND pg_relation_size(s.indexrelid) &amp;gt; %(min_bytes)s
ORDER BY pg_relation_size(s.indexrelid) DESC;
"""&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;DUPLICATE_INDEXES = """&lt;br&gt;
-- Indexes whose column list is a PREFIX of another index's column list.&lt;br&gt;
-- (a) is redundant if (a, b) exists: the composite serves both.&lt;br&gt;
SELECT a.indexrelid::regclass AS redundant,&lt;br&gt;
       b.indexrelid::regclass AS covered_by,&lt;br&gt;
       pg_size_pretty(pg_relation_size(a.indexrelid)) AS wasted&lt;br&gt;
FROM pg_index a&lt;br&gt;
JOIN pg_index b&lt;br&gt;
  ON a.indrelid = b.indrelid&lt;br&gt;
 AND a.indexrelid &amp;lt;&amp;gt; b.indexrelid&lt;br&gt;
 AND array_to_string(b.indkey, ' ') LIKE array_to_string(a.indkey, ' ') || '%%'&lt;br&gt;
WHERE NOT a.indisunique AND NOT a.indisprimary;&lt;br&gt;
"""&lt;/p&gt;

&lt;p&gt;WRITE_AMPLIFICATION = """&lt;br&gt;
-- Index count and total index bytes per write. The ratio of index size to&lt;br&gt;
-- table size is a rough proxy for how much each write costs you.&lt;br&gt;
SELECT&lt;br&gt;
  t.relname AS table_name,&lt;br&gt;
  count(i.indexrelid) AS index_count,&lt;br&gt;
  pg_size_pretty(pg_relation_size(t.oid)) AS table_size,&lt;br&gt;
  pg_size_pretty(sum(pg_relation_size(i.indexrelid))) AS index_size,&lt;br&gt;
  round(sum(pg_relation_size(i.indexrelid))::numeric&lt;br&gt;
        / NULLIF(pg_relation_size(t.oid), 0), 2) AS index_to_table_ratio,&lt;br&gt;
  s.n_tup_ins + s.n_tup_upd + s.n_tup_del AS writes&lt;br&gt;
FROM pg_class t&lt;br&gt;
JOIN pg_stat_user_tables s ON s.relid = t.oid&lt;br&gt;
LEFT JOIN pg_index i ON i.indrelid = t.oid&lt;br&gt;
WHERE t.relkind = 'r'&lt;br&gt;
GROUP BY t.relname, t.oid, s.n_tup_ins, s.n_tup_upd, s.n_tup_del&lt;br&gt;
HAVING count(i.indexrelid) &amp;gt; 3&lt;br&gt;
ORDER BY index_to_table_ratio DESC;&lt;br&gt;
"""&lt;/p&gt;

&lt;p&gt;def audit(conn, min_scans: int = 50, min_bytes: int = 10 * 1024 * 1024) -&amp;gt; dict:&lt;br&gt;
    """An index-to-table ratio above ~1.5 with high write volume is a strong&lt;br&gt;
    candidate for pruning. Read &lt;code&gt;idx_scan&lt;/code&gt; with care: stats reset on&lt;br&gt;
    restart and on pg_stat_reset(), so a 'never used' index may just be&lt;br&gt;
    new. Check pg_stat_get_db_stat_reset_time() before believing it."""&lt;br&gt;
    with conn.cursor() as cur:&lt;br&gt;
        cur.execute("SELECT stats_reset FROM pg_stat_database WHERE datname = current_database()")&lt;br&gt;
        reset_at = cur.fetchone()[0]&lt;br&gt;
        cur.execute(UNUSED_INDEXES, {"min_scans": min_scans, "min_bytes": min_bytes})&lt;br&gt;
        unused = cur.fetchall()&lt;br&gt;
        cur.execute(DUPLICATE_INDEXES)&lt;br&gt;
        dupes = cur.fetchall()&lt;br&gt;
        cur.execute(WRITE_AMPLIFICATION)&lt;br&gt;
        amplification = cur.fetchall()&lt;br&gt;
    return {&lt;br&gt;
        "stats_since": reset_at,&lt;br&gt;
        "unused": unused,&lt;br&gt;
        "wasted_bytes": sum(r[4] for r in unused),&lt;br&gt;
        "redundant": dupes,&lt;br&gt;
        "write_amplification": amplification,&lt;br&gt;
    }&lt;br&gt;
The stats_since field is not decoration. Dropping an index because idx_scan = 0 when statistics were reset two days ago is how you delete the index that serves the monthly reporting job.&lt;br&gt;
Dropping safely&lt;br&gt;
Postgres 12+ can deactivate an index without dropping it:&lt;br&gt;
UPDATE pg_index SET indisvalid = false&lt;br&gt;
WHERE indexrelid = 'idx_orders_status'::regclass;&lt;br&gt;
-- Planner now ignores it. Writes still maintain it, so this tests the READ&lt;br&gt;
-- impact only - but that is the risky half. Watch for a week, then DROP.&lt;br&gt;
And always build with CONCURRENTLY on a live table:&lt;br&gt;
CREATE INDEX CONCURRENTLY idx_new ON orders (customer_id);&lt;br&gt;
-- No ACCESS EXCLUSIVE lock. Slower, two table passes, and it can leave an&lt;br&gt;
-- INVALID index behind if it fails - check indisvalid afterwards and drop&lt;br&gt;
-- the corpse, or the next CREATE INDEX CONCURRENTLY will trip over it.&lt;br&gt;
The habit worth building&lt;br&gt;
Add index review to the same cadence as dependency updates. Every quarter: run the audit, drop what is genuinely unused, and convert full indexes to partial where a predicate is stable.&lt;br&gt;
The reason this matters more over time is that indexes are added in response to a slow query and essentially never removed when that query changes. After three years a hot table has eleven indexes, four of which serve queries that no longer exist, and every write pays for all eleven.&lt;/p&gt;

&lt;p&gt;We do data platform and database performance work at SoluLab — the &lt;a href="https://www.solulab.com/case-study/data-empowerment-with-infusenet/" rel="noopener noreferrer"&gt;InfuseNet case study&lt;/a&gt; covers one at scale.&lt;/p&gt;

</description>
      <category>database</category>
      <category>performance</category>
      <category>postgres</category>
      <category>sql</category>
    </item>
    <item>
      <title>Model Extraction: Rate Limiting Buys Time, It Does Not Stop Distillation</title>
      <dc:creator>Clara Decker</dc:creator>
      <pubDate>Mon, 05 Oct 2026 06:47:12 +0000</pubDate>
      <link>https://dev.to/clara_decker_99c61ab5fe0e/model-extraction-rate-limiting-buys-time-it-does-not-stop-distillation-1d9i</link>
      <guid>https://dev.to/clara_decker_99c61ab5fe0e/model-extraction-rate-limiting-buys-time-it-does-not-stop-distillation-1d9i</guid>
      <description>&lt;p&gt;Model extraction is straightforward in principle: query the API, collect input-output pairs, train a student model on them. For a classifier with a well-defined decision boundary, a few tens of thousands of well-chosen queries can produce a student that agrees with the victim on the large majority of inputs.&lt;br&gt;
What makes it efficient is active learning. The attacker does not sample randomly — they concentrate queries near the decision boundary, where each answer is maximally informative. Uncertainty sampling can cut the query budget by an order of magnitude versus random sampling.&lt;br&gt;
Two things follow that shape every defence:&lt;br&gt;
What leaks scales with output granularity. A full probability vector leaks far more per query than a top-1 label. Confidence scores are the attacker’s gradient signal.&lt;br&gt;
Prevention is not achievable. If the API is useful, it is extractable — the information the legitimate user needs is the information the attacker needs. The goal is to make extraction cost more than licensing, and to detect it while it is in progress.&lt;br&gt;
Defence 1: reduce what each answer reveals&lt;br&gt;
from &lt;strong&gt;future&lt;/strong&gt; import annotations&lt;/p&gt;

&lt;p&gt;import hashlib&lt;br&gt;
import math&lt;br&gt;
from dataclasses import dataclass&lt;/p&gt;

&lt;p&gt;@dataclass&lt;br&gt;
class OutputPolicy:&lt;br&gt;
    top_k: int = 3                 # never return the full vector&lt;br&gt;
    round_to: int = 2              # decimal places on probabilities&lt;br&gt;
    temperature: float = 1.0       # &amp;gt;1 flattens, leaking less boundary detail&lt;br&gt;
    deterministic_noise: bool = True&lt;/p&gt;

&lt;p&gt;def defend_output(&lt;br&gt;
    probs: dict[str, float], request_key: str, policy: OutputPolicy&lt;br&gt;
) -&amp;gt; dict[str, float]:&lt;br&gt;
    """Truncate, round, and add small DETERMINISTIC noise.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Determinism matters: random noise per call is averaged away by an
attacker who queries the same point repeatedly, and it costs you
reproducibility. Derive the perturbation from the input so the same
input always gets the same answer."""
if policy.temperature != 1.0:
    t = policy.temperature
    logits = {k: math.log(max(v, 1e-9)) / t for k, v in probs.items()}
    m = max(logits.values())
    exp = {k: math.exp(v - m) for k, v in logits.items()}
    total = sum(exp.values())
    probs = {k: v / total for k, v in exp.items()}

top = dict(sorted(probs.items(), key=lambda kv: -kv[1])[: policy.top_k])

if policy.deterministic_noise:
    seed = int.from_bytes(
        hashlib.blake2b(request_key.encode(), digest_size=4).digest(), "big")
    for i, k in enumerate(top):
        # +/- half a rounding unit, stable per input.
        jitter = (((seed &amp;gt;&amp;gt; (i * 5)) % 1000) / 1000 - 0.5) * (10 ** -policy.round_to)
        top[k] = max(0.0, min(1.0, top[k] + jitter))

return {k: round(v, policy.round_to) for k, v in top.items()}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Rounding to 2 decimal places sounds trivial. It is not — it removes most of the fine-grained boundary information that uncertainty sampling depends on, and it is invisible to nearly every legitimate consumer. This is the highest value-per-effort defence on the list, and the one most often skipped because it feels like degrading the product.&lt;br&gt;
Defence 2: detect the query distribution, not the volume&lt;br&gt;
An extractor’s query pattern is statistically distinct from a user’s. Legitimate traffic clusters around a few common cases. Extraction traffic is unusually uniform over the input space, unusually close to the decision boundary, and has unusually low repetition.&lt;br&gt;
@dataclass&lt;br&gt;
class ExtractionSignals:&lt;br&gt;
    """Per-principal, over a rolling window. None of these is conclusive&lt;br&gt;
    alone; the combination is."""&lt;br&gt;
    queries: int = 0&lt;br&gt;
    unique_inputs: int = 0&lt;br&gt;
    near_boundary: int = 0      # confidence in [0.4, 0.6]&lt;br&gt;
    mean_pairwise_distance: float = 0.0&lt;br&gt;
    distinct_classes_returned: int = 0&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;def score(self) -&amp;gt; tuple[float, list[str]]:
    reasons: list[str] = []
    s = 0.0
    if self.queries &amp;lt; 200:
        return 0.0, ["insufficient volume to assess"]

    uniqueness = self.unique_inputs / self.queries
    if uniqueness &amp;gt; 0.97:
        s += 0.3
        reasons.append(f"near-zero repetition ({uniqueness:.2f})")

    boundary_frac = self.near_boundary / self.queries
    if boundary_frac &amp;gt; 0.35:
        s += 0.4
        reasons.append(f"{boundary_frac:.0%} of queries near the decision boundary")

    # Real users do not sweep the whole input space evenly.
    if self.mean_pairwise_distance &amp;gt; 0.8:
        s += 0.2
        reasons.append("queries uniformly spread across input space")

    if self.distinct_classes_returned &amp;gt;= 0.9 * 10:
        s += 0.1
        reasons.append("coverage of nearly all output classes")

    return min(s, 1.0), reasons
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The boundary-concentration signal is the strongest. A user classifying real documents gets confident predictions most of the time; an attacker doing uncertainty sampling is deliberately querying where the model is unsure, and that shows up immediately as an anomalous confidence histogram.&lt;br&gt;
Defence 3: watermark the decision boundary&lt;br&gt;
You cannot stop extraction, so make a stolen model provable. Train the victim to produce specific, arbitrary outputs on a secret set of trigger inputs — points far from the natural data distribution, where the behaviour costs nothing on real traffic. A student distilled from the API inherits them.&lt;br&gt;
def verify_ownership(suspect_predict, triggers: list[tuple], alpha: float = 1e-6) -&amp;gt; dict:&lt;br&gt;
    """Under the null hypothesis (independent model), matching the trigger&lt;br&gt;
    labels is chance. With enough triggers the binomial p-value is tiny."""&lt;br&gt;
    n = len(triggers)&lt;br&gt;
    matches = sum(1 for x, y in triggers if suspect_predict(x) == y)&lt;br&gt;
    p_chance = 1 / 10                   # number of classes&lt;br&gt;
    p_value = sum(&lt;br&gt;
        math.comb(n, k) * p_chance*&lt;em&gt;k * (1 - p_chance) *&lt;/em&gt; (n - k)&lt;br&gt;
        for k in range(matches, n + 1)&lt;br&gt;
    )&lt;br&gt;
    return {"triggers": n, "matches": matches, "p_value": p_value,&lt;br&gt;
            "claim_supported": p_value &amp;lt; alpha}&lt;br&gt;
With 100 triggers over 10 classes, chance gives about 10 matches; 40 matches produces an astronomically small p-value. That is evidence you can put in front of a lawyer, which is the actual remedy — technical defences delay extraction, legal ones address it.&lt;br&gt;
What to actually deploy&lt;br&gt;
In order of value per unit of effort:&lt;br&gt;
Round and truncate outputs. Free, instant, removes most of the leak.&lt;br&gt;
Authenticate everything and rate limit per principal. An unauthenticated endpoint cannot be defended at all, because the attacker rotates IPs and there is nothing to bind the budget to.&lt;br&gt;
Log the confidence distribution per principal and alert on anomalies. Cheap, and it catches extraction while it is running rather than after.&lt;br&gt;
Watermark. Moderate effort, and the only one that gives you a remedy after the fact.&lt;br&gt;
Differential privacy on outputs. A real guarantee, a real accuracy cost. Reach for it when the model itself is the product.&lt;br&gt;
The honest limitation&lt;br&gt;
None of this stops a determined, well-resourced attacker who is patient. They can spread queries across accounts, sample slowly, mimic legitimate distributions and accept a lower-fidelity student.&lt;br&gt;
What the defences do is change the economics. Extraction that needs 500,000 queries across 50 accounts over three months, with a detection risk at every step, is a different proposition from one that needs 20,000 queries in an afternoon. For most models that difference is the whole game — and for the handful where it is not, the answer is not to serve the model over a public API at all.&lt;/p&gt;

&lt;p&gt;We build and secure ML systems at SoluLab — more on our &lt;a href="https://www.solulab.com/machine-learning-development-company/" rel="noopener noreferrer"&gt;machine learning development&lt;/a&gt; work.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cybersecurity</category>
      <category>machinelearning</category>
      <category>security</category>
    </item>
    <item>
      <title>Canary Deployments on Error Budgets: A Fixed Error-Rate Threshold Is a Coin Flip</title>
      <dc:creator>Clara Decker</dc:creator>
      <pubDate>Mon, 05 Oct 2026 06:19:05 +0000</pubDate>
      <link>https://dev.to/clara_decker_99c61ab5fe0e/canary-deployments-on-error-budgets-a-fixed-error-rate-threshold-is-a-coin-flip-5a92</link>
      <guid>https://dev.to/clara_decker_99c61ab5fe0e/canary-deployments-on-error-budgets-a-fixed-error-rate-threshold-is-a-coin-flip-5a92</guid>
      <description>&lt;p&gt;The standard canary: shift 5% of traffic to the new version, watch for five minutes, abort if the error rate exceeds 1%.&lt;br&gt;
Consider what that actually measures. 5% of 200 requests/minute over five minutes is 50 requests. One error is a 2% error rate — abort. Zero errors “proves” nothing; a version with a genuine 1.5% regression has roughly a 47% chance of producing zero errors in 50 requests.&lt;br&gt;
You are flipping a coin and calling it a deployment gate. It fails both ways: false aborts on healthy deploys, and false confidence on broken ones. Teams respond by loosening the threshold, which only removes the false aborts and keeps the false confidence.&lt;br&gt;
Two things fix it: compare against the baseline rather than an absolute number, and require enough samples for the comparison to mean anything.&lt;br&gt;
Sample size first&lt;br&gt;
from &lt;strong&gt;future&lt;/strong&gt; import annotations&lt;/p&gt;

&lt;p&gt;import math&lt;br&gt;
from dataclasses import dataclass&lt;/p&gt;

&lt;p&gt;def required_samples(&lt;br&gt;
    baseline_rate: float,&lt;br&gt;
    min_detectable_effect: float,&lt;br&gt;
    power: float = 0.8,&lt;br&gt;
    alpha: float = 0.05,&lt;br&gt;
) -&amp;gt; int:&lt;br&gt;
    """Samples per arm needed to detect a rate change of &lt;code&gt;min_detectable_effect&lt;/code&gt;.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Two-proportion z-test. The output is usually much larger than teams
expect, and that is the actual finding: at low traffic you cannot canary
on error rate at all, and pretending otherwise is theatre."""
p1 = max(baseline_rate, 1e-6)
p2 = p1 + min_detectable_effect
p_bar = (p1 + p2) / 2

z_alpha = 1.959963985 if alpha == 0.05 else abs(_z(1 - alpha / 2))
z_beta = 0.841621234 if power == 0.8 else abs(_z(power))

num = (
    z_alpha * math.sqrt(2 * p_bar * (1 - p_bar))
    + z_beta * math.sqrt(p1 * (1 - p1) + p2 * (1 - p2))
) ** 2
return math.ceil(num / (min_detectable_effect ** 2))
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;def _z(p: float) -&amp;gt; float:&lt;br&gt;
    """Inverse normal CDF, Acklam's rational approximation."""&lt;br&gt;
    a = [-3.969683028665376e+01, 2.209460984245205e+02, -2.759285104469687e+02,&lt;br&gt;
         1.383577518672690e+02, -3.066479806614716e+01, 2.506628277459239e+00]&lt;br&gt;
    b = [-5.447609879822406e+01, 1.615858368580409e+02, -1.556989798598866e+02,&lt;br&gt;
         6.680131188771972e+01, -1.328068155288572e+01]&lt;br&gt;
    c = [-7.784894002430293e-03, -3.223964580411365e-01, -2.400758277161838e+00,&lt;br&gt;
         -2.549732539343734e+00, 4.374664141464968e+00, 2.938163982698783e+00]&lt;br&gt;
    d = [7.784695709041462e-03, 3.224671290700398e-01, 2.445134137142996e+00,&lt;br&gt;
         3.754408661907416e+00]&lt;br&gt;
    plow, phigh = 0.02425, 1 - 0.02425&lt;br&gt;
    if p &amp;lt; plow:&lt;br&gt;
        q = math.sqrt(-2 * math.log(p))&lt;br&gt;
        return (((((c[0]*q+c[1])*q+c[2])*q+c[3])*q+c[4])*q+c[5]) / \&lt;br&gt;
               ((((d[0]*q+d[1])*q+d[2])*q+d[3])*q+1)&lt;br&gt;
    if p &amp;gt; phigh:&lt;br&gt;
        q = math.sqrt(-2 * math.log(1 - p))&lt;br&gt;
        return -(((((c[0]*q+c[1])*q+c[2])*q+c[3])*q+c[4])*q+c[5]) / \&lt;br&gt;
                ((((d[0]*q+d[1])*q+d[2])*q+d[3])*q+1)&lt;br&gt;
    q = p - 0.5&lt;br&gt;
    r = q * q&lt;br&gt;
    return (((((a[0]*r+a[1])*r+a[2])*r+a[3])*r+a[4])*r+a[5])*q / \&lt;br&gt;
           (((((b[0]*r+b[1])*r+b[2])*r+b[3])*r+b[4])*r+1)&lt;/p&gt;

&lt;h1&gt;
  
  
  A 0.5% baseline and wanting to catch a 0.5pp regression needs
&lt;/h1&gt;

&lt;h1&gt;
  
  
  4,673 requests per arm. At 5% canary traffic and 200 rps that is
&lt;/h1&gt;

&lt;h1&gt;
  
  
  ~8 min. At 20 rps it is ~1.3 hours - and that is the
&lt;/h1&gt;

&lt;h1&gt;
  
  
  honest answer, not 5 minutes.
&lt;/h1&gt;

&lt;p&gt;The gate itself&lt;br&gt;
@dataclass&lt;br&gt;
class ArmStats:&lt;br&gt;
    requests: int&lt;br&gt;
    errors: int&lt;br&gt;
    latency_p99_ms: float&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;@property
def rate(self) -&amp;gt; float:
    return self.errors / self.requests if self.requests else 0.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;@dataclass&lt;br&gt;
class Verdict:&lt;br&gt;
    action: str          # "promote" | "hold" | "abort"&lt;br&gt;
    reason: str&lt;br&gt;
    confidence: float = 0.0&lt;/p&gt;

&lt;p&gt;def two_proportion_z(a: ArmStats, b: ArmStats) -&amp;gt; float:&lt;br&gt;
    """Positive z means the canary (b) is worse."""&lt;br&gt;
    n1, n2 = a.requests, b.requests&lt;br&gt;
    if n1 == 0 or n2 == 0:&lt;br&gt;
        return 0.0&lt;br&gt;
    p_pool = (a.errors + b.errors) / (n1 + n2)&lt;br&gt;
    se = math.sqrt(p_pool * (1 - p_pool) * (1 / n1 + 1 / n2))&lt;br&gt;
    return 0.0 if se == 0 else (b.rate - a.rate) / se&lt;/p&gt;

&lt;p&gt;def evaluate(&lt;br&gt;
    baseline: ArmStats,&lt;br&gt;
    canary: ArmStats,&lt;br&gt;
    budget_remaining: float,      # fraction of the SLO error budget left&lt;br&gt;
    min_samples: int,&lt;br&gt;
    hard_abort_rate: float = 0.05,&lt;br&gt;
) -&amp;gt; Verdict:&lt;br&gt;
    # Hard circuit breaker, no statistics required. A canary returning 5%&lt;br&gt;
    # errors is broken; do not wait for significance.&lt;br&gt;
    if canary.requests &amp;gt;= 50 and canary.rate &amp;gt;= hard_abort_rate:&lt;br&gt;
        return Verdict("abort", f"canary error rate {canary.rate:.1%} exceeds hard limit")&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;if canary.requests &amp;lt; min_samples:
    return Verdict("hold", f"{canary.requests}/{min_samples} samples")

z = two_proportion_z(baseline, canary)
if z &amp;gt; 1.96:
    return Verdict("abort", f"canary significantly worse (z={z:.2f})", 0.95)

# Error budget governs how much risk you may take, not the canary itself.
# Nearly exhausted budget means small steps and long bakes.
if budget_remaining &amp;lt; 0.1:
    return Verdict("hold", "error budget nearly exhausted; deploys frozen")

if canary.latency_p99_ms &amp;gt; baseline.latency_p99_ms * 1.25:
    return Verdict("abort", "p99 latency regression &amp;gt;25%")

return Verdict("promote", f"no significant regression (z={z:.2f})", 0.95)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Why baseline comparison, not an absolute threshold&lt;br&gt;
Your error rate is not constant. It varies with time of day, with a flaky upstream, with a bot crawl. An absolute 1% threshold aborts a perfectly good canary during a period when the baseline is at 1.2%, and it passes a bad canary during a quiet period when the baseline is 0.05%.&lt;br&gt;
Comparing arms at the same moment cancels all of that. Both versions see the same upstream, the same traffic mix, the same hour. That is the entire reason to run a canary rather than deploying to staging.&lt;br&gt;
Traffic splitting has to be sticky&lt;br&gt;
A user who hits the canary on request one and the baseline on request two experiences a version-mismatch bug you will never reproduce. Split on a stable hash of a session or user identifier, not per request:&lt;br&gt;
import hashlib&lt;/p&gt;

&lt;p&gt;def in_canary(session_id: str, percentage: float) -&amp;gt; bool:&lt;br&gt;
    """Stable per session. Also means increasing the percentage only ever&lt;br&gt;
    ADDS sessions to the canary - nobody gets moved back."""&lt;br&gt;
    h = hashlib.blake2b(session_id.encode(), digest_size=8).digest()&lt;br&gt;
    return (int.from_bytes(h, "big") % 10_000) &amp;lt; percentage * 100&lt;br&gt;
The monotonicity property matters. With a naive random split, raising the canary from 5% to 10% reshuffles everyone; with a stable hash, the original 5% stay put and 5% more join. That makes the comparison valid across steps and makes rollback clean.&lt;br&gt;
What the error budget is actually for&lt;br&gt;
The budget does not decide whether this canary is healthy — the statistics do. It decides how much risk you are permitted to take right now.&lt;br&gt;
Full budget: 20% steps, short bakes, automatic promotion. Half consumed: 5% steps, longer bakes, a human approves promotion. Under 10%: no deploys except fixes. That policy converts the SLO from a dashboard into something that changes behaviour, which is the only reason to have one.&lt;br&gt;
If your traffic is too low&lt;br&gt;
Accept it. At 20 requests per second you cannot statistically canary a 0.5pp regression in a useful timeframe, and no amount of dashboard sophistication changes that.&lt;br&gt;
What works instead: shadow traffic (mirror production requests to the new version and compare responses without serving them), heavier pre-production testing, and feature flags that let you turn off behaviour without redeploying. Those are the tools for low-volume services, and reaching for a canary there produces a gate that looks rigorous and decides nothing.&lt;/p&gt;

&lt;p&gt;We build delivery pipelines and platform tooling at SoluLab — more on our &lt;a href="https://www.solulab.com/software-consulting-company/" rel="noopener noreferrer"&gt;software consulting&lt;/a&gt; work.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>sre</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
