DEV Community

ArthurFinley2291
ArthurFinley2291

Posted on

Feature Flag Gradual Rollout Explained: 6 Percentage Canary Controls

A feature flag gradual rollout by percentage needs stable account assignment and a compact decision record for every notification attempt; for a Node.js backend API or any other runtime, that is the least complex canary release design that lets an operator reconstruct which rule exposed a SaaS user, what the service attempted, and whether a retry could send a duplicate.

Short answer: a percentage is only an allocation rule. A defensible gradual rollout also needs deterministic assignment, an immutable configuration version, an idempotency key, outcome telemetry, explicit promotion criteria, and a bounded retention policy. The example below implements those six controls in Go; the same boundary works when the surrounding API happens to be written in Node.js, because assignment is a language-independent data contract rather than a framework feature.

What does incident reconstruction actually cost?

The dominant storage term is usually the number of delivery attempts, not the number of flags. If a notification service processes N attempts per day, retains each compact record for D days, and stores B bytes after encoding and indexing, its approximate hot footprint is N × D × B, before replication and index overhead. Payload bodies, stack traces, and unbounded attributes make B grow quickly; raising a canary from 1% to 10% does not help if the service records every full request twice.

The useful change is to split evidence by purpose. Keep a small, structured decision record for every attempt, while sampling verbose diagnostic detail according to a separate policy. A decision record needs identifiers and outcomes: delivery ID, account bucket key or a privacy-preserving derivative, flag key, configuration version, assigned variant, attempt number, provider outcome class, and timestamps. It does not need the notification body, access token, or a raw player profile.

This distinction matters under compliance constraints. The OWASP Logging Cheat Sheet says logs should generally exclude access tokens, authentication passwords, sensitive personal data, and payment data; NIST SP 800-92 treats retention as an organizational policy question rather than prescribing one universal duration. The retention period therefore comes from incident response, legal, and privacy requirements, then the storage calculation follows from that decision.

Keep the arithmetic visible. For example, an operator can calculate capacity with an observed encoded record size and measured daily attempt count, but those values belong in a capacity worksheet, not as universal constants in application code.

I would deliberately stop keeping message bodies and successful-provider response bodies in the general event stream. The cost is real: if the provider accepted malformed content, the compact record can prove what configuration and template revision were selected, but it cannot reproduce the rendered message byte for byte. A narrowly controlled, shorter-lived payload archive may be justified for that investigation; retaining every body indefinitely is not.

How should a Node.js backend API run a gradual percentage rollout?

Because assignment and delivery are different state machines. A flag evaluator may consistently put an account into the 5% cohort while two workers concurrently process the same queue item. Both workers can make the same correct flag decision and still send the same tournament reminder twice.

Exactly-once delivery across a database, a queue, and an external notification provider is not created by a boolean flag. The practical target is an exactly-once effect within a defined boundary: derive an idempotency key from the immutable logical notification ID and channel, claim that key atomically before the external side effect, and make retries reuse it. If the provider supports its own idempotency contract, pass a stable key there as a second line of defense; the local ledger remains necessary for audit and reconciliation.

Do not bucket on a process-local random number. Do not bucket on a mutable email address. A stable account or tenant identifier prevents users from moving in and out of exposure on each request, while a salt that includes the flag key prevents unrelated flags from producing perfectly correlated cohorts. Record the configuration version as well as the variant. Without it, a later incident review cannot distinguish “the account evaluated differently” from “the rule changed between attempts.”

Here is a minimal evaluator. It uses the first eight bytes of SHA-256 as an unsigned integer and maps that value into 10,000 basis-point buckets, giving percentage changes a resolution of 0.01%. The algorithm and text encoding must be specified identically in every implementation that evaluates the same flag.

package rollout

import (
    "crypto/sha256"
    "encoding/binary"
)

type Decision struct {
    FlagKey       string
    ConfigVersion string
    SubjectID     string
    Bucket        uint16
    Enabled       bool
}

func Evaluate(flagKey, configVersion, subjectID string, basisPoints uint16) Decision {
    if basisPoints > 10_000 {
        basisPoints = 10_000
    }

    sum := sha256.Sum256([]byte(flagKey + "\x00" + subjectID))
    bucket := uint16(binary.BigEndian.Uint64(sum[:8]) % 10_000)

    return Decision{
        FlagKey:       flagKey,
        ConfigVersion: configVersion,
        SubjectID:     subjectID,
        Bucket:        bucket,
        Enabled:       bucket < basisPoints,
    }
}
Enter fullscreen mode Exit fullscreen mode

The null separator removes an avoidable ambiguity between concatenated fields. More important, the evaluator returns evidence rather than only true: the bucket and configuration version can travel with the delivery command and later enter the audit record.

The event is a ledger entry, not a debug sentence

Free-form log lines are poor reconciliation records. They are hard to validate, their field names drift, and a message such as “canary enabled” says nothing about which rule revision produced that answer. OpenTelemetry's log data model provides fields for timestamps, trace and span identifiers, severity, body, resource, attributes, and an observed timestamp; its trace semantic conventions also define common naming rules. Those standards are useful boundaries even if the storage system is a plain append-only table.

For notification delivery, use a schema that can answer a join without parsing prose. The following interface intentionally separates the atomic idempotency claim from append-only evidence. Its persistence implementation could be a relational transaction or another store with equivalent conditional-write semantics.

package delivery

import (
    "context"
    "time"
)

type Attempt struct {
    DeliveryID    string
    IdempotencyKey string
    AccountID     string
    FlagKey       string
    ConfigVersion string
    Variant       string
    Bucket        uint16
    Attempt       int
    Outcome       string
    OccurredAt    time.Time
}

type Ledger interface {
    Claim(ctx context.Context, idempotencyKey string) (claimed bool, err error)
    Append(ctx context.Context, attempt Attempt) error
}
Enter fullscreen mode Exit fullscreen mode

There is an uncomfortable failure window: the provider may accept a notification and the service may crash before appending the accepted outcome. A local idempotency claim prevents another local worker from racing, but it cannot prove the external side effect. Resolve that uncertainty through provider-side idempotency where available, delivery receipts, or a reconciliation job that compares pending ledger entries with provider evidence. Mark the state unknown until reconciliation; do not silently translate uncertainty into failure and resend.

That single state is valuable. Unknown is honest.

Audit integrity also depends on write permissions and clocks. Application workers should be able to append attempts but not rewrite historical outcomes; corrective events should refer to prior event IDs. Record both event time and ingestion time so clock skew and delayed queues do not rewrite the apparent sequence. Protect the log store from unauthorized access and modification, as both OWASP guidance and NIST log-management guidance require.

Promotion is an evidence decision

A gradual release is a controlled comparison, not a timer that eventually reaches 100%. Before changing exposure, define an observation window and compare canary with control on outcomes that represent the user-visible delivery path. In a gaming notification service, useful measures include accepted, rejected, rate-limited, timed out, and confirmed delivery outcomes, partitioned by channel and configuration version. Queue age and retry count expose slower failure modes that an aggregate success ratio can hide.

Use minimum sample requirements and uncertainty bounds appropriate to the decision. A cohort with three deliveries and zero errors does not establish safety. Nor does a global average protect a small region or channel. Promotion rules should state which dimensions may block advancement, who approves an override, and which audit event records that override.

The operational sequence can stay short:

  1. Validate the evaluator against fixed test vectors in every runtime that consumes the contract.
  2. Start with internal or explicitly selected accounts, then open a small percentage cohort.
  3. Compare canary and control over the declared window; inspect unknown outcomes and retry amplification separately.
  4. Increase exposure by changing a versioned configuration, never by editing historical events.
  5. On rollback, set new evaluations to the control variant while allowing already claimed deliveries to reconcile under their recorded version.

This is where the audit trail earns its storage. During an incident, operators can group failed attempts by configuration version and bucket, find the first exposure change preceding the divergence, and distinguish evaluation errors from downstream delivery failures. A trace ID can connect the decision to queue handling and provider calls, but it should supplement, not replace, the durable delivery ID. The trade-off is operational weight: the team must own schema evolution, access control, retention, reconciliation, and index capacity. This design does not fit a low-volume, reversible presentation change whose outcome has no external side effect and can be understood from ordinary request metrics; a simple deterministic evaluator plus aggregate counters is enough there. It also has a limitation for anonymous traffic because a durable subject key may not exist, in which case a short-lived session assignment offers weaker reconstruction and must not be presented as account-level evidence.

Testing the boundary before production

Hash functions are deterministic, yet rollout systems still fail through inconsistent input normalization, integer conversion, and configuration races. Publish test vectors containing the flag key, subject ID, expected bucket, and expected decision at selected thresholds. Test boundary values at 0 and 10,000 basis points. Run distribution tests over a large synthetic identifier set, but do not mistake statistical balance for a contract test: exact vectors catch cross-language disagreement. Then test the failure windows. Start two workers against the same idempotency key and assert that only one claims it. Inject a timeout after an external acceptance and verify that the attempt becomes unknown rather than immediately resent. Replay an old queue message after a configuration promotion and verify that the recorded decision remains tied to its original version. Finally, test rollback while requests are in flight. Deployment should make evaluator changes independently reviewable from percentage changes. A new hash algorithm remaps the whole population even when the configured percentage is unchanged, so such a change requires a migration plan, dual evaluation, and explicit versioning. A percentage adjustment within an unchanged algorithm preserves the nesting property: accounts in the smaller cohort remain in the larger one.

A compact decision rule

Use percentage allocation when subjects need stable, broadly representative exposure and the service can measure outcomes by cohort. Use an allowlist when the blast radius must be tied to known accounts. Combine them when internal accounts should enter first and a deterministic public cohort should follow.

For incident reconstruction, reject any design that cannot answer these questions from durable records: Which subject and bucket were evaluated? Which immutable configuration produced the result? Which logical notification was claimed? Which external outcome is known, and which remains uncertain? Which later event reconciled it?

The best implementation is the smallest one that preserves those answers. Retain compact decision and outcome records for the policy-defined investigation window, keep sensitive payloads out of the general stream, and treat removal of old evidence as an explicit loss of reconstruction depth. Percentage rollout then becomes a controlled state transition with an audit trail, rather than a random gate attached to a send call.

Further reading

Top comments (0)