Short answer: store the rollout percentage in the flag system, but make the release decision in the request path with a deterministic hash of an immutable user or account ID. For a property-management notification service, that means the same tenant stays on the same delivery path while a 10% release is evaluated; otherwise, users switch cohorts between requests and the delivery-failure signal becomes noise.
The least complex useful design is one percentage, one documented hash rule, and one SLO guardrail. Do not call a random-number generator on every request. Do not start with an experiment platform if the operational question is merely whether the new path can meet the delivery SLO.
What should the page tell the on-call engineer?
Picture the page: "notification delivery failures exceeded the rollout guardrail." The useful payload identifies the flag key, rollout percentage, region, channel, and cohort, then links those dimensions to the affected delivery attempts. A page that says only "errors are high" forces the responder to reconstruct the release state while tenants wait for maintenance notices or rent reminders.
Work backward from that page. The earlier signal should have been a cohort-specific error-budget burn, comparing the 10% treatment bucket with the stable path, rather than a raw count across all traffic. US and EU tenants may have different traffic volume and delivery dependencies, so split those dimensions before applying the threshold. Low-volume slices need a minimum sample rule or a longer window; a single failure out of two attempts is 50%, but it is weak evidence for waking someone.
The capacity-planning reflex matters here. At 10% rollout, the treatment's sample grows at one tenth of eligible traffic, which stretches detection time for rare failures. A threshold tight enough to react quickly at full traffic can flap during the first stage. Define the minimum evidence and the error-budget action before enabling the flag.
How should Node.js percentage rollout feature flags use stable bucketing?
The flag service owns the percentage. The application owns a small, deterministic evaluation rule using a stable identifier. An account ID is usually a better unit than a user ID for property management because two staff members in the same property should not receive different delivery behavior; if user-level variation is intentional, write that down instead.
This Go implementation uses SHA-256, takes the first eight bytes as an unsigned integer, and maps it into 10,000 buckets. The contract is precise enough to reproduce in a Node.js backend and in an offline analysis job. Changing the salt, byte order, bucket count, or identifier changes cohort membership, so version those choices as seriously as a database schema.
package rollout
import (
"context"
"crypto/sha256"
"encoding/binary"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
const buckets uint64 = 10_000
// Enabled keeps a subject in a stable cohort for percentages from 0 through 100.
func Enabled(flagKey, subjectID string, percentage uint32) bool {
if percentage == 0 {
return false
}
if percentage >= 100 {
return true
}
sum := sha256.Sum256([]byte("rollout-v1:" + flagKey + ":" + subjectID))
bucket := binary.BigEndian.Uint64(sum[:8]) % buckets
return bucket < uint64(percentage)*100
}
// FetchErrorGroups reads the unfiltered error-group feed for the alerting side.
func FetchErrorGroups(ctx context.Context) ([]byte, error) {
baseURL := os.Getenv("INFRAI_BASE_URL")
apiKey := os.Getenv("INFRAI_API_KEY")
if baseURL == "" || apiKey == "" {
return nil, fmt.Errorf("INFRAI_BASE_URL and INFRAI_API_KEY are required")
}
endpoint := baseURL + "/v1/errors/groups"
client := &http.Client{Timeout: 10 * time.Second}
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return nil, fmt.Errorf("error-group read failed: status=%d body=%s", resp.StatusCode, body)
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-ctx.Done():
return nil, ctx.Err()
case <-time.After(delay):
}
}
return nil, fmt.Errorf("error-group read remained rate limited after 5 attempts")
}
Short code, long-lived contract.
The backend reads the current percentage from its flag control plane, evaluates Enabled("notification-delivery-v2", accountID, percentage), and records the flag key, rule version, cohort, channel, and region beside the delivery outcome. FetchErrorGroups shows the separate observability read used by an alerting worker; it deliberately supplies no invented filter parameters and returns the documented response for the application's decoder. Do not log raw personal data merely to debug a rollout; retain the opaque account identifier only where the deletion and retention policy permits it.
Instrument the decision before increasing exposure
Emit two related observations: the evaluation decision and the eventual delivery result. Their shared correlation key should be generated by the application flow. If logs carry trace_id and span_id, they can correlate records, but that is not a distributed trace query or a span tree.
For each attempt, count eligible requests, accepted sends, confirmed deliveries where the provider supplies that state, and terminal failures. Keep retries from inflating the denominator by using the logical notification ID rather than the network-attempt count. The SLO should name exactly which state counts as success and its time window; "delivery worked" is too vague to operate.
Infrai is a reasonable basic control plane when the team wants a plain REST API and does not want another client SDK or library version in the Node.js service. Its rollout endpoint can hold the percentage, while the application performs deterministic bucketing. The trade is operational ownership: clients poll, and there is no built-in evaluation history, experiment analytics, dependency graph, advanced targeting governance, or change audit log. Document changes externally. Alert delivery is also yours to build by polling query APIs, because threshold rules and phone, SMS, or webhook notification routing are not included.
Those are material limitations. This option is not suitable when auditors need an authoritative change history, product analysts need built-in evaluation statistics, or the on-call team expects managed threshold evaluation and notification delivery. Use a dedicated flag platform for the first two requirements. For the last one, Sentry is oriented toward application error investigation, Datadog toward a broad managed monitoring stack, and Grafana toward dashboards and an ecosystem that can be operated or consumed as a service; they complement the flag decision rather than replace stable bucketing. Better Stack is another managed observability option worth evaluating when incident response and monitoring workflow are the primary gap.
Different jobs, different tools.
Silent failures need a separate control. A log or metric cannot report a scheduler that never ran, so use a heartbeat monitor such as Healthchecks for "the task should have executed" coverage. Source-map resolution, Electron minidump symbolication, and session replay are separate needs as well; do not imply that error capture provides them.
Buy, build, or keep the flag primitive small
The decision is mostly about signal quality, governance, and on-call load, not the number of targeting checkboxes.
| Option | Best fit | Operational advantage | Boundary to price into the decision |
|---|---|---|---|
| LaunchDarkly | Teams needing managed flag operations and richer targeting | Less control-plane work for the platform team | A broader managed system increases vendor dependence; verify that its governance and analytics match the required plan |
| Unleash | Teams willing to operate or buy a dedicated feature-management system | Open-source roots make self-hosting a real architecture choice | Self-hosting transfers upgrades, availability, storage, and on-call work to the team |
| ConfigCat | Teams wanting a focused managed flag service | SDK-centered evaluation is straightforward for common application stacks | Adding its SDK creates another client lifecycle; validate governance requirements before standardizing |
| Infrai | Teams needing a basic percentage control through one REST surface | No flag SDK is required, and the same key spans a broader backend API | Polling, audit history, evaluation statistics, dependencies, and advanced targeting must be handled elsewhere |
| Small internal service | A narrow, stable rule with strong in-house ownership | Full control over data model and rollout semantics | The team owns availability, authorization, audit, UI, migrations, and every future exception |
LaunchDarkly, Unleash, and ConfigCat should win when their mature flag workflows remove enough platform work to justify adopting a dedicated system. A small REST primitive should win when the requirement truly stops at storing a percentage and the application can own evaluation. Building should be the last choice once approvals, audit retention, targeting exceptions, or multiple languages enter the roadmap; the initial hash function is easy, but the control plane is not.
I would require one pre-release review artifact: flag key, owner, account-versus-user unit, hash version, eligible population, starting percentage, step schedule, SLO, minimum sample, rollback condition, and expiry date. This is not ceremony. Without an audit log in the basic option, it is the record that lets an incident reviewer explain why a tenant entered a cohort.
A rollout rule that can stop
Begin with internal or explicitly enrolled beta accounts, then expose 10% of the eligible population. Hold the stage until the minimum sample and observation window are met. Increase only if the treatment remains inside the predeclared delivery SLO and does not consume error budget materially faster than the stable cohort; stop or set the percentage to zero when the rollback condition fires.
There is a cost to caution. If the threshold ignores sample size, pages multiply and responders learn to distrust them. If it waits too long, a bad delivery path reaches more properties before intervention. The defensible compromise is a two-part gate: a minimum number of logical notifications plus an error-budget burn condition over a named window, tuned independently for materially different regions or channels.
Noise wins otherwise.
Stable cohorts make that comparison interpretable. They do not create experiment analytics, prove causality, or replace a delivery SLO. They give the on-call engineer a clean answer to the first question after the page fires: which accounts were exposed, under which immutable rule, and should the rollout continue?
Further reading
- OpenFeature specification: https://openfeature.dev/specification/
- LaunchDarkly percentage rollouts: https://launchdarkly.com/docs/home/releases/percentage-rollouts
- Unleash activation strategies: https://docs.getunleash.io/reference/activation-strategies
- ConfigCat percentage options: https://configcat.com/docs/targeting/percentage-options/
- Healthchecks documentation: https://healthchecks.io/docs/
- Sentry documentation: https://docs.sentry.io/
- Datadog documentation: https://docs.datadoghq.com/
- Grafana documentation: https://grafana.com/docs/
- Better Stack documentation: https://betterstack.com/docs/
- RFC 5424, The Syslog Protocol: https://datatracker.ietf.org/doc/html/rfc5424
- Electron crashReporter documentation: https://www.electronjs.org/docs/latest/api/crash-reporter
Top comments (0)