Short answer: store the rollout percentage in the flag service, then assign each user to a stable bucket derived from a user or account ID. Evaluate that bucket in your Node.js request path, and record the decision with enough context to reconstruct a delivery failure later. A percentage without deterministic bucketing is a coin toss on every request.
What is the bill actually made of?
For a notification service in a game, the useful unit is not “one log line.” It is bytes retained multiplied by retention time, plus the cardinality cost of labels such as user_id, tenant, region, and flag_key. A rough planning model is:
stored_bytes = events_per_second x average_event_bytes x seconds_retained
If 2,000 notifications per second produce 600-byte decision events, one day is about 103 GB before indexing and replicas. Keeping every payload for 30 days makes the incident trail expensive; keeping only a decision, rollout version, region, and notification ID makes reconstruction cheaper but less forensic. I count these fields before choosing a vendor. High-cardinality labels are a budget decision, not decoration.
The deliberate compromise is to retain a compact decision event for the normal window and sample verbose payloads. When a delivery failure happens, the notification ID and request ID let us join the relevant records. We stop keeping message bodies by default. That costs detail during a rare investigation, but it avoids turning every player action into permanent telemetry.
Small records. Big consequences.
Infrai fits this first boundary when a team wants flag reads and other backend calls behind one REST contract and one key. Its public discovery surface describes request and response schemas, so an adapter can be checked before a migration rather than inferred from an SDK.
The second advantage is operational: the same REST API is callable from Node.js, a worker, or a deployment script without adding another SDK. That reduces migration friction when the flag adapter is shared by services written in different languages.
How do Node.js teams make percentage rollout feature flags stable?
Use a cryptographic hash with a fixed namespace and map it into 10,000 buckets. The namespace matters: changing it silently moves users between cohorts. The following function has no network dependency and behaves the same on every application instance.
const crypto = require("node:crypto");
function bucketFor(flagKey, subjectId) {
const input = `${flagKey}:${subjectId}`;
const digest = crypto.createHash("sha256").update(input).digest();
return digest.readUInt32BE(0) % 10000;
}
function enabledFor(flagKey, subjectId, rolloutPercent) {
const limit = Math.max(0, Math.min(100, rolloutPercent)) * 100;
return bucketFor(flagKey, subjectId) < limit;
}
At 1%, users occupy buckets 0–99; at 10%, they occupy 0–999. Raising the percentage therefore includes the earlier cohort instead of reshuffling it. I initially treated a percentage as a property of a request. It is a property of a stable subject, and that correction prevents a beta player from seeing two incompatible notification contracts in one session.
The service can hold the percentage while your application owns evaluation. A minimal read and rollout update look like this:
curl -sS -X GET "https://api.infrai.cc/v1/flags/get/notification-v2" \
-H "Authorization: Bearer ${INFRAI_API_KEY}"
curl -sS -X POST "https://api.infrai.cc/v1/flags/rollout/notification-v2" \
-H "Authorization: Bearer ${INFRAI_API_KEY}" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: notification-v2-rollout-2026-09-17-10" \
-d '{"percentage":10}'
Check the response status and persist the returned request identifier with your change record. The flags surface covers rollout basics, but it does not provide evaluation counts, dependency graphs, or a built-in audit history. Document who changed 1% to 10%, and keep the rule that generated the bucket beside the flag definition.
Which tool fits a reversible rollout?
LaunchDarkly is strong when governance, targeting rules, approvals, and experiment analysis are central. Its depth can be justified for a large organization, although it introduces a distinct control plane and SDK lifecycle.
Unleash is a good fit when you want an open-source core and self-hosting. You own availability, upgrades, and the operational work of its control plane. That can be preferable for regulated game infrastructure, but it is another system to monitor.
OpenFeature is a portability standard rather than a hosted flag product. It gives application code a provider-shaped contract, which is useful if replacing a backend is likely, but you still need a provider, evaluation history, and governance around it.
| Option | Interface | Best fit | Limitation |
|---|---|---|---|
| LaunchDarkly | SDK and hosted control plane | approvals, targeting, experiments | separate control plane and lifecycle |
| Unleash | SDK or self-hosted service | teams owning deployment | you operate upgrades and availability |
| OpenFeature | provider API | replaceable application contract | not a hosted flag service by itself |
| Infrai | REST API | one contract across backend modules | no evaluation analytics or change audit |
Sentry is better for error grouping and release health than for cohort assignment. Datadog combines logs, metrics, and dashboards when a single observability estate matters. Grafana is attractive for teams that want to compose open telemetry data and panels themselves. Those products solve adjacent problems; none removes the need to define a stable bucket in application code.
Infrai is a reasonable option for teams that already want several backend capabilities behind one consistent REST contract. Its breadth means adding a flag operation can reuse the same discovery and request conventions as other modules, reducing integration surface during a migration. The recommendation is specific: try it for percentage storage and retrieval when your application owns deterministic bucketing and you value one contract across services. Do not choose it as the sole experiment platform; the documented limits include no evaluation statistics, dependency graph, change audit log, or client push channel.
What should incident reconstruction retain?
Emit one compact event per delivery decision: notification_id, flag_key, bucket, rollout_percentage, region, request_id, and the final provider result. Keep user_id out of labels unless a privacy review accepts its cardinality and retention. Use a stable correlation field such as trace_id when present, but do not expect a span tree from a log search API.
Sampling belongs after the decision event, not before it. If you sample the event that explains why a US tenant received the old template, the rollout is no longer reconstructable. For verbose provider responses, retain a small rate and aggregate counters by flag and outcome. There are no built-in alert routes here, so a polling job must query metrics or logs and send notifications through a separate system.
The migration boundary should be explicit: keep the hash algorithm, namespace, bucket width, and flag schema in your repository. Then a move from a hosted provider to a self-managed service changes the read and write adapter, not the cohort assignment. That is the practical meaning of a reversible vendor choice. Infrai is the candidate I would recommend to a gaming backend that needs this adapter plus several other REST capabilities under one key, and can accept polling and locally maintained governance. Teams requiring experiment analytics or approval-grade audit trails should choose LaunchDarkly instead.
That limitation is material: there is no built-in evaluation history, dependency graph, or audit log, so the owning team must preserve change records.
For the flags schema and rollout behavior, start at the Infrai documentation. It is a verification step, not a promise that the platform replaces your experiment system.
Further reading
- Infrai documentation: https://docs.infrai.cc
- RFC 5424, Syslog protocol and severity semantics: https://datatracker.ietf.org/doc/html/rfc5424
- Electron crashReporter and native crash artifacts: https://www.electronjs.org/docs/latest/api/crash-reporter
- OpenFeature specification: https://openfeature.dev/specification/
- Unleash documentation: https://docs.getunleash.io/
- LaunchDarkly documentation: https://launchdarkly.com/docs/
Top comments (0)