A React or Next.js pricing rollout has a stricter requirement than "the browser eventually sees the feature flags." When a shopper reports that a cart total changed, the team must be able to reconstruct the exact decision the shopper received at that moment. My choice is a small frontend polling client backed by versioned, non-secret snapshots, plus one decision event at the pricing boundary. The flag value alone isn't enough.
TL;DR: poll a compact public configuration document, retain the last valid snapshot during transient failures, assign shoppers with a deterministic key, and record the snapshot version alongside the pricing result. Evaluate the design by replaying incidents, not by admiring how few lines the hook contains.
The tempting first pass is a timer that fetches {"new_price_rule": true} every few seconds and stores the feature toggle in component state. It demos beautifully in a notebook-sized prototype. It also throws away the evidence needed to answer the first serious production question: which rule, configuration revision, and assignment produced this displayed price? A free or cheap delivery mechanism doesn't change that requirement.
What must be captured to replay a pricing decision?
A useful snapshot is an immutable input to a decision, not a bag of mutable booleans. For this rollout, I would give every published configuration a snapshot_id, an issue time, and a rule payload. The browser decision event then joins five values: an opaque subject key, flag key, variant, snapshot ID, and pricing outcome. Keep the event narrow. A trace does not need the shopper's email, full cart, or prompt transcript to explain a percentage rule.
This distinction matters because observability systems commonly group related errors rather than preserving a magical, complete story automatically. Sentry documents that grouping uses fingerprints and that an event can override its fingerprint. That mechanism is useful for grouping repeated evaluation failures, but grouping identity and business-decision identity solve different problems. A fingerprint such as pricing-flag-evaluation can collect a class of failures; snapshot_id=pricing-2026-04 identifies the configuration involved in one decision. Conflating the two makes search convenient and reconstruction vague.
For an e-commerce rollout, the event shape can stay modest:
{
"event": "pricing_rule_evaluated",
"flag_key": "pricing_rule",
"variant": "candidate",
"snapshot_id": "pricing-2026-04",
"subject_key": "opaque_7f3a",
"result_code": "applied"
}
The sample values are illustrative, not a production schema. The important constraint is referential: the snapshot ID must resolve to the exact rule inputs used by the client. If publishing replaces data in place, the ID is decoration.
That join matters.
Record decisions, not just fetches. A successful poll says configuration arrived. It does not say which branch rendered a price or whether the shopper was reassigned after hydration.
How should a React or Next.js polling client evaluate feature flags?
A browser polling client has four meaningful inputs: the last valid snapshot, a newly fetched candidate, the current clock, and an opaque assignment key. Treating those inputs explicitly makes the React or Next.js implementation almost boring, which is a good outcome. The rendering layer should subscribe to resolved decisions; it shouldn't contain retry policy, parsing, or bucketing math.
The following Python reference model is deliberately small. I use this kind of executable model to pin down semantics before translating them into a UI hook. It rejects malformed candidates, ignores older revisions, and keeps the last valid state when a poll fails. No network library or vendor SDK is implied.
from dataclasses import dataclass
from typing import Any, Mapping, Optional
@dataclass(frozen=True)
class Snapshot:
snapshot_id: str
revision: int
issued_at: str
flags: Mapping[str, Mapping[str, Any]]
@dataclass(frozen=True)
class PollState:
snapshot: Optional[Snapshot]
last_attempt_ms: int
error_code: Optional[str]
def accept_candidate(
state: PollState, candidate: Mapping[str, Any], now_ms: int
) -> PollState:
try:
incoming = Snapshot(
snapshot_id=str(candidate["snapshot_id"]),
revision=int(candidate["revision"]),
issued_at=str(candidate["issued_at"]),
flags=dict(candidate["flags"]),
)
except (KeyError, TypeError, ValueError):
return PollState(state.snapshot, now_ms, "invalid_snapshot")
if state.snapshot and incoming.revision <= state.snapshot.revision:
return PollState(state.snapshot, now_ms, None)
return PollState(incoming, now_ms, None)
def retain_after_fetch_error(state: PollState, now_ms: int) -> PollState:
return PollState(state.snapshot, now_ms, "fetch_failed")
There is a sharp trade-off here. Retaining a known snapshot avoids changing the rendered price merely because one request failed, but it also means freshness must be bounded by policy. The configuration should therefore carry enough information for the application to decide when an old snapshot is no longer acceptable. What happens after that bound is product policy: preserve the established price, use a documented baseline rule, or stop presenting the affected offer. The team must choose before launch.
Short code is not the goal. Stable semantics are.
The client should also avoid overlapping polls. Schedule the next attempt after the current attempt completes, add bounded randomization so clients do not move in lockstep, and cancel work when its owner is disposed. Those are implementation decisions rather than pricing facts, so they belong in tests as observable behavior. I would test them with a fake clock and scripted fetch outcomes instead of waiting on real timers.
Stable assignment prevents phantom price changes
A percentage rollout needs a repeatable mapping from an opaque subject key to a bucket. Random choice on every render is unusable because the same shopper can cross branches without any configuration change. A stable hash creates a reproducible assignment, provided the hash input and algorithm version are part of the contract.
import hashlib
def bucket(subject_key: str, flag_key: str, algorithm_version: str) -> int:
material = f"{algorithm_version}:{flag_key}:{subject_key}".encode("utf-8")
digest = hashlib.sha256(material).digest()
return int.from_bytes(digest[:8], "big") % 10_000
def choose_variant(
subject_key: str, flag_key: str, rollout_basis_points: int
) -> str:
if not 0 <= rollout_basis_points <= 10_000:
raise ValueError("rollout_basis_points must be between 0 and 10,000")
assigned = bucket(subject_key, flag_key, "v1")
return "candidate" if assigned < rollout_basis_points else "baseline"
Ten thousand buckets make percentages expressible as basis points in this example; they do not promise any particular statistical result. Before copying that number, run the actual assignment population through the function and inspect balance, repeatability, and behavior when keys are missing. Changing the salt, input order, hash, or algorithm version can move shoppers. Such a change deserves its own migration plan and telemetry.
A server-rendered page adds another constraint: the initial decision and the hydrated client decision must use the same snapshot and subject key. Otherwise the displayed price can flip during hydration even though both evaluators are individually correct. Pass the initial decision metadata with the rendered data, then let polling replace it only after the client has accepted a newer valid snapshot.
No guesswork.
Do not put secrets in the public snapshot. Browser-delivered rules and keys are inspectable by the shopper. The backend must remain authoritative for the amount charged, while the frontend decision controls presentation and experiment exposure. This boundary also gives incident review two records to compare: what the browser presented and what the order service accepted.
Make the event stream useful without making it permanent
Incident reconstruction encourages retention; privacy law constrains it. GDPR Article 17 describes a right to erasure and also lists circumstances in which that right does not apply. The engineering consequence is not a universal retention number. It is a requirement to define purpose, deletion handling, and the relationship between an opaque subject key and any identifiable account data with counsel and the data owner.
Keep high-cardinality values out of error grouping keys. Group evaluation failures by a stable failure class and attach snapshot ID as searchable context. Put the decision event in the analytics or audit path designed for that cardinality. Sampling deserves special care: an error sample without its decision event, or a decision sample without the associated pricing result, weakens the reconstruction. Decide what must be joined first, then configure sampling around that unit.
I also separate operational payloads from eval fixtures. The former should be minimal and governed; the latter can be synthetic rows such as expired snapshot, reordered response, missing subject key, and hydration mismatch. This keeps the eval harness repeatable without copying shopper data into a notebook. It also keeps token-heavy AI analysis out of the request path. Summaries can be produced later from bounded, redacted fields if the team chooses to use a model during incident review.
Observability has a cost shape. Poll frequency multiplies request volume, decision events multiply ingestion volume, and unbounded context multiplies investigation noise. Measure all three. I would rather retain a compact join key at every decision than emit a giant payload occasionally and hope it contains the right clue.
What should you measure before copying this design?
Start with a replay drill. Given an order or support timestamp, can an engineer locate the client decision, resolve its immutable snapshot, reproduce the assignment, and compare the presented rule with the charged result? Record whether the drill succeeds and which join failed. This is the primary evaluation constraint.
Replay it cold.
Then exercise the ugly sequences: first load without configuration, malformed response after a valid response, older revision arriving late, repeated failure past the freshness bound, subject key appearing after hydration, and a rollout percentage changing between polls. Assert both the rendered variant and the emitted event. A test that checks only UI text misses the evidence trail; a test that checks only telemetry can bless the wrong price.
The operating dashboard needs rates and age distributions rather than one triumphant availability number. Track accepted snapshot age, invalid snapshot count, fetch failure count, assignment changes per subject, hydration disagreements, decision events missing snapshot IDs, and pricing outcomes that cannot be joined to a client decision. Alert thresholds depend on traffic and business tolerance, so derive them from a baseline rather than importing someone else's percentages.
Finally, load-test the polling schedule with the expected active-client distribution. Compare a shorter interval with the incident benefit it actually provides. Faster refresh can shrink configuration propagation time while increasing requests and synchronized bursts; it does not improve reconstruction if snapshots are mutable or decision events lack IDs. That is why price is not the deciding axis. The useful question is how much evidence and propagation control the system buys for its operational load.
A simple client can be the right choice. Keep its contract precise: versioned public snapshots, deterministic assignment, last-valid-state behavior, an explicit freshness policy, and a compact decision record that joins browser presentation to backend outcome. Ship the flag only after the replay drill works.
Top comments (0)