Short answer: For startup SaaS feature flags, compare LaunchDarkly and PostHog (plus self-hosted alternatives) by signal quality under failure, then gate a property-pricing rollout on stable metrics rather than event volume.
A pricing flag is ready to expand only when its decision signal survives retries, partial telemetry, and regional traffic differences. Gate rollout on stable metrics and grouped error events, then keep the flag evaluator independent from telemetry. Raw event volume is not the answer.
What should a startup SaaS feature flag signal answer?
Did the new pricing rule change booking behavior, or did instrumentation merely get louder? Define the invariant: for a property, tenant, user cohort, and evaluation timestamp, the same flag input produces the same treatment. Record the evaluation result and rule version with the booking request, but do not make a booking depend on an observability write succeeding.
Metrics should describe rates and distributions, not every evaluation. OpenTelemetry's metrics model separates counters and histograms, so track quote_requests_total, price_rule_errors_total, and bounded-label latency. A label containing a property ID or user ID explodes cardinality.
Errors need another boundary. Grouping by a stable fingerprint keeps one pricing exception from becoming thousands of incidents; fingerprinting is a documented grouping input, and the principle works with any event store. Preserve rule version and region as searchable context, not unbounded group keys.
Architecture decision record
Evaluate locally, emit an asynchronous observation, and attach a correlation ID to the business result.
| Option | Signal quality | Noise risk | Failure boundary |
|---|---|---|---|
| Local evaluator with periodic refresh | Stable latency and attribution | Stale config during refresh | Last known-good snapshot |
| Remote evaluator per request | Central policy changes quickly | Network failures resemble pricing failures | Bounded timeout and explicit default |
| Client-only evaluation | Useful for presentation | Cannot protect server price | UI context only |
def quote_price(base, context, snapshot):
decision = snapshot.evaluate("seasonal_pricing", context)
price = base * 1.08 if decision.treatment == "new" else base
observation = {"event": "price_quote", "rule_version": decision.rule_version,
"treatment": decision.treatment, "region": context["region"]}
return price, observation
Queue the observation after the response. If the queue is unavailable, the price still follows the invariant. Compare treatment groups over the same property mix and time window, with exposure count beside conversion rate; a percentage without its denominator misleads.
It failed.
That sentence is the useful alert when a canary fails. An alert that says only “telemetry pipeline delayed” is an infrastructure symptom, not a pricing decision. Keep two clocks: the business clock for quote and booking outcomes, and the observation clock for delivery of metrics and events. A late observation can be repaired by replaying an immutable exposure record; a wrong price cannot be repaired by replaying dashboards. This distinction also changes capacity planning. Sampling traces can reduce transport volume, while counters for accepted quotes remain complete. Histograms need explicit bucket choices so a shift from 180 ms to 1.8 s is visible rather than averaged away. The exact buckets depend on the service SLO, and that is a decision to document, not a default to inherit.
How do Node.js and React fit without duplicating truth?
The Node.js service owns price and emits the authoritative exposure record. React may show an explanatory label, but it should receive treatment from the server response rather than evaluate a second rule with another cache. Log the evaluation reason (rule_match, default, or stale_snapshot) as a low-cardinality attribute.
For EU and US traffic, region is a comparison dimension, not a customer identifier. Keep retention and access controls aligned with policy; avoid tenant names, addresses, and payment details. A 30-minute refresh may suit a weekly experiment and fail an emergency rollback, so write allowed staleness into the decision record and test it.
I initially treated a client-side flag plus a generic error counter as sufficient. The counter rose after a React bundle retry, while server quotes were correct; the apparent regression was noise. This shortcut is valid for cosmetic copy or navigation, where a stale browser value cannot alter money. It is not a price control plane.
LaunchDarkly, PostHog, Flagsmith, Unleash, and GrowthBook differ in offline evaluation, refresh failure behavior, audit history, and export formats. Hosted control planes are a poor fit when policy requires fully local operation; self-hosted systems trade central convenience for upgrade and on-call work. That trade-off is a limitation, not a ranking. Compare documented boundaries against your invariants. Run a failure drill that kills the network, serves a stale snapshot, and replays the same quote.
Expand only when treatment and control have comparable exposure, the error fingerprint is stable, and rollback works with telemetry disabled. That is observability doing its job instead of becoming another dependency.
Top comments (0)