DEV Community

FluxH91
FluxH91

Posted on

Cheap LaunchDarkly Alternative API Feature Flags for Rollback-Safe Percentage Rollouts

TL;DR: A cheap LaunchDarkly alternative with a simple API can handle feature flags for a Node.js or Next.js edtech application, but only if pricing evaluation stays on the server, percentage rollout uses deterministic account cohorts, and every policy decision is reconstructable. Rollback should mean selecting the previous policy version, not rewriting history or trusting a browser to calculate a price. Basic US and EU targeting is useful only if those invariants survive stale configuration, retries, and a control-plane outage.

That is the decision rule. Cost belongs in the comparison, but it is not the first filter: a low-cost flag that can silently quote two prices to the same school is an expensive ambiguity.

Should a cheap LaunchDarkly alternative use simple API feature flags?

The application scenario is narrow: an education platform is introducing a pricing rule behind a flag, starting with a percentage of eligible accounts in the US and EU. The dangerous interpretation is that a flag chooses which price to display. The safer interpretation is that a flag selects a versioned pricing policy, while the pricing service remains the sole authority for calculating and recording the quote.

Four invariants follow.

  1. The same stable account identifier, flag key, and allocation salt produce the same cohort until an operator deliberately changes the allocation.
  2. Region eligibility is derived from trusted account data on the server, not a client-supplied country string.
  3. A committed quote or purchase retains the policy version and monetary inputs used at that moment; disabling the experiment affects new decisions, not historical records.
  4. Every evaluation records the configuration version, result, reason, and correlation context needed to join it to the pricing request.

Stop there for a moment. None of those invariants requires a large feature-management suite. All of them require more than a boolean endpoint.

Rollback first.

The failure boundary should also be explicit. If configuration cannot be loaded or validated, a pricing request uses a previously validated local snapshot or the defined control policy. It must not alternate between policies according to network luck. If the account lacks a trustworthy region or stable identifier, it enters the control path. If logging is unavailable, the pricing decision may proceed only when the business has consciously accepted that observability failure mode; blocking every quote because a telemetry sink is slow couples revenue availability to diagnostics.

Rollback itself has two layers. The exposure layer stops assigning new requests to the candidate policy. The transaction layer preserves already committed quotes and orders. Confusing them creates a particularly ugly failure: the UI shows the old offer while a retry recalculates an in-flight transaction under the new rule.

Decision record and option boundaries

The choice is not “full platform or homemade toggle.” There are at least three operating models, and each moves responsibility rather than eliminating it.

Model Rollback path Consistency boundary Operational burden Appropriate use
Managed flag control plane Publish a prior version or disable the candidate rule SDK cache and configuration propagation Lower service operation, but policy governance still belongs to the team Many services, frequent changes, delegated approvals
Self-hosted flag service Revert versioned configuration under team control Service availability, cache behavior, and data-store semantics Patching, backup, capacity, and recovery are local obligations Data-control requirements and an established platform team
Application-owned evaluator Deploy or atomically select a signed, versioned policy document The application's snapshot loader and release process Small surface area, high responsibility for correctness and tooling Few flags, simple targeting, infrequent policy changes

These are not product rankings. A managed service can reduce operational work without resolving the application-level question of whether an accepted quote may be repriced. A self-hosted service increases control while placing durability and recovery tests on the owning team. An application-owned evaluator can be pleasantly small, yet its apparent simplicity disappears if every team invents a different hashing rule, audit schema, or approval mechanism. LaunchDarkly represents the managed-control-plane category; Unleash and Flagsmith document self-hosting paths as well as hosted offerings. Those names locate the operating models, not a recommendation: their relevant boundaries still come from SDK cache behavior, deployment topology, data handling, and the application's own transaction rules, all of which must be verified against the chosen edition and configuration before a pricing rollout.

I initially considered a request-time API lookup the cleanest design because it centralizes the answer. Following the rollback sequence exposed the mistake: centralization does not help when a quote retry observes a newer policy than the original request. For the stated scope, I would accept the application-owned model only if the organization can keep one evaluator library, one policy schema, and one promotion workflow; otherwise, a dedicated control plane is the more legible boundary. The deciding metric is rollback confidence under failure, not the number of dashboard features and certainly not an advertised entry price.

The evidence store deserves similar skepticism. Evaluation events can be buffered and written in batches, but a batch is not evidence merely because it reached object storage. Define the record format, partitioning, retention, access controls, and reconciliation process. Keep the mutable policy document separate from append-only decision records, and include a schema version so that a later parser does not have to guess what yesterday's fields meant.

That trade-off is deliberate.

The critical path in Python

The evaluator below illustrates the contract rather than a deployable service. It deliberately accepts validated configuration and trusted account attributes as inputs. SHA-256 supplies stable bytes for cohort assignment; this is deterministic allocation, not encryption of the account identifier and not a substitute for access control.

from __future__ import annotations

import hashlib
from dataclasses import dataclass
from typing import Literal


Region = Literal["US", "EU", "OTHER"]


@dataclass(frozen=True)
class PricingFlag:
    key: str
    version: int
    rollout_basis_points: int
    eligible_regions: frozenset[Region]
    salt: str
    enabled: bool
    candidate_policy: str
    control_policy: str


@dataclass(frozen=True)
class Decision:
    policy: str
    flag_version: int
    candidate: bool
    reason: str
    bucket: int | None


def cohort_bucket(account_id: str, key: str, salt: str) -> int:
    material = f"{salt}:{key}:{account_id}".encode("utf-8")
    digest = hashlib.sha256(material).digest()
    return int.from_bytes(digest[:8], "big") % 10_000


def select_policy(flag: PricingFlag, account_id: str, region: Region) -> Decision:
    if not 0 <= flag.rollout_basis_points <= 10_000:
        raise ValueError("rollout_basis_points must be between 0 and 10000")

    if not flag.enabled:
        return Decision(flag.control_policy, flag.version, False, "disabled", None)

    if region not in flag.eligible_regions:
        return Decision(flag.control_policy, flag.version, False, "region_ineligible", None)

    bucket = cohort_bucket(account_id, flag.key, flag.salt)
    candidate = bucket < flag.rollout_basis_points
    return Decision(
        flag.candidate_policy if candidate else flag.control_policy,
        flag.version,
        candidate,
        "cohort_match" if candidate else "cohort_control",
        bucket,
    )
Enter fullscreen mode Exit fullscreen mode

Ten thousand buckets let configuration express percentages in basis points without floating-point comparisons. The exact bucket count is a local protocol choice, not a universal standard; once production assignments depend on it, changing the hash input, modulus, text encoding, or salt is a migration because it can reshuffle accounts.

One byte changed is enough.

The pricing service should persist the returned policy, flag_version, and a request correlation identifier beside the quote. It should also emit a structured evaluation event. OpenTelemetry's logs model describes how logs can carry trace and span identifiers for correlation, including the case where existing log formats are collected through an agent rather than rewritten at once. That makes the flag decision inspectable alongside the pricing request without pretending that a log is the transaction record.

from datetime import datetime, timezone
from typing import Any


def evaluation_event(
    decision: Decision,
    *,
    flag_key: str,
    account_subject: str,
    region: Region,
    trace_id: str,
) -> dict[str, Any]:
    return {
        "schema_version": 1,
        "observed_at": datetime.now(timezone.utc).isoformat(),
        "event_name": "pricing_flag_evaluated",
        "flag_key": flag_key,
        "flag_version": decision.flag_version,
        "selected_policy": decision.policy,
        "candidate": decision.candidate,
        "reason": decision.reason,
        "bucket": decision.bucket,
        "account_subject": account_subject,
        "region": region,
        "trace_id": trace_id,
    }
Enter fullscreen mode Exit fullscreen mode

account_subject should be a deliberately designed pseudonymous subject, not an email address copied into every telemetry sink. Privacy and access policy still apply to pseudonymous data. The event time is observation time; if the system also tracks when a decision occurred or when a collector received it, those timestamps need distinct field names rather than overloaded semantics.

Log severity is another place where teams manufacture confusion. RFC 5424 defines ordered numerical severity values from Emergency through Debug, with lower numbers representing higher severity. Do not casually map an ordinary control-cohort decision to an error because “false” looks unsuccessful. A normal evaluation is informational telemetry; an invalid policy document or impossible version transition is an operational fault. Preserve the original severity when forwarding records, and document any mapping performed by a collector.

Proving rollback before increasing exposure

A percentage slider is not a rollout plan. Before the first candidate assignment, test the evaluator with fixed vectors: known account identifiers, regions, policy versions, and expected buckets. Those vectors should run in every implementation that evaluates the flag. A browser and server producing different cohorts is avoidable; for pricing, the cleaner answer is not to let the browser decide at all.

Then rehearse the sequence that matters. Load version 41 with the control policy, publish version 42 with a small eligible cohort, create both candidate and control quotes, and select version 41 again. New requests must use version 41. Existing version-42 quotes must retain their recorded terms according to the product's quote policy. Duplicate requests must not create conflicting transactions. Delayed evaluation events must still be attributable to version 42 rather than relabeled with whatever configuration is current when the batch arrives.

Observability should answer concrete rollback questions: Which policy version produced this quote? Was the account eligible by trusted region? Was the result a cohort match, an explicit override, a disabled flag, or a fallback after invalid configuration? How many requests used a stale snapshot, and for how long? Aggregate exposure counts are useful, but they cannot replace a decision record linked to the affected request.

There is a cost trade-off here. Recording every evaluation can create much more telemetry than recording one pricing decision per quote. For a pricing rule, instrument the authoritative decision point and sample ancillary reads only if the resulting evidence still supports investigation. Retention should follow the period in which the business must explain pricing decisions, while raw high-volume diagnostics may have a shorter life. Those periods are policy choices; inventing universal numbers would be irresponsible.

Promotion should be gated by evidence, not a timer alone. Compare candidate and control on predefined business and technical signals, check that region and cohort distributions match the intended policy, and verify that missing-configuration and stale-snapshot counters remain within an explicitly approved envelope. Keep an operator-visible version history. Require a reason for changes. Small systems benefit from boring controls.

The rejected request-time lookup

The rejected design calls a remote flag API synchronously for every pricing request and treats the response as the price-selection authority. It looks simple on a sequence diagram. In operation, its timeout, retry, cache, and partial-outage behavior become part of checkout correctness, while a later configuration response can disagree with the version that produced an earlier quote. Node.js middleware does not remove that boundary, and a Next.js server action does not make a client-derived region trustworthy; both are integration locations, while policy provenance and transaction persistence remain architectural obligations.

It still has a valid use case: low-consequence, non-transactional presentation choices where a missed evaluation can safely fall back and no durable decision must be reconstructed. Pricing is different because the selected rule crosses into a transaction boundary. Put remote configuration outside that critical request path, validate snapshots before activation, and make the locally evaluated version part of the durable quote.

The final selection among managed, self-hosted, and application-owned control planes can now be made without pretending they are interchangeable. Reject any option that cannot demonstrate deterministic assignment, explicit fallback behavior, versioned rollback, exportable decision evidence, and a tested recovery path. After that, compare operational effort, developer ergonomics, regional deployment constraints, and total cost. Cheap and simple are useful properties. They are not invariants.

References

Top comments (0)