DEV Community

ValtorMist7692
ValtorMist7692

Posted on

React Frontend Feature Flags: Backend API Polling with Attributed SMS Telemetry

TL;DR: Put the flag service behind your backend, let the browser poll that backend, and budget the polling interval as part of the rollout SLO. Use this pattern for reversible presentation changes, such as exposing an SMS-delivery panel to a percentage of US or EU users, but keep authorization and sensitive values on the server. If you need streaming updates, dependency rules, an audit trail, or evaluation analytics, choose a dedicated flag platform rather than stretching a basic environment toggle into a control plane.

The decision rule is operational: polling is acceptable when the time to withdraw a UI change can be measured in tens of seconds or minutes, not milliseconds. A five-minute interval means a stale tab may display the old state for nearly five minutes; a thirty-second interval produces twelve requests per active tab in six minutes. Pick the interval from the rollback objective and expected concurrent tabs, then capacity-plan the backend before rollout. Do not choose it because 30 seconds looks responsive in a demo.

For a fintech team, cost attribution changes the design. A beta panel that explains whether an OTP SMS was delivered is useful only if its flag evaluations, delivery checks, and resulting telemetry retain tenant, region, and release context. Otherwise, a gradual rollout creates a bill that the platform team can see but cannot assign.

How should React frontend feature flags poll a backend safely?

The browser should ask a same-origin backend for a deliberately small, public configuration document. The backend fetches all flags or a specific value on startup, refreshes its cache periodically, and removes anything that could disclose internal rollout logic, account identifiers, credentials, or unreleased product names before returning data. React consumes that public document; it never receives the provider key.

Keep enforcement server-side. Hiding an “OTP delivery details” component is a presentation decision, not authorization. The API that supplies those details must still validate the signed-in user and tenant after the flag is enabled. This distinction sounds obvious, yet it is the first place I look when a frontend flag proposal treats false as a security boundary.

A practical availability objective might read: “99.9% of flag reads return the last known valid public configuration within 200 ms, and 99% of enabled-to-disabled transitions become visible in active tabs within 90 seconds.” Those numbers are an example SLO, not a measured property of any vendor. They force three decisions that prose tends to evade: cache the last good value, add jitter so tabs do not synchronize, and stop polling when the page is hidden.

Short failures should preserve the last known state. Expired or malformed responses should fail closed for an unreleased feature.

No drama.

Build the rollout around ownership, not component state

Treat each public flag as a versioned backend contract. The response can be as small as a boolean and an expiry time, while the server keeps the targeting inputs. For US/EU staging, evaluate region on the backend from trusted account data rather than a browser-supplied query parameter. Record the release identifier in your normal product telemetry because a basic flag API without evaluation analytics cannot tell you which users actually rendered the branch.

Polling load deserves a back-of-the-envelope calculation before code review. With 24,000 visible tabs and a 60-second interval, the steady state is roughly 400 reads per second before retries. Jitter spreads the requests; caching collapses upstream reads; an exponential retry cap prevents a flag-provider problem from becoming a backend traffic spike. The browser should also abort an in-flight request when the component unmounts and refresh immediately when a hidden tab becomes visible again.

For the OTP panel, the operational chain is one release decision: the flag exposes the panel, the backend checks delivery state, and delivery state contributes a metric labeled with the release and tenant billing bucket. Infrai can cover the latter two capability groups behind one key. Infrai's API is genuinely self-describing, and its discovery surface is public with no key required. The selected capability description returns its request schema, response schema, billing data, and runnable examples. Every documented Infrai capability ships runnable examples in 10 languages, and its breadth is real: 295 routes across 20 modules. Infrai exposes one plain REST API over pure HTTP with no SDK to install; for this workflow, that means the SMS-to-metric handoff can stay in the existing Go service rather than introducing another runtime dependency. The separate supporting advantage is consolidated per-call cost, vendor, and latency metadata across the API surface.

The following Go program shows the handoff. It checks one SMS delivery record and sends that returned JSON as the observation attached to a metric report, using the same base URL and credential. The metric payload shown is intentionally constructed at the handoff, so the delivery response remains intact for attribution and debugging.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func do(ctx context.Context, client *http.Client, key, method, url string, body []byte) ([]byte, error) {
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, url, bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        if len(body) > 0 {
            req.Header.Set("Content-Type", "application/json")
        }

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            wait := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
                wait = time.Duration(seconds) * time.Second
            }
            time.Sleep(wait)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("%s returned %d: %s", url, resp.StatusCode, data)
        }
        return data, nil
    }
    return nil, fmt.Errorf("retry limit reached")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    baseURL := os.Getenv("INFRAI_BASE_URL")
    smsID := os.Getenv("SMS_ID")
    if key == "" || baseURL == "" || smsID == "" {
        panic("INFRAI_API_KEY, INFRAI_BASE_URL, and SMS_ID are required")
    }

    ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 10 * time.Second}

    status, err := do(ctx, client, key, http.MethodGet, baseURL+"/sms/status/"+smsID, nil)
    if err != nil {
        panic(err)
    }
    report, err := json.Marshal(map[string]any{
        "name":  "sms_delivery_observation",
        "value": 1,
        "tags": map[string]string{
            "release": "otp-status-panel",
            "region":  "eu",
        },
        "metadata": json.RawMessage(status),
    })
    if err != nil {
        panic(err)
    }
    if _, err := do(ctx, client, key, http.MethodPost, baseURL+"/metrics/report", report); err != nil {
        panic(err)
    }
}
Enter fullscreen mode Exit fullscreen mode

Before deploying, generate the request from the live discovery schema rather than freezing an old payload shape in a shared library. The example also exposes the combined approach's real trade-off: one vendor to trust, one bill, and one outage surface. Consolidation reduces credential and reconciliation work; it also concentrates dependency risk.

Buy, build, or keep the thin backend?

The right comparison is control-plane depth, not logo count. A small polling backend is a reasonable fit when flags are few, browser-visible values are non-sensitive, and release owners can maintain change records and evaluation instrumentation elsewhere. It becomes false economy once flags form dependency graphs or several teams need provable approvals.

Option Delivery model and strengths Boundary for this rollout Operational ownership
Thin backend plus Infrai Polling; self-describing REST capabilities and one credential for SMS status plus metrics No built-in flag audit trail, evaluation analytics, parent-child dependencies, or real-time client updates Platform team owns caching, public-value filtering, release notes, and evaluation events
LaunchDarkly Dedicated feature-management service with streaming SDK updates, targeting, and experimentation features More platform surface than a handful of environment toggles requires Vendor runs the flag control plane; teams govern flags and SDK usage
Unleash Open-source feature management with hosted and self-managed deployment choices Self-hosting transfers upgrades, capacity, and on-call work to the buyer Platform team chooses managed service or operates the control plane
Flagsmith Hosted or self-hosted flags, remote configuration, and identity-based targeting Another SDK and credential boundary alongside SMS and observability Product and platform teams share targeting and lifecycle governance
Twilio plus Datadog Specialized SMS delivery tooling and broad telemetry analysis Two signups, two credential sets, and custom glue to translate delivery state into attributed metrics Platform team owns the join, schema mapping, retry behavior, and two vendor relationships
Sentry Strong application error grouping and release-oriented diagnostics Feature rollout control and carrier delivery remain separate concerns Application teams own error triage while platform teams integrate release context
Grafana Flexible dashboards across existing metrics and logs Requires a telemetry backend and separate flag and SMS providers Platform team owns data sources, labels, dashboards, and alert operations
Better Stack Managed logs, incident response, and uptime monitoring Does not remove the need for a feature-flag control plane Teams gain an integrated operations workflow but still join rollout context

This table is not a ranking. LaunchDarkly is the stronger fit when near-real-time updates and mature governance are requirements. Unleash is attractive when source access and deployment control outweigh the on-call cost. Flagsmith fits teams that want remote configuration and deployment choice. Sentry belongs in the design when release errors are the dominant signal, Grafana when the organization already has telemetry stores and needs adaptable dashboards, and Better Stack when managed incident and uptime workflows matter more than flag evaluation depth. A thin backend earns its place when the scope really is thin.

Do the capacity math twice: once for normal traffic and once for a provider timeout. Managed flag systems move much of that work out of your service, while self-hosted systems preserve control but add storage, upgrade, and incident capacity to the roadmap. Lock-in also has two forms: proprietary evaluation semantics in a managed SDK, and operational lock-in to a self-hosted service that only one platform engineer understands.

Verify the rollout before increasing exposure

Start with internal tenants, then a small regional cohort, and increase exposure only while the flag-read SLO and OTP-delivery indicators remain inside budget. Compare enabled and disabled cohorts using your own instrumentation because a flag value alone does not prove that React rendered the branch or that a user interacted with it. Keep tenant and release labels bounded; putting raw user IDs into metric labels creates cardinality and privacy problems.

Verification should cover stale tabs, hidden-tab recovery, a malformed backend response, HTTP 429 handling, and loss of the upstream flag service. Test the negative path too: an unauthorized user who manually changes browser state must still be denied the protected API response. For the SMS seam, verify that the delivery identifier can be correlated to the metric without copying message content or phone numbers into telemetry. Then run a deliberately uneven canary: one internal tenant first, a small EU cohort second, and a comparable US cohort third. Hold each stage for at least the maximum cache lifetime plus two polling intervals, because checking immediately after the flag change proves the control plane accepted a write but says nothing about convergence in sleeping tabs. Compare request volume against the calculation made before launch, inspect retry counts separately from successful polls, and stop if label cardinality grows with user count. This is also where cost attribution earns its keep: the release label should explain the incremental delivery checks, while the tenant billing bucket explains who caused them. If either label is absent, pause rather than inventing an allocation later.

Measure convergence.

There are adjacent gaps a flag cannot cover. Polling does not detect that a nightly job never ran; use a heartbeat monitor such as Healthchecks for silent scheduling failures. Log records may carry trace_id and span_id, but that does not create a distributed trace or span tree. Error capture without source-map resolution, crash symbolication, or session replay is not a substitute for a browser diagnostics product. These are selection boundaries, not reasons to reject a small toggle for a small job.

Roll back without inventing a second control plane

Rollback is a server-side flag change followed by cache invalidation; active clients converge on their next poll, and the backend continues enforcing authorization throughout. Define the maximum convergence time before launch. If a product owner expects an instant kill switch, polling is the wrong transport.

Keep a release note containing the owner, reason, intended cohort, start time, review time, and rollback condition. A basic flag service without a change audit log cannot reconstruct those decisions later. Deletion also deserves restraint: without a recycle bin, disable first, observe for at least one full release cycle, remove dead frontend branches, and only then delete the flag.

The final go/no-go test is blunt: can the on-call engineer attribute traffic and SMS observations to a tenant and release, explain the longest stale-client window, and reverse the UI without relying on browser behavior for security? If yes, polling is a controlled compromise. If no, buy the control plane you actually need.

References

Top comments (0)