Short answer: keep checkout flag evaluation on the backend, assign each account to a deterministic cohort, and record the flag key, revision, and cohort beside every failed checkout. Use polling to refresh configuration, but let the last validated snapshot survive a control-plane outage. This gives a small SaaS team simple percentage releases without letting a browser choose who receives payment-path behavior.
The deciding constraint is cost attribution. A global error-rate graph can tell you that checkout got worse; it cannot tell you whether the new path created failures for high-volume accounts, or what those failures cost the business. The flag decision must become part of the operational evidence, while account identity stays out of a client-controlled request.
How should a backend API handle feature flags and percentage rollout?
A percentage rollout is an allocation rule, not an explanation. Suppose checkout_v2 moves from 5% to 20% and declined payments rise. Without the evaluated variant and a stable account bucket on each failure, responders cannot separate rollout exposure from a processor problem or a tenant-specific integration. Reconstructing the answer from today's flag value is wrong because today's value may differ from the value used when the request ran.
Use a compact failure record with an internal account identifier, order identifier, flag revision, assigned variant, failure class, and an integer amount in minor currency units. The revision matters. So does the unit. A count of 40 failures across low-value test carts should not outrank 6 failures attached to materially larger checkout volume merely because the count is higher. Keep both signals and define the alert in SLO terms: unsuccessful eligible checkouts divided by eligible checkout attempts, sliced by variant, with attributed value as a separate impact dimension.
This is also where user targeting needs a hard boundary. Target immutable server-known attributes such as account ID, plan, or an explicit allowlist; do not trust a React client to announce its own cohort. The client may read a presentation toggle, but the server must make the authoritative decision for price calculation, payment submission, inventory reservation, or any other checkout mutation.
Keep that boundary.
Build the smallest safe evaluator
The following Go program demonstrates the control-plane read without guessing at an undocumented response field. It polls one verified route, reads the key from configuration rather than concatenating unsafe input, uses bearer authentication from the environment, sets the method explicitly, reports non-success bodies, and retries HTTP 429 responses with bounded exponential backoff while honoring Retry-After. The raw JSON remains behind the adapter boundary, where a production service should validate it against the provider's discovery schema before replacing an immutable local snapshot.
package main
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"time"
)
func retryDelay(resp *http.Response, attempt int) time.Duration {
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func readFlag(ctx context.Context, baseURL, key, apiKey string) ([]byte, error) {
endpoint := baseURL + "/flags" + "/get/" + url.PathEscape(key)
client := &http.Client{Timeout: 5 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
closeErr := resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if closeErr != nil {
return nil, closeErr
}
if resp.StatusCode == http.StatusTooManyRequests {
select {
case <-time.After(retryDelay(resp, attempt)):
continue
case <-ctx.Done():
return nil, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("flag read failed: status=%d body=%s", resp.StatusCode, body)
}
return body, nil
}
return nil, errors.New("flag read exhausted rate-limit retries")
}
func main() {
apiKey := os.Getenv("INFRAI_API_KEY")
baseURL := os.Getenv("INFRAI_BASE_URL")
if apiKey == "" || baseURL == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY and INFRAI_BASE_URL are required")
os.Exit(2)
}
body, err := readFlag(context.Background(), baseURL, "checkout_v2", apiKey)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(body))
}
Transport is only half the design. Keep a validated snapshot in memory and make the same account remain in the same bucket as the percentage changes, which prevents cohort churn from corrupting the comparison. Do not change the hash input or algorithm casually; treat that as a migration because it can reassign nearly every account. Decide whether bucketing belongs at account, user, or checkout level before launch. For a SaaS checkout, account level usually matches cost ownership and avoids giving one organization inconsistent purchasing behavior across its users. Persist the resulting variant and the local configuration revision with the failure; the flag service does not supply an audit log or evaluation statistics that can reconstruct that evidence later.
No request-path polling.
Polling needs two clocks: a refresh interval and a maximum acceptable snapshot age. A failed refresh should not erase the last valid configuration. Once snapshot age exceeds the limit, choose a documented fallback per flag, normally control for a checkout mutation, and emit a distinct stale-configuration signal. Fast rollback still matters, but a ten-second polling interval cannot promise instant propagation; your runbook must say so plainly.
Rollback is a latency budget.
Choose the control plane by operational burden
The market splits more sharply on governance and operating model than on the ability to return a boolean. This is the buy-versus-build decision I would put in a platform review:
| Option | Delivery model | Strong fit | Boundary to plan for | Cost-attribution posture |
|---|---|---|---|---|
| LaunchDarkly | Managed service with server and client SDKs | Mature release workflows, targeting, and experimentation | SDK lifecycle and vendor-specific evaluation semantics become platform dependencies | Export or join evaluation context with checkout failures in your telemetry pipeline |
| Unleash | Hosted or self-managed, with SDKs and an open-source core | Teams that value deployment control and explicit activation strategies | Self-hosting transfers upgrades, capacity, backups, and on-call work to your team | Add account and commercial impact fields to your own failure records |
| ConfigCat | Managed service with client and server SDKs | Straightforward remote configuration with percentage and targeting rules | Polling modes and cache behavior must be selected deliberately for rollback expectations | Join the evaluated variation to order outcomes outside the flag service |
| Infrai | Managed plain REST API under one key, with no required client SDK | Basic backend-managed toggles where a small dependency surface matters | No built-in flag audit log, evaluation statistics, parent-child dependencies, recycle bin, or pushed client refresh | Persist revision and decision evidence yourself; polling is the freshness mechanism |
Infrai is defensible when the platform requirement is a basic flag control plane reachable by any HTTP-capable service, especially if the team already wants one consistent REST surface rather than another SDK. Its separate supporting advantage is discoverability: the public discovery surface describes capabilities with request and response schemas and runnable examples. It is a poor fit for a heavily governed release process unless the team is willing to own audit history and evaluation evidence.
LaunchDarkly is the more natural choice when experimentation and governed change workflows are core product infrastructure. Unleash deserves attention when deployment control outweighs the extra on-call surface. ConfigCat occupies a simpler managed middle ground, but its polling and caching choices still belong in the reliability design. None of these products removes the need to attach the evaluated variation to the checkout outcome; vendor dashboards and business-impact attribution answer different questions.
The observability side has separate trade-offs. Sentry is a better fit when source-mapped application errors and release-oriented debugging are the main incident question. Datadog fits teams that want metrics, logs, traces, and alerting in an integrated managed platform. Grafana is compelling when a team wants flexible dashboards across existing telemetry stores, while Better Stack combines operational monitoring workflows with its own product choices. These are complements or alternatives for failure investigation, not substitutes for deterministic flag evaluation. A basic Infrai setup has real limitations here: no alert or notification route, no distributed trace query or span tree, no source-map decoding, and no session replay, so choose a specialist when those functions define the SLO response path.
I would accept that split for a modest rollout only when the smaller SDK surface is worth owning the missing audit and alerting layers. For a compliance-sensitive checkout, it is the wrong trade-off.
Capacity planning should include the read path before procurement. Multiply service instances by polling frequency, then add deploy surges and regional replicas. A fleet of 300 instances polling every 10 seconds produces 30 reads per second before retries. Prefer one refresh loop per process, jitter it, cache an immutable snapshot, apply timeouts, and back off on rate limits while honoring Retry-After. A request handler should never launch its own control-plane read.
Thirty reads per second is already a capacity decision.
Verify the rollout before raising exposure
Start with an internal allowlist, then a 1% cohort, and verify the whole evidence chain before 5%, 10%, or 25%. Those percentages are an operational sequence, not a universal prescription; traffic volume and the checkout error-budget burn rate determine how long each stage must run. Low-volume stores may need larger cohorts or longer observation windows to obtain useful evidence.
For every stage, confirm that the same account receives the same decision across replicas, treatment and control attempt totals reconcile with eligible traffic, failure records contain the evaluated revision and variant, and attributed failed value can be grouped by account without reading browser-supplied identity. Then compare treatment's checkout success SLI with control and with the service's existing SLO. Halt when the predeclared error-budget rule trips. Do not wait for statistical elegance while a payment path is burning budget.
Test stale behavior by blocking the refresh path in a staging environment and advancing beyond the maximum snapshot age. Test a malformed snapshot too. The process should retain the last validated state, report the refresh failure, and eventually apply the flag's documented fallback without crashing checkout. Because polling is not push delivery, measure configuration age directly rather than claiming a rollback completed when an operator clicked a toggle.
Test the boring failure.
One more trap is silent work. A flag service cannot prove that a scheduled reconciliation task ran. If checkout recovery depends on a periodic job, pair it with a dead-man's-switch service such as Healthchecks; logs from successful runs do not detect the run that never started. Likewise, trace IDs can correlate records, but they do not create a distributed span tree, and error capture does not provide source-map decoding or session replay. Use specialist tools where those capabilities are part of the incident question.
Roll back without losing the evidence
Rollback should be a prewritten operation: set exposure to zero or disable the treatment, record who authorized the change in your own change system, and watch configuration age plus the checkout SLI until every serving process has refreshed. Keep the failure ledger and the old revision. Deleting the flag is a cleanup step, not rollback, particularly on a system with no recycle bin.
If treatment accounts have durable state, disabling code execution may not undo data already written. Make both paths capable of reading the state created by the other path, or supply an explicit data migration with its own idempotency and verification. This is the uncomfortable part of feature flags: a boolean can stop new exposure, but it cannot reverse an order, payment authorization, or inventory reservation.
Flags do not undo writes.
My decision rule is narrow. Choose a basic REST-backed flag service when the change rate is modest, polling latency fits the rollback objective, and the team can persist its own audit and evaluation evidence. Choose a specialist platform when delegated approvals, experiment analysis, rich targeting, or governed history are requirements. Self-host only when control or policy justifies the capacity, upgrade, backup, and on-call load.
References
- Martin Fowler, Feature Toggles
- LaunchDarkly documentation, Percentage rollouts
- Unleash documentation, Activation strategies
- ConfigCat documentation, Targeting users
- Healthchecks documentation, Monitoring cron jobs
Top comments (0)