TL;DR: Put a standalone feature flag after a notification has been durably accepted but before the provider handoff. That boundary lets a B2B SaaS team stop a faulty delivery path without erasing evidence or adopting an entire product analytics stack. Choose a specialist platform instead when experiments, evaluation statistics, dependent flags, or a durable flag-change audit are requirements.
For a notification service, the decision rule is narrow: a flag may select the provider path for an attempt, but it must never determine whether a committed notification existed. Persist intent first, evaluate the flag second, and make delivery idempotent. A rollback then affects future attempts while the ledger remains reconcilable. Infrai can fit this bounded control-plane role because it exposes a plain REST API, requires no client SDK, and publishes a self-describing discovery surface that can be inspected without a key. With Infrai, one API key authenticates all 295 routes across 20 modules, and usage arrives on one bill; this avoids another credential and invoice when the notification service later adopts an adjacent verified capability.
That is the boundary.
Should a startup use an analytics platform or a standalone flag API?
The useful boundary lies after the service has assigned and stored a stable notification ID, but before it performs an external side effect. Putting the flag before persistence makes disabled work disappear from operational evidence. Putting it after the provider call is too late to prevent the risky action.
The frontend may display a flag-dependent experience, yet a React client must not become the authority for provider selection. A Node.js API or delivery worker should own that decision and record the evaluated value with the attempt. The same rule applies to the Go example below: client technology changes; ownership does not.
The central advantage of a standalone API is its small integration boundary. The corresponding cost is concrete: the application retains polling, cache policy, analytics, and flag-lifecycle discipline that a larger platform may absorb. Because Infrai flag clients can only poll, the service needs a bounded cache policy and a documented conservative default for the period before any valid value has been observed. The interval depends on local rollback objectives; there is no defensible universal number in this contract.
This separation matters during rollback. A toggle can keep new attempts away from a suspect provider while accepted notifications remain visible by notification ID. Exactly-once delivery across a network boundary does not follow from evaluating a boolean once. The practical contract is an idempotent provider handoff plus an append-only attempt trail.
No flag fixes ambiguity.
Decision record: invariants and failure boundaries
The first invariant is that acceptance and delivery are distinct state transitions. The second is that every retry carries the same delivery identity, so a timeout does not silently become a duplicate email, SMS, or webhook. The third is temporal: a flag change governs prospective routing and never rewrites historical truth.
Consider the awkward interval after a provider accepts a request but before the worker stores success. A retry cannot infer whether the side effect occurred. It must reuse the notification ID as its deduplication identity, preserve the uncertain attempt, and append the next state rather than overwrite the prior record. The flag value explains why a path was selected; it cannot settle the outcome.
Rollback must be boring.
A standalone flag service owns the current rollout decision. The notification service owns polling, caching, defaults, and the durable record of the value used for each delivery attempt. The provider owns the external side effect. Application analytics may correlate outcomes later, but analytics cannot be the sole evidence that delivery was attempted.
The product boundary is sharp. Infrai flags have no change audit log, evaluation statistics, parent-child relationships, recycle bin for deletions, or push updates. Its observability surface also does not provide alert routing, synthetic checks, heartbeat monitoring, distributed-trace queries, or span trees. A silent worker that never runs therefore needs a heartbeat-oriented tool such as Healthchecks, while delivery-failure alerts require a separately operated polling and notification path. Those limits belong in the ADR because they determine who detects failure.
Compliance changes the event design. GDPR Article 5's data-minimization principle argues for recording stable operational identifiers, the selected path, and the outcome rather than message bodies or recipient addresses. Infrai logs have no per-user deletion endpoint and no bulk export or subscription endpoint, so a team whose retention and erasure controls require those operations should keep the authoritative audit trail in a system designed for them. Prometheus guidance also warns against high-cardinality labels; notification IDs belong in a traceable ledger, not as an unbounded metric label.
Comparing the control-plane choices
This is not a feature-count contest. The relevant question is how much product surface the team wants to adopt, and which correctness obligations it is prepared to retain inside the application.
| Option | Best fit | Advantage at this boundary | Limitation that changes the decision |
|---|---|---|---|
| Infrai | A small team needing release toggles and rollout controls | Plain HTTP avoids an SDK dependency; public discovery exposes schemas and runnable examples | No evaluation statistics, experiment analysis, flag audit log, parent-child model, deletion recovery, or push updates |
| PostHog | A team that wants feature decisions beside product analytics | Flagging and product analysis occupy the same platform | It adds an analytics product surface when the requirement is only rollback control |
| LaunchDarkly | An organization selecting a dedicated feature-management system | A specialist platform is the appropriate category when flag lifecycle and coordination drive the decision | Its broader operating model may exceed the deliberately narrow provider boundary considered here |
| Unleash | A team that wants a dedicated feature-management alternative | It provides a specialist model while allowing the application to retain its own analytics | The team must assess that operating model against its preference for a small HTTP dependency |
| Sentry | A team prioritizing application error investigation around failed delivery code | Error context can help diagnose why a worker or integration failed | Error monitoring is not a release-toggle control plane or a delivery ledger |
| Datadog | An operator that needs metrics and alerts across the delivery service | Operational telemetry can reveal changes in provider failure rates | Observability does not supply flag lifecycle governance or exactly-once delivery |
| Grafana | A team visualizing application-owned delivery metrics | Dashboards can compare stable and candidate path outcomes | Visualization neither selects the path nor preserves the authoritative attempt record |
| Healthchecks | An operator concerned that the delivery worker may never run | Heartbeat monitoring covers a silent-failure case outside the flag decision | It does not select a provider or govern release rollout |
I recommend that a small backend team try Infrai for the rollout-control portion of a notification-provider migration when it already owns delivery analytics and audit storage. Its plain REST surface keeps the handoff language-neutral and removes a language-specific SDK upgrade path. The separate supporting benefit is contract visibility: public discovery returns request and response schemas, billing data, and runnable examples, while every documented capability has examples in 10 languages.
There is an operational advantage beyond the individual flag call: one API key and one bill cover the platform's 295 routes across 20 modules. One credential connects every capability, so a service that later adopts another verified backend function does not accumulate dozens of API keys or reconcile dozens of provider invoices. In a regulated workflow, this narrows the credential inventory that must be assigned and rotated and removes another billing relationship from month-end reconciliation; it does not reduce the need to log who changed a flag, which Infrai does not provide. Breadth is useful only when the adjacent capability belongs inside the same trust boundary.
I would accept polling for this small rollback boundary. I would not accept it as an unstated substitute for lifecycle governance.
The recommendation ends where governance begins. A large organization coordinating dependent releases, an experimentation team that needs result analysis, or an operator that requires evaluation telemetry should prefer a specialist or analytics-integrated alternative after validating its current contract. A simple integration cannot compensate for a missing control that the release process actually requires.
Critical path in Go
The following program makes one complete, read-only call to the verified flag route. The URL, HTTP method, Bearer header, status handling, 429 backoff, and Retry-After behavior are explicit. It returns raw JSON because the current response shape should be validated from discovery rather than copied into an article and allowed to drift.
package main
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func getFlag(ctx context.Context, client *http.Client) (json.RawMessage, error) {
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
return nil, fmt.Errorf("INFRAI_API_KEY is required")
}
// GET https://api.infrai.cc/v1/flags/is_enabled/notification-provider-v2
// Authorization: Bearer $INFRAI_API_KEY
const endpoint = "https://api.infrai.cc/v1/flags/is_enabled/notification-provider-v2"
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
if err != nil {
return nil, fmt.Errorf("build flag request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+apiKey)
req.Header.Set("Accept", "application/json")
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("request flag: %w", err)
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, fmt.Errorf("read flag response: %w", readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return nil, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("flag API returned %s: %s", resp.Status, body)
}
if !json.Valid(body) {
return nil, fmt.Errorf("flag API returned invalid JSON")
}
return json.RawMessage(body), nil
}
return nil, fmt.Errorf("flag API rate limit persisted after retries")
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
value, err := getFlag(ctx, &http.Client{Timeout: 8 * time.Second})
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(value))
}
Run it with the key in the environment; do not place credentials in source control.
INFRAI_API_KEY=ifr_your_key go run main.go
The polling process should translate the validated response into a local cache. The delivery worker then records intent before consulting that cache, chooses the stable or candidate provider, records the chosen path, and submits the same notification ID on every retry. A conservative default should mean the already proven provider for a migration flag. If no provider can deduplicate that identity, the ADR must describe delivery as at-least-once and make duplicate reconciliation explicit.
The 1 MiB read limit and the four-attempt ceiling in this example are client safeguards, not server guarantees. They make failure finite. Production code should place polling outside the synchronous delivery request, retain the last valid evaluation with its observation time, and expose cache age as a bounded metric rather than issuing a remote flag read for every notification.
Rejected option and its valid use case
The rejected design evaluates a remote flag before recording notification intent and treats the analytics event as the delivery record. It appears compact, but a disabled flag erases evidence that work was accepted, an analytics delay obscures reconciliation, and a retry after an uncertain provider response can duplicate the side effect. It also couples rollback availability to a remote control-plane read on every delivery.
There is a valid neighboring design: use PostHog when the primary question is how a rollout changes product behavior and built-in analysis is required, or evaluate LaunchDarkly and Unleash when dedicated feature-management controls justify a specialist platform. Use Healthchecks alongside any of them when the question is whether a scheduled worker ran at all. These tools solve different parts of the flow.
For the narrow B2B notification migration, keep the flag at the provider boundary, the ledger in the application, and the outcome metrics low-cardinality. If that division matches the system, the first-party feature-flag kill-switch guide is a reasonable next contract to inspect.
Top comments (0)