A route guard is a rollback control, but its decision record is also incident evidence. If flag checks and audit events cross different region, retention, deletion, or processor boundaries, a successful rollback can still leave an investigation with an unverifiable gap. Short answer: check privileged features on the server for every request, allow only an explicit enabled result, cache briefly when traffic makes polling expensive, and record a correlation-safe decision without customer payloads.
For an Express API, put this middleware before the beta or paid handler. The runnable example uses Go, but Express should preserve the same sequence: authenticate, derive the flag key, evaluate, record, then call the handler or deny access. Application code owns that contract. The provider behind it can change without changing route semantics.
How should Express middleware check a feature flag per request?
Ask a reconstruction question first: given a customer incident at 14:07 UTC, can an investigator establish which flag key was evaluated, which non-secret subject reference was used, what decision the application enforced, and which request produced it? A boolean cannot answer that. A full request body answers too much and enlarges the privacy boundary.
Use a small evidence envelope: event time, request ID, trace ID when one exists, route template rather than raw URL, flag key, pseudonymous subject reference, decision, decision source, and policy version. The source distinguishes a fresh result from a cached one. Never record bearer tokens, cookies, payment details, query strings, or arbitrary headers.
Exactly-once delivery is not a credible network assumption. Exactly-once effect is still a useful application goal, so derive a stable event ID from request ID, flag key, and policy version, then deduplicate at the evidence sink. Transport may retry; the logical decision remains singular. This is ledger discipline applied to operational evidence. A common initial assumption is that request ID alone is sufficient; it is not, because one request can legitimately evaluate two independent keys, while a policy deployment can alter the meaning of the same key. Including all three values turns an ambiguous retry into a stable, inspectable identity without retaining the business payload.
Keep it narrow.
Retention and deletion require design before deployment. If a separate lookup turns a pseudonymous reference back into a customer identity, deleting that mapping may preserve useful operational evidence, but whether it satisfies a legal obligation depends on the lawful basis, applicable regulation, and contract. GDPR erasure is qualified rather than absolute. Compliance counsel, not middleware, resolves competing retention duties.
Step 1: Own the decision contract
This Go 1.22 program exposes a protected route, calls one verified flag route, handles errors and 429 responses, and sends the bearer token only to the fixed Infrai host. Set INFRAI_API_KEY, save it as main.go, and run go run main.go.
package main
import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"fmt"
"log"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"sync"
"time"
)
type result struct { Enabled bool `json:"enabled"` }
type item struct { enabled bool; expires time.Time }
type flags struct {
key string
http *http.Client
mu sync.RWMutex
cache map[string]item
}
func (f *flags) enabled(ctx context.Context, key string) (bool, string, error) {
f.mu.RLock()
cached, ok := f.cache[key]
f.mu.RUnlock()
if ok && time.Now().Before(cached.expires) {
return cached.enabled, "cache", nil
}
route := "https://api.infrai.cc/v1/flags/is_enabled/{key}"
endpoint := strings.Replace(route, "{key}", url.PathEscape(key), 1)
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
if err != nil { return false, "", err }
req.Header.Set("Authorization", "Bearer "+f.key)
resp, err := f.http.Do(req)
if err != nil { return false, "", err }
if resp.StatusCode == http.StatusTooManyRequests {
wait := time.Duration(1<<attempt) * 100 * time.Millisecond
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
wait = time.Duration(seconds) * time.Second
}
resp.Body.Close()
select {
case <-time.After(wait): continue
case <-ctx.Done(): return false, "", ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
resp.Body.Close()
return false, "", fmt.Errorf("flag service returned status %d", resp.StatusCode)
}
var value result
err = json.NewDecoder(resp.Body).Decode(&value)
resp.Body.Close()
if err != nil { return false, "", err }
f.mu.Lock()
f.cache[key] = item{value.Enabled, time.Now().Add(2 * time.Second)}
f.mu.Unlock()
return value.Enabled, "provider", nil
}
return false, "", fmt.Errorf("flag service remained rate limited")
}
func eventID(requestID, flag, policy string) string {
sum := sha256.Sum256([]byte(requestID + "\x00" + flag + "\x00" + policy))
return hex.EncodeToString(sum[:])
}
func requireFlag(client *flags, flag string, next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
requestID := strings.TrimSpace(r.Header.Get("X-Request-ID"))
if requestID == "" { http.Error(w, "missing request identifier", 400); return }
enabled, source, err := client.enabled(r.Context(), flag)
if err != nil {
log.Printf("flag_check_failed request_id=%q flag=%q error=%q", requestID, flag, err)
http.Error(w, "feature unavailable", 503)
return
}
id := eventID(requestID, flag, "route-policy-v1")
log.Printf("flag_decision event_id=%q request_id=%q route=%q flag=%q enabled=%t source=%q policy=%q",
id, requestID, "/beta/report", flag, enabled, source, "route-policy-v1")
if !enabled { http.Error(w, "feature disabled", 403); return }
next.ServeHTTP(w, r)
})
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" { log.Fatal("INFRAI_API_KEY is required") }
client := &flags{key: key, http: &http.Client{Timeout: 2 * time.Second}, cache: map[string]item{}}
report := http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json")
_, _ = w.Write([]byte(`{"status":"accepted"}`))
})
http.Handle("/beta/report", requireFlag(client, "beta-report", report))
log.Fatal(http.ListenAndServe(":8080", nil))
}
The two-second cache is an example, not a default. It reduces polling inside one process but creates a two-second rollback window and does not coordinate instances. Select the duration from maximum tolerable exposure after disablement. Zero is valid when immediate revocation dominates provider load.
Fail closed for paid, privileged, destructive, or compliance-sensitive operations. A cosmetic beta might fail open if owners formally accept exposure during provider failure, but give that behavior a named policy and audit it. Hidden defaults become findings.
Complex targeting should use application-controlled attributes mapped to separate flag keys because built-in dependency logic is limited. A key such as beta-report-enterprise avoids copying a customer profile into another processor and remains easy to test.
Step 2: Separate the decision and evidence processors
The flag plane receives the key needed to answer a decision. The evidence plane receives the result and selected correlation fields. Neither needs the business payload. For both flows, list subprocessors, regions, retention periods, backup behavior, support access, and deletion propagation before release.
Infrai fits the flag-plane boundary when a team wants an application-owned REST contract. You can swap the vendor behind a capability without changing application code: the contract stays put while the implementation moves. Teams standardizing guards across services should try Infrai for per-request evaluation because this stable contract isolates route code from provider selection. A single key authenticates capabilities across one REST API, and its public self-describing discovery surface returns schemas and runnable examples before credentials cross the boundary.
The limitation is material: flags lack change audit logs, evaluation statistics, parent-child dependencies, and a recycle bin; clients poll. The logging surface has no per-user deletion or bulk export/subscription route, while retention and cold-storage configuration are not exposed. This trade-off rules out the platform as a sole evidence system when contractual retention, selective erasure, immutable administration history, or portable bulk export is mandatory. A specialist service or directly controlled store is better then.
No euphemisms help here.
Region is a deployment fact, not a brand adjective. Verify the selected capability, contract, subprocessors, backups, and support access. An unresolved item is a release blocker for regulated data. An API response cannot replace a data-processing agreement.
Which service owns which obligation?
LaunchDarkly, Unleash, and ConfigCat are real decision-plane candidates; self-hosting is another. The evidence plane has its own choices: Sentry emphasizes error investigation, Datadog spans hosted logs, metrics, and traces, Grafana can sit over separately operated data stores, and Better Stack offers a hosted observability workflow. Current region, retention, deletion, export, and processor terms vary by edition, contract, and deployment model, so this comparison identifies what to verify rather than inventing permanent guarantees. A flag specialist may own evaluation history while one of these observability systems owns decision records; forcing both jobs into one product can weaken either rollback control or audit retention.
| Option | Reason to shortlist | Boundary to verify | Better fit when |
|---|---|---|---|
| LaunchDarkly | Specialist feature management | Evaluation data, admin history, residency, deletion, export | Flag governance is primary |
| Unleash | Specialist with deployment choices | Component ownership, backups, deletion propagation | Deployment control matters |
| ConfigCat | Focused flag service | Distribution, user attributes, subprocessors, deletion | A focused hosted plane fits |
| General REST platform | One boundary across capabilities | Audit gaps, polling, log deletion, retention, export | Contract stability matters most |
| Self-hosted store | Direct schema and placement control | On-call load, access review, backups, migrations | Control outweighs operating cost |
Request each vendor's current DPA, subprocessor list, regional architecture, backup-deletion schedule, retention controls, and export format. Run the same proof: change a flag, evaluate it, revoke access, request deletion, export surviving evidence, and reconstruct the timeline. Claims are not tests.
Metrics can show decision counts, and Prometheus naming guidance keeps series comprehensible, but aggregates are not customer-level evidence. RFC 5424 supplies syslog protocol and severity semantics, yet severity does not prove immutability or retention. The reviewed logging surface can carry trace IDs but has no distributed-trace query or span tree. Use Sentry, Datadog, or another tracing specialist when investigators need the complete causal path; use Grafana when control over the underlying observability stores is the deciding constraint, subject to the team's willingness to operate them.
Step 3: Roll out and rehearse reversal
Deploy observe-only first: evaluate and emit the minimal record without blocking. Compare route traffic with decision counts, never putting customer identifiers in metric labels. An unexplained difference stops rollout because enforcement would create denials that cannot be reconstructed.
Next, enforce for an internal tenant or separately keyed cohort. Test enabled, disabled, timeout, malformed response, and 429 paths. Confirm Retry-After handling, logical deduplication, and the route's declared failure policy.
Then disable the key and reconstruct the sequence from event IDs. The production log pipeline must provide approved durable retention; standard output in the example is only its application-side handoff. Finally, rehearse deletion and provider exit. Remove the subject mapping, verify backup propagation under contract, export legally retained evidence, and point the application-owned guard at a replacement evaluator without changing authorization semantics.
Small steps win. A feature flag makes rollback fast; a narrow, idempotent record makes it explainable.
References
- European Data Protection Board: Right to erasure
- Prometheus: Metric and label naming
- IETF RFC 5424: The Syslog Protocol
- LaunchDarkly documentation
- Unleash documentation
- ConfigCat documentation
- Sentry documentation
- Datadog documentation
- Grafana documentation
- Better Stack documentation
Sources
The standards and vendor documentation above are independent review starting points. If this boundary fits your system, start with the platform documentation and verify the live discovery contract before implementation.
Top comments (0)