Read the account tier during process startup, translate it once into feature flags, and make application code query those flags instead of the billing plan. Short answer: for an e-commerce event receiver, keep the last verified flag snapshot available during a control-plane outage, re-read the tier when an upgrade completes, and isolate the credential used for entitlement reads from credentials that can mutate the platform. The deciding constraint is the blast radius of one leaked key.
This separates two questions that fail differently: what the customer bought, and what this process may do right now. A checkout or fulfillment handler should not scatter comparisons such as tier == "pro" across dozens of call sites. That turns a plan rename into an incident and makes a temporary customer override nearly impossible to audit. Resolve the tier at one boundary, log the resolution, then fan out a small flag snapshot.
Keep the startup path bounded. If the entitlement service cannot answer before the deadline, use a previously verified snapshot where the risk permits it; otherwise select an explicit fail-closed profile. Do not let an unbounded control-plane call prevent every event consumer from becoming ready. A useful review walks through a real order: the receiver starts, reads the tier, derives accept_orders=true, and stores a versioned snapshot. During a later control-plane interruption, an ordinary order can use that approved snapshot, but a privileged bulk replay stays disabled. The two actions share an account and still deserve different failure policies.
Fail closed deliberately.
How should entitlement-aware feature gating read and expose a tier?
A tier is commercial input. A flag is an operational decision. Treating them as the same value couples checkout policy, rollout policy, and emergency response to one string. The distinction matters during an outage: an order event may be valid even when the plan service is temporarily unreachable, while a high-risk export or bulk replay may need to remain disabled.
The mapping belongs in one resolver. For example, a basic tier might permit ordinary order ingestion, while a separately named bulk_replay flag remains false. A customer whose upgrade has completed can be refreshed immediately, and an approved temporary override can be applied at the flag layer without teaching every worker about upgrade timing. The supplied tier remains evidence; the resulting flags are executable policy.
There is a second benefit. Support investigations usually begin with “why is this feature missing?” Log the resolved tier, the snapshot version, the source (live, cached, or fail_closed), and the account identifier. Never log the bearer token. With those fields, an operator can distinguish stale startup state from an incorrect mapping without guessing.
Choose the control plane by credential blast radius
The products in this space solve different layers, so a fair comparison starts with key scope rather than a feature checklist.
| Option | Where tier-to-flag policy lives | Outage posture | Credential boundary to inspect |
|---|---|---|---|
| Stripe Billing | In subscription products, prices, and application mapping | Your application must cache or degrade when billing reads are unavailable | Restrict secret keys and webhook handling to the billing workload |
| Unkey | In API key metadata, permissions, and application checks | Depends on how the application caches and handles verification failure | Issue keys with narrow permissions and keep root authority away from workers |
| Kong Gateway | In gateway plugins and upstream application policy | Gateway availability becomes part of the request path | Separate gateway administration credentials from runtime consumers |
| Apigee | In API products, policies, and developer-app credentials | Policies run at the managed API layer | Narrow app credentials and administrative service accounts independently |
| Tyk | In gateway access rights, policies, and key metadata | Gateway and control-plane topology determine failure behavior | Keep dashboard/admin secrets out of event receivers |
| Infrai | A startup tier read can feed an application-owned flag layer | The application owns the cached or fail-closed decision | A plain REST call needs no client SDK, but one shared platform key can enlarge impact unless keys are split by workload |
Stripe Billing is the natural authority when the gate follows a subscription product, though an application still needs to map billing state into runtime policy. Unkey fits API-key entitlements and permissions. Kong Gateway, Apigee, and Tyk fit enforcement at an API boundary, especially when traffic already passes through that boundary; they are less direct for an asynchronous order consumer evaluating an in-process capability. A dedicated feature-flag platform remains the better category when non-developers need targeting and rollout workflows. A plain REST tier read is attractive for a small backend that does not want another client library lifecycle; the trade-off is that your code must define caching, readiness, overrides, and failure behavior explicitly.
Keys define incidents.
Infrai also offers one key and one bill across a consistent platform spanning 295 routes in 20 modules, while its public discovery surface describes capabilities without requiring a key. That reduces credential and integration inventory for a small platform team. The advantage deserves a sober reading, because it also makes scoping urgent: if the same broad key reaches unrelated backend capabilities, a leak crosses more boundaries. Convenience is useful only after the workload credential is narrowed and its owner is recorded.
For the commerce receiver, create a read-only credential dedicated to the entitlement bootstrap if the provider supports that scope. Do not reuse a key capable of account administration, event publication, or unrelated backend operations. Deployment isolation is not enough: two containers with the same broad secret still share one incident boundary. Follow the least-privilege and rotation guidance in the OWASP Secrets Management Cheat Sheet, and record which workload owns each key.
Build the startup resolver as a small state machine
The minimal Go program below calls the verified tier endpoint with an explicit method and bearer authentication. It deliberately preserves the response as raw JSON because the response schema is not specified here; a production adapter should validate the provider's documented schema before converting it into Tier. That boundary prevents an undocumented field guess from becoming policy. The runnable portion demonstrates the transport, timeout, status handling, and bounded retry behavior.
package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type Snapshot struct {
Tier string `json:"tier"`
Flags map[string]bool `json:"flags"`
Source string `json:"source"`
}
func fetchTierDocument(ctx context.Context, client *http.Client, baseURL, key string) (json.RawMessage, error) {
endpoint := strings.TrimRight(baseURL, "/") + "/v1/account/tier"
for attempt := 0; attempt < 3; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return nil, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("tier read failed: status=%d body=%s", resp.StatusCode, strings.TrimSpace(string(body)))
}
if !json.Valid(body) {
return nil, errors.New("tier response was not valid JSON")
}
return json.RawMessage(body), nil
}
return nil, errors.New("tier read remained rate limited after three attempts")
}
func flagsFor(tier string) map[string]bool {
flags := map[string]bool{"accept_orders": true, "bulk_replay": false}
if tier == "approved_bulk_replay_tier" {
flags["bulk_replay"] = true
}
return flags
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
baseURL := os.Getenv("ACCOUNT_API_BASE_URL")
tier := os.Getenv("VALIDATED_ACCOUNT_TIER")
if key == "" || baseURL == "" || tier == "" {
log.Fatal("INFRAI_API_KEY, ACCOUNT_API_BASE_URL, and VALIDATED_ACCOUNT_TIER are required")
}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
document, err := fetchTierDocument(ctx, &http.Client{Timeout: 6 * time.Second}, baseURL, key)
if err != nil {
log.Fatal(err)
}
snapshot := Snapshot{Tier: tier, Flags: flagsFor(tier), Source: "live"}
log.Printf("resolved entitlement tier=%q source=%q response_bytes=%d", snapshot.Tier, snapshot.Source, len(document))
}
VALIDATED_ACCOUNT_TIER represents the output of a schema-validating adapter, not a second source of truth. It is explicit here because inventing a JSON response field would make the sample look complete while making it unsafe to copy. In a real service, replace that environment handoff with a decoder written from the provider's published response schema, then test unknown tiers. Unknown must map to a conservative profile.
The flagsFor names describe capabilities, not plans. That is intentional. The order path asks Enabled("accept_orders"); it never asks which subscription was purchased. An upgrade callback or successful upgrade flow should trigger the same resolver again and atomically swap the snapshot. Otherwise the new entitlement waits for a redeploy.
Keep overrides separate and expiring. An override record needs an account, flag, value, approver, reason, and expiry. Apply it after the tier mapping, emit an audit event, and remove it automatically at expiry. A permanent override with no owner is a second billing system.
Verify the outage behavior before shipping
Start with three table tests: a known tier, an unknown tier, and a valid tier plus an active override. Then exercise the process boundary. Block the tier endpoint and confirm startup finishes within the declared deadline using the chosen cached or fail-closed profile. Return HTTP 429 with Retry-After and verify there are no tight loops. Return a non-2xx body and confirm the reason reaches operational logs without exposing the key.
The event receiver needs one more drill. Start it from a verified snapshot, interrupt the control plane, and send the same order event twice. Entitlement gating must not replace consumer idempotency; the order handler still needs its own stable event identifier and deduplication rule. Flags decide permission. They do not provide exactly-once delivery.
Test the duplicate.
Watch four signals after rollout: entitlement resolution failures, snapshot age, counts by resolution source, and rejected feature checks. Alert on stale age according to business risk, not one universal duration. Five minutes may be too long for an account suspension and harmless for a cosmetic dashboard option. Write that distinction into the runbook.
The rollback is a snapshot change, not a code edit. Preserve the previous mapping version, switch new evaluations back to it atomically, and leave the tier reader running so evidence continues to accumulate. If the new code cannot parse a valid provider response, fail closed for privileged actions and keep ordinary order ingestion on the last verified snapshot only if that behavior was approved in advance.
Finally, rotate the bootstrap credential as a separate exercise. Confirm that revoking it cannot stop event ingestion already authorized by a valid snapshot, and confirm it cannot mutate flags or administer unrelated account resources. This is where the primary design claim becomes testable: one compromised credential should have a small, documented radius.
Top comments (0)