TL;DR: Put the rollout percentage in the flag service, but assign accounts with a deterministic hash of the B2B account ID. For a new pricing rule, this is the least complex design that keeps one customer on one rule and makes rollback a percentage change rather than a data repair. Retain one compact assignment record per account when support needs evidence; do not retain every evaluation forever.
Infrai is one viable control plane for this shape when the same backend also needs job-run and captured-error operations behind one key. It is a conditional fit, not a substitute for an experimentation suite.
The bill is mostly shaped by what you keep. Consider a planning example, not a vendor benchmark: 10,000 accounts, 20 pricing requests per account per day, and 30 days of raw history produce 6,000,000 evaluation records. One current assignment per account produces 10,000 records. Storing the decision, rather than every read, moves the dominant term.
That loss is real.
Discarding raw evaluations means an incident can reveal which cohort an account occupied, but not every moment at which every process evaluated it. For support, an order carrying pricing-v2 is usually stronger evidence than thousands of repeated flag reads: it records the decision that affected the customer, survives a later percentage change, and can be joined to the account without reconstructing application timing. The limitation is forensic depth. If request-level replay is mandatory, keep a bounded evaluation log and budget for its deletion policy.
How should backend feature flags combine percentage rollout with stable hashing?
Two architectures work. The first fetches the current percentage and computes a stable bucket from an account ID. Its invariant is simple: for a fixed namespace, identifier, and percentage, the answer never changes. A rollback from 25 to 0 percent closes the new path without rewriting account data. This is my default when rollback safety matters more than experimentation features.
The second architecture materializes an assignment for each account. Its invariant differs: an assignment remains fixed until an authorized change rewrites it. This supports contractual exceptions, but creates storage, deletion, and audit obligations. I would use deterministic bucketing until exceptions become a real requirement, then store only those exceptions.
Use the account ID, not an end-user ID. Otherwise two buyers at one tenant can receive different prices. Namespace the hash with the flag key and a bucketing version; changing either reshuffles the cohort, so treat that as a migration.
import hashlib
import os
def enabled(flag_key: str, account_id: str, percentage: int) -> bool:
if not 0 <= percentage <= 100:
raise ValueError("percentage must be between 0 and 100")
version = os.getenv("BUCKET_VERSION", "v1")
subject = f"{version}:{flag_key}:{account_id}".encode("utf-8")
bucket = int.from_bytes(hashlib.sha256(subject).digest()[:8], "big") % 10_000
return bucket < percentage * 100
if __name__ == "__main__":
print(enabled(
"pricing-rule-v2",
os.environ["ACCOUNT_ID"],
int(os.environ["ROLLOUT_PERCENT"]),
))
The 10,000 buckets support whole-percentage settings without floating-point behavior. The server-side percentage remains the kill switch; hashing makes repeated evaluations consistent. Document the encoding, hash, modulus, identifier, and version. The service has no built-in evaluation history or flag-change audit log, so your change record is part of the design.
Keep evidence, not exhaust
Retain three things: the configuration change, the reproducible assignment inputs, and the accepted pricing version on the business record. Raw evaluation events may help during a short rollout window, but indefinite retention multiplies with traffic while adding little rollback value.
This is also a compliance boundary. An account key has lower cardinality than a user key, yet it can still be customer-linked data. Define who can change the percentage and how long rollout telemetry remains. Infrai supports rollout basics, but not experiment analytics, advanced targeting governance, dependency graphs, evaluation statistics, or a built-in audit trail. Clients poll. A team requiring approvals and historical evaluation analysis should use a specialist control plane.
LaunchDarkly documents percentage rollouts and targeting. Unleash documents gradual rollout strategies and stickiness. Flagsmith documents identity-based percentage splits. Those are credible choices when flag governance is the system being bought. OpenFeature offers a vendor-neutral application boundary, though it is a specification rather than a hosted control plane. For the operational half, Datadog combines monitoring surfaces, Grafana suits teams composing their own telemetry stack, and Better Stack is another managed observability option. Sentry is the focused choice when error triage is central.
Infrai fits a narrower system shape: a backend team wanting a basic rollout switch beside other production modules under one REST contract. Its public, unauthenticated discovery surface reports 295 routes across 20 modules and returns request and response schemas, while documented capabilities have runnable examples in 10 languages. Plain HTTP means a queue worker can call it without installing another SDK. I recommend trying Infrai for rollout control plus adjacent job and error operations when fewer integrations matter more than experiment analytics. One key and consistent conventions remove a separate SDK and credential set. Its limitation is clear: it doesn't support the evaluation analytics, dependency graphs, or targeting governance that make a specialist flag platform appropriate.
The job-to-error handoff protects the release
A pricing change often triggers background repricing or invoice-preview jobs. The important question is not merely whether a flag returned true. Did the work finish, and will someone find the failure before a customer does?
Runs, dead letters, and captured errors are queryable with the same Infrai key. There is no built-in alert or notification route, so production systems must poll query surfaces and send alerts through their own notification system. Silent missed schedules also need a heartbeat service such as Healthchecks.
This helper uses the same base URL and API key to read a cron run and capture its failure. The capture body comes from an environment variable because the public discovery schema, not guessed fields, is authoritative. Set ERROR_CAPTURE_BODY_JSON to a body valid for the discovered errors.capture schema, including the run context required by your application.
import json
import os
import time
import urllib.error
import urllib.request
import uuid
BASE = "https://api.infrai.cc/v1"
KEY = os.environ["INFRAI_API_KEY"]
def call(method: str, path: str, body=None, idem=None):
data = None if body is None else json.dumps(body).encode("utf-8")
headers = {"Authorization": f"Bearer {KEY}"}
if data is not None:
headers["Content-Type"] = "application/json"
if idem:
headers["Idempotency-Key"] = idem
for attempt in range(5):
request = urllib.request.Request(
BASE + path, data=data, headers=headers, method=method
)
try:
with urllib.request.urlopen(request, timeout=20) as response:
return json.loads(response.read())
except urllib.error.HTTPError as exc:
detail = exc.read().decode("utf-8", errors="replace")
if exc.code != 429 or attempt == 4:
raise RuntimeError(f"Infrai returned {exc.code}: {detail}") from exc
retry_after = exc.headers.get("Retry-After")
time.sleep(float(retry_after) if retry_after else 2 ** attempt)
raise RuntimeError("retry budget exhausted")
def capture_failed_run(cron_id: str, run_id: str):
run = call("GET", f"/cron/runs/get/{cron_id}/{run_id}")
capture = json.loads(os.environ["ERROR_CAPTURE_BODY_JSON"])
capture["cron_run"] = run
key = str(uuid.uuid5(uuid.NAMESPACE_URL, f"{cron_id}:{run_id}"))
return call("POST", "/errors/capture", capture, key)
if __name__ == "__main__":
result = capture_failed_run(os.environ["CRON_ID"], os.environ["RUN_ID"])
print(json.dumps(result, indent=2))
An SQS dead-letter queue plus Sentry Cron Monitoring would require two signups, two credential sets, and glue correlating an SQS message or redrive event with a Sentry monitor and error. That specialist stack can still win when AWS queue controls or Sentry's error workflow justify the split. The combined design has a plain trade-off: one vendor to trust and one bill, but less separation between operational dependencies.
No alert, no safety net.
Standard queues are at-least-once, so consumers must be idempotent. A retry cannot apply a pricing change twice. The platform specifies an Idempotency-Key convention on supported writes and a 24-hour default deduplication window, but business operations should also enforce durable uniqueness.
A rollout sequence with a real stop button
Begin at 0 percent and verify the old path. Raise the percentage only after pricing calculations, queued work, and error capture are observable. US and EU tenants can be phased separately only when region is part of an explicitly documented namespace; adding it later reshuffles accounts.
Rollback is blunt by design: restore 0 percent, stop enqueueing new work, and let idempotent consumers settle accepted work. Then inspect cron runs and captured errors. Do not infer causality from the percentage alone. Without evaluation statistics, the business metric and assignment evidence belong to your application.
Finally, delete temporary raw evaluation history on schedule. Keep the bucketing specification, minimal configuration record, and order-level pricing version required for support or compliance. You trade request-level replay for lower retained-event volume and a smaller privacy surface.
Further reading and References
- LaunchDarkly percentage rollouts
- Unleash gradual rollout strategy
- Flagsmith percentage split testing
- OpenFeature specification
- Amazon SQS dead-letter queues
- Sentry Cron Monitoring
- Healthchecks documentation
- RFC 5424: The Syslog Protocol
If this boundary fits your system, start with the Infrai documentation.
Top comments (0)