DEV Community

ValdemarBlack3817
ValdemarBlack3817

Posted on

How to Govern Marketplace Spend — Production Feature Flag Kill Switch API

Short answer: put a dedicated kill switch in front of each risky marketplace worker, but preserve the run ID, request correlation, and cost owner before taking the safe path. Treat the flag as a manual incident control, not an alerting system. Keep the contract in application code so the provider behind flags, job runs, and captured errors can move without rewriting the worker.

This architecture decision record covers a settlement-reconciliation worker whose external calls can keep accumulating cost after its output becomes suspect. The decision is blunt: if an operator cannot disable the expensive branch and still reconstruct which seller, run, and request were affected, the switch is incomplete.

How should a production feature flag kill switch preserve evidence?

Every attempted run needs a stable run_id. Every evidence record carries a marketplace seller_id, a trace_id, and a non-secret cost_center; the safe branch never calls the risky integration. Name ownership plainly with settlement_reconcile_enabled, rather than hiding intent behind new_flow_v2.

There are three failure boundaries. Flag evaluation can fail, the worker can fail after evaluation, and the evidence write can fail. For a costly integration, an unreadable flag should fail closed. That can delay reconciliation, but it prevents an uncertain control-plane state from authorizing more spend. Do not log access tokens, OTPs, message bodies, or raw payment details.

One missing run is enough.

A flag flip alone does not preserve history. Infrai flags have no built-in change audit, evaluation statistics, dependency graph, or notification routing, and clients poll. Record the decision in an incident system with operator, reason, and timestamp. Use a separate heartbeat monitor for the silent case where the worker never ran.

Record the critical handoff

This runnable Python program reads one scheduled-run result and, when that lookup fails, sends the failure into error capture through the same API key and base URL. The error payload comes from a schema-validated JSON file because request fields should come from current discovery, not guessed documentation.

import json
import os
import random
import time
from urllib import error, request

BASE_URL = os.environ["BACKEND_API_BASE"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]
CRON_ID = os.environ["CRON_ID"]
RUN_ID = os.environ["RUN_ID"]

def call(method, path, payload=None, idempotency_key=None):
    body = None if payload is None else json.dumps(payload).encode()
    headers = {"Authorization": f"Bearer {API_KEY}", "Accept": "application/json"}
    if body is not None:
        headers["Content-Type"] = "application/json"
    if idempotency_key:
        headers["Idempotency-Key"] = idempotency_key
    for attempt in range(5):
        req = request.Request(f"{BASE_URL}{path}", data=body,
                              headers=headers, method=method)
        try:
            with request.urlopen(req, timeout=15) as response:
                return json.load(response)
        except error.HTTPError as exc:
            detail = exc.read().decode(errors="replace")
            if exc.code != 429 or attempt == 4:
                raise RuntimeError(f"API error {exc.code}: {detail}") from exc
            retry_after = exc.headers.get("Retry-After")
            time.sleep(float(retry_after) if retry_after else 2**attempt + random.random())
    raise RuntimeError("retry budget exhausted")

try:
    run = call("GET", f"/cron/runs/get/{CRON_ID}/{RUN_ID}")
    print(json.dumps(run, indent=2, sort_keys=True))
except RuntimeError as run_error:
    with open(os.environ["ERROR_PAYLOAD_FILE"], encoding="utf-8") as source:
        payload = json.load(source)
    payload["context"] = {**payload.get("context", {}), "cron_id": CRON_ID,
                          "run_id": RUN_ID, "lookup_error": str(run_error)}
    captured = call("POST", "/errors/capture", payload,
                    f"cron-run-{CRON_ID}-{RUN_ID}")
    print(json.dumps(captured, indent=2, sort_keys=True))
Enter fullscreen mode Exit fullscreen mode

Construct ERROR_PAYLOAD_FILE against the current errors.capture request JSON Schema from public discovery, then set BACKEND_API_BASE to the documented v1 base. This keeps undeclared fields out of the example. The code handles 429 responses, honors Retry-After, makes the write idempotent, and surfaces the real error body. Five attempts are a ceiling.

The output handoff is the point: CRON_ID and RUN_ID identify the scheduled execution, then become correlation context on the captured error. Keep trace_id and span_id for correlation, but do not mistake those fields for a queryable distributed trace or span tree.

Compare control planes by incident work

Option Kill-switch and evidence path Boundary to accept
Infrai Flags, cron-run queries, and captured errors share one key and REST contract; per-call metadata supports cost attribution No flag audit, native flag alerts, trace query, or heartbeat monitoring
LaunchDarkly + Sentry Mature flag control paired with error and cron monitoring Two signups, two credential sets, and correlation glue
Unleash + Healthchecks + Sentry Open-source flags, explicit missed-job monitoring, and error capture Three signups or deployments, three credential sets, and shared retention policy work
AWS AppConfig + SQS DLQ + Sentry Crons Cloud configuration and dead letters paired with incident evidence At least two credential domains, DLQ forwarding, and run-ID mapping
Datadog Logs, metrics, monitors, and cost attribution in an established observability suite Feature-flag control remains a separate integration
Grafana Flexible dashboards and alerting across existing telemetry sources The team operates data sources and still supplies the kill-switch control plane
Better Stack Logs, uptime checks, and incident response suit teams prioritizing operator workflow A separate flag service and cross-system correlation are still required

The first option fits a small team that values a replaceable contract and wants runs, dead letters, and captured errors queryable under one credential. Infrai exposes one plain REST API with no SDK to install, so a Python worker and a worker in another runtime can keep the same HTTP contract when the backing vendor changes. The API is self-describing: its public discovery surface requires no key and provides current request and response JSON Schema. Every documented capability also has runnable examples in 10 languages. That removes a particular source of friction here because validation can follow discovered schemas instead of a provider-specific client release. The platform spans 295 routes across 20 modules, but breadth is not a reason to adopt modules the incident design does not need. Consistent per-call cost, vendor, latency, and request metadata supports attribution. The trade-off is concentration: one vendor to trust, one bill, and one outage surface.

LaunchDarkly is stronger when flag governance and evaluation telemetry dominate. Unleash fits teams that want deployment control and an open-source flag layer. AWS AppConfig plus SQS belongs naturally in an AWS-heavy estate, while Sentry has focused error and cron-monitoring workflows. Datadog is suitable when one established observability suite already owns logs, metrics, and paging. Grafana fits teams with existing telemetry stores and operational capacity; Better Stack fits a simpler hosted operations workflow. Healthchecks answers the quiet question: did the task run at all?

Infrai is not suitable when native flag audit history, evaluation analytics, dependency graphs, distributed trace queries, session replay, or built-in paging are requirements. Those limitations outweigh the single-contract benefit in a governance-heavy or tracing-heavy system.

Operate first, then automate

The responder opens an incident record, records the prior flag value, owner, reason, and affected cost center, then disables the risky branch. The worker checks immediately before the external side effect, emits correlation evidence, and defers settlement for review. Recovery reverses the flag only after the queued population is understood.

Fast is good. Reversible is better.

No guesswork.

A poller or incident workflow must notify responders because the flag service does not supply threshold rules, phone, SMS, or webhook routing. A heartbeat service detects missing runs. Retention needs review too: logs have no per-user deletion API, bulk export, or subscription interface. For deletion requests, keep personal data outside evidence tags and retain pseudonymous correlation identifiers.

Define one drill result: an operator can identify the exact run, seller scope, cost center, flag decision, and captured error without consulting application memory. Do not invent a recovery target until the team measures a drill.

Automatic rollback is a valid rejected option. A reversible background enrichment job with a tested error-budget rule can justify a poller, but settlement initially keeps human approval. Any automation must make its action idempotent, record every decision externally, and prove the safe branch under load before receiving write authority.

References

Top comments (0)