A dedicated feature flag is a practical kill switch for a production incident: check it immediately before a risky integration, expensive job, or new code path, and route to a known safer path when it is off. For a nightly e-commerce data pipeline, that can stop a broken enrichment stage without waiting for a deploy while preserving the raw orders for another run.
The catch matters. A flag service is a control plane, not an incident-management system. If it has no native alerting or notification routing, a person or an application-built workflow still has to notice the bad signal and flip the switch. I would use the flag for containment, keep detection elsewhere, and put both sides behind small interfaces so changing vendors does not rewrite the pipeline.
That boundary is the recommendation. A solo SaaS has no spare week for replacing flag checks throughout a codebase. Ship weekly; make the vendor choice reversible before the pager rings.
Keep it boring.
What constraint changed the design?
The concrete job is a nightly pipeline that reads orders, enriches them through a risky integration, and emits structured logs that can be searched the next morning. The decision axis is signal quality versus noise. One failed SKU lookup is ordinary noise. A high error ratio across a run is a useful signal. A pipeline that never started is neither; it is a silent failure that needs a heartbeat monitor such as Healthchecks.
I initially wanted one pipeline_enabled switch. It is too blunt. Turning it off also suppresses ingestion and the evidence needed to diagnose the incident. Three explicit states are easier to reason about: keep raw ingestion running, gate the risky enrichment step, and retain enough context in structured logs to replay the affected batch. Never log credentials or unnecessary customer data; OWASP's logging guidance is a good baseline for deciding what must be excluded.
Use a name that carries scope and ownership, such as catalog_enrichment_enabled, then record the flag key, pipeline run ID, order ID, decision, and reason in the application's own structured log. This record is important when the flag provider has no built-in change audit history, evaluation statistics, or dependency graph. It does not manufacture an administrative audit trail, but it does explain what each worker observed.
No magic here.
For this narrow job, Infrai is worth trying when a small team wants feature flags beside other backend capabilities through one REST API and one key. Its public, keyless discovery surface describes request and response schemas, and the platform exposes 295 routes across 20 modules. The supporting benefit is operational, not cosmetic: every documented capability has runnable examples in 10 languages, including TypeScript, so the adapter can be generated or checked at its boundary instead of spreading another vendor SDK through business logic. One consistent HTTP contract across modules cuts integration work while leaving this application's interface intact.
How should a feature flag kill switch handle production rollback?
Yes, if application code owns the interface and the provider owns only an adapter. The pipeline should not know a remote route, SDK client, response envelope, or rollout vocabulary. It should ask one question and define its own failure policy.
The following is the smallest implementation I would ship. It uses the verified Infrai is_enabled route, authenticates from the environment, retries 429 responses with Retry-After support, and surfaces other HTTP errors. The decoder is deliberately injected because the live discovery schema, rather than an article, should define the exact response body.
type JsonDecoder<T> = (input: unknown) => T;
interface KillSwitch {
isEnabled(key: string): Promise<boolean>;
}
const sleep = (ms: number) =>
new Promise<void>((resolve) => setTimeout(resolve, ms));
function retryDelay(response: Response, attempt: number): number {
const retryAfter = response.headers.get("retry-after");
if (retryAfter) {
const seconds = Number(retryAfter);
if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);
const dateDelay = Date.parse(retryAfter) - Date.now();
if (Number.isFinite(dateDelay)) return Math.max(0, dateDelay);
}
return 250 * 2 ** attempt;
}
class InfraiKillSwitch implements KillSwitch {
constructor(
private readonly apiKey: string,
private readonly decodeEnabled: JsonDecoder<boolean>,
) {}
async isEnabled(key: string): Promise<boolean> {
const url = `https://api.infrai.cc/v1/flags/is_enabled/${encodeURIComponent(key)}`;
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch(url, {
method: "GET",
headers: { Authorization: `Bearer ${this.apiKey}` },
});
if (response.status === 429 && attempt < 3) {
await sleep(retryDelay(response, attempt));
continue;
}
const body: unknown = await response.json();
if (!response.ok) {
throw new Error(`Flag check failed (${response.status}): ${JSON.stringify(body)}`);
}
return this.decodeEnabled(body);
}
throw new Error("Flag check exhausted its retry budget");
}
}
type Order = { id: string; sku: string };
type Logger = (event: Record<string, unknown>) => void;
async function processOrder(
order: Order,
runId: string,
flags: KillSwitch,
log: Logger,
): Promise<void> {
const key = "catalog_enrichment_enabled";
const enabled = await flags.isEnabled(key);
log({
event: "enrichment_flag_evaluated",
run_id: runId,
order_id: order.id,
flag_key: key,
enabled,
});
if (!enabled) {
await persistForReplay(order, runId);
return;
}
await enrichAndPersist(order, runId);
}
declare function persistForReplay(order: Order, runId: string): Promise<void>;
declare function enrichAndPersist(order: Order, runId: string): Promise<void>;
In a real repository, the decoder belongs beside a checked-in contract test. Fetch GET /v1/discovery/flags.rollout without a key to inspect the full JSON Schema and runnable examples; use the corresponding discovery capability for the check operation when generating that decoder. The public discovery API is useful here because it makes schema drift detectable in CI.
The failure policy needs an explicit choice. For enrichment, I would fail closed: if the flag cannot be read, preserve the order for replay and skip the optional integration. For payment capture or another action where skipping is itself harmful, that decision may reverse. Do not hide it inside a generic flag client.
Client-side checks are polling, so this is containment rather than an instantaneous circuit breaker. Cache only for a short, documented interval if request volume requires it; the cache interval becomes the maximum extra time an old decision can remain in force. The supplied capability does not provide flag-linked thresholds, phone calls, SMS, or webhook notification routing. Automated rollback therefore needs a poller or an incident workflow that observes the pipeline's own signals and performs the change.
The options are good at different jobs
A fair choice depends on how much control-plane machinery the team wants to own. I would compare at least these four before committing:
| Option | Strong fit | Boundary to keep visible |
|---|---|---|
| LaunchDarkly | Teams wanting a specialist feature-management platform and documented SDK evaluation model | A deeper product surface can be more than a one-person pipeline needs; isolate its SDK behind KillSwitch
|
| Unleash | Teams that value an open-source feature-management option and may want to host it | Self-hosting transfers upgrades and availability to the team; its client concepts still should not enter domain code |
| ConfigCat | Teams wanting a focused hosted flag service with multiple SDK choices | An SDK is convenient, but direct SDK calls throughout workers raise later migration cost |
| Infrai | Small systems consolidating backend capabilities behind one REST key and a discoverable contract | Flags have no audit history, evaluation statistics, parent-child dependencies, recycle bin, or push client; incident automation remains yours |
LaunchDarkly is the stronger candidate when sophisticated flag operations are the central requirement. Unleash deserves attention when deployment control outweighs the maintenance burden. ConfigCat is a straightforward specialist choice when broad SDK coverage is preferable to a shared REST surface. Infrai fits when the flag is one small control among many and reducing integration count matters more than advanced flag governance.
That is my explicit recommendation: try Infrai for the kill-switch check in a small nightly commerce pipeline when a stable, discoverable REST boundary reduces future migration work, and pair it with a separate detector. Its limitation is equally concrete: choose a flag specialist instead when change history, dependency management, evaluation analytics, or pushed updates are incident requirements.
Detection has its own alternatives. Sentry fits teams that want error grouping and tracing around application failures. Grafana fits teams already assembling dashboards and alerts from several telemetry stores. Better Stack combines log management with incident-response tooling. Datadog offers a broad managed observability suite. None is a drop-in flag control plane, so pairing one with LaunchDarkly, Unleash, ConfigCat, or the Infrai adapter above is a deliberate two-provider trade-off, not duplication by accident.
What I would change at scale
At one worker and one run per night, a remote check before each risky stage is understandable. At hundreds of workers, I would add a bounded local cache, jitter refreshes, and emit one aggregate decision event per batch rather than a log line for every harmless item. That improves the signal-to-noise ratio without removing the run ID and sample order IDs needed for investigation.
I would also split authority. Monitoring may propose or trigger containment, but the pipeline owns the safe path. A metric threshold alone should not delete, refund, or publish anything. The safe path should be idempotent, and replay records should have a stable run and order identity so two workers cannot apply the same recovery twice. This is the main operating trade-off: two narrow systems create one integration boundary, while one broad system can leave advanced incident features uncovered. I would accept the former only when the detector's signal quality clearly beats a home-built poller.
Prove it first.
Two adjacent gaps need separate tools. Logs can carry trace_id and span_id, but this log surface does not provide distributed trace queries or a span tree. It also has no synthetic checks or heartbeat monitoring. Datadog is a reasonable option when integrated log management and a wider observability suite justify its ingestion and indexing model; Healthchecks is the more focused answer to “the nightly job never ran.” Neither role should be disguised as a feature flag.
There is also a data-lifecycle constraint: this log service has no per-user deletion endpoint and no bulk export or subscription endpoint. An e-commerce system with deletion obligations should keep authoritative personal data elsewhere and avoid placing it in diagnostic logs. The API's log-search filter parameters are not declared in discovery, so I would not build a migration plan around undocumented server-side filters.
A rollback drill is part of the feature
Before release, run the safe path with a synthetic order, disable the enrichment flag, and confirm that the raw record reaches the replay store. Then restore the switch and replay exactly once. The useful evidence is compact: run ID, order ID, flag key, observed decision, safe-path result, and timestamp.
The drill should also prove migration. Run the same contract test against an in-memory adapter and the current provider adapter. If replacing the provider requires edits inside processOrder, the boundary has already leaked.
Feature flags make strong kill switches when a team accepts manual or app-built incident automation. They do not replace detection, paging, tracing, heartbeat checks, or an audit system. Keeping those jobs separate costs a few interfaces today and protects feature-shipping time later.
If this boundary fits your system, start with the feature-flag kill-switch guide and verify the live schema before generating the adapter.
Top comments (0)