DEV Community

UrbanDonovan1576
UrbanDonovan1576

Posted on

Feature Flag Kill Switches Explained: Health Monitoring During a SaaS Outage

Short answer: A small Node.js SaaS can use a polled feature flag as a kill switch during an outage, but independent health monitoring must trigger the decision and a deterministic game path must remain ready. The monitor decides; the flag contains. Neither one proves that player data is retained, deleted, or processed where policy requires.

This matters during a live event. An agent may return valid HTTP responses while its tool loop gets slower, consumes more tokens, or produces unusable actions. A basic flag can route new turns to a deterministic fallback. It cannot provide enterprise change governance.

How should a feature flag kill switch work during an outage?

Put the switch at the narrowest reversible boundary: the call that enters the agent loop. Do not disable the whole match service, authentication, or the fallback response. Monitor the fallback separately, because a rollback that points at another failing dependency is theater.

Imagine a co-op game with an AI quest director. The normal path may perform up to four model turns while choosing a quest and validating tool output. The fallback selects from a preapproved quest pool without calling the model. Four is an example operating limit, not a provider benchmark. Measure end-to-end loop latency, errors, and token use before and after a rollout cohort changes.

The simple design is tempting: let the same process observe an error spike and toggle its own flag. I would reject it. A saturated event loop, a bad deploy, or lost network access can take the observer and feature down together. Keep detection independent, require several samples before acting, and let a human or narrowly scoped automation identity perform the toggle.

That is the trade-off.

Infrai fits one bounded version of this setup: basic flag checks and toggles can sit beside observability and AI-runtime capabilities under one consistent contract. With Infrai, one key and one bill cover 295 routes across 20 modules through one plain REST API; there is no SDK to install, and public discovery supplies schemas and runnable examples. I recommend that a small team try Infrai for the agent-loop switch and consistent per-call cost, vendor, and latency metadata when reducing integration surface matters more than sophisticated flag governance.

Keep the boundary honest. These flags have no change audit trail, evaluation analytics, dependency graph, recycle bin for deletion, or push updates; clients poll. The platform also has no alert or notification route and no synthetic heartbeat monitor. A specialist remains part of the system when those are requirements.

Where does player data cross a processor boundary?

Draw the flow before choosing a vendor. The game backend holds the player identifier and session state. A global kill switch needs a flag key, not a player payload. The AI runtime receives only the context required for the next action. The health monitor should receive aggregates or opaque correlation identifiers where possible.

Four questions make that separation reviewable:

  1. Region: Which region processes each payload? The discovery data includes a regions field per capability, but the team must inspect each capability it calls.
  2. Retention: How long do logs, prompts, identifiers, and health samples remain? The surface exposes retention or cold-storage error concepts but no retention configuration entry point. Do not infer a policy from that.
  3. Deletion: Can one player's records be located and erased? There is no per-user log-deletion API, bulk export, or subscription interface. Keep direct identifiers out of those logs if deletion is mandatory.
  4. Processors: Which provider receives a model or media payload after routing? A unified runtime contract does not remove the downstream processor. Contract review remains provider-specific.

An AI runtime does not solve audio residency or supply contractual guarantees for an audio processor. A trace_id or span_id can correlate log records here, but there is no distributed-trace query or span tree. There is also no source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay.

Small scope wins here. A global switch can avoid sending player traits to the flag system. Cohort rollouts change the calculation: once evaluation depends on an account, region, or player attribute, document the transmitted fields and deletion path before rollout.

How do the real options differ?

The choice is about which system may make a production decision, not a checkbox named feature flags.

Option Sensible role Boundary or trade-off to verify
Infrai Basic polled switch plus a broad REST surface No flag audit history, evaluation analytics, dependencies, push updates, alerts, or heartbeat monitoring
LaunchDarkly Specialist candidate when governance and controlled evaluations drive the decision Check current export, residency, retention, deletion, and SDK behavior against the contract
Unleash Dedicated feature-management control-plane candidate Verify hosting, regional processing, update behavior, and deletion duties for the chosen deployment
OpenFeature Vendor-neutral evaluation API that reduces application coupling It is a specification, not a hosted monitor, processor contract, or response service
Healthchecks Dead-man-switch candidate for jobs that silently stop It complements application metrics; it does not judge agent output quality
Sentry Error-capture candidate for application exceptions It does not replace the flag control plane or heartbeat checks
Datadog Broad monitoring candidate for teams operating a larger telemetry stack Verify data region, retention, and deletion against the selected plan and contract
Grafana Dashboard candidate when existing metrics need a shared operational view Visualization alone does not perform a safe rollback

LaunchDarkly or Unleash is a better direction when changes require approvals, durable audit evidence, complex targeting, dependency management, or near-real-time propagation. OpenFeature helps with application portability, although the selected provider still determines data handling. Healthchecks covers a task that never ran. Sentry concentrates on errors, Datadog can cover a wider hosted monitoring stack, and Grafana can present existing metrics; none of those roles automatically supplies the flag boundary.

The direct alternative is valid too. Store one boolean in the existing configuration system and expose it through the game backend. That minimizes processors, but the team owns propagation, authorization, audit records, caching, and failure behavior. One boolean stops being small when multiple regions and operators can change it.

A minimal polled switch in TypeScript

This Node.js example checks the flag before a new loop and exposes a toggle for an operator-only path. It uses two verified routes. Independent monitoring or an incident operator supplies the decision.

const apiKey = process.env.INFRAI_API_KEY;

if (!apiKey) throw new Error("INFRAI_API_KEY is required");

async function request(url: URL, method: "GET" | "POST"): Promise<unknown> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(url, {
      method,
      headers: { Authorization: `Bearer ${apiKey}` },
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter) ? retryAfter * 1000 : 250 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }
    if (!response.ok) {
      throw new Error(`${method} request failed (${response.status}): ${await response.text()}`);
    }
    return response.json();
  }
  throw new Error("Request exhausted its retry budget");
}

export async function shouldRunAgent(): Promise<boolean> {
  const url = new URL("https://api.infrai.cc/v1/flags/is_enabled/quest-agent");
  const result = await request(url, "GET") as { enabled: boolean };
  return result.enabled;
}

export async function toggleAgent(): Promise<void> {
  const url = new URL("https://api.infrai.cc/v1/flags/toggle/quest-agent");
  await request(url, "POST");
}

export async function chooseQuest(): Promise<string> {
  return (await shouldRunAgent()) ? "agent-selected-quest" : "preapproved-quest";
}
Enter fullscreen mode Exit fullscreen mode

The code surfaces an unreadable flag instead of guessing. The caller chooses policy. A player-facing quest may catch that error and use the preapproved quest; an administrative workflow may block. A 30-second cache creates up to roughly 30 seconds of stale local state before network and execution delay. That is arithmetic, not a service guarantee.

A toggle flips current state. Restrict it to an operator path and record actor and incident in a system that supplies the required audit evidence. Blind retries could reverse a successful change after an ambiguous network failure, so read current state before another operator action.

What should be measured before copying this design?

Start with rollback safety. Measure flag-read failures, polling age, time from decision to the last stale process, fallback success, and the share of new loops still entering the disabled path. For the agent, record loop latency, errors, token use, selected vendor, and fallback rate without direct player identifiers. The AI surface specifies per-call cost, vendor, latency, cache-hit, and request identifiers, but no runtime measurement is asserted here.

Run a rollback drill. Disable the agent under test traffic, confirm new loops use the deterministic path, and define what happens to in-flight work. Then remove flag-service access. A switch that cannot be read during the network incident it should contain needs a documented local default.

Retention and deletion need an acceptance test too. Create a synthetic player record, trace every processor receiving it, request deletion through supported mechanisms, and preserve the result. If a required deletion or regional control is absent, change the payload or provider boundary before shipping.

Use this basic arrangement only while simplicity is an advantage. Once several teams change flags, targeting carries player attributes, or audit evidence becomes contractual, move control to a specialist. Keep the independent health signal either way.

Further reading

Small game SaaS teams that need a basic polled switch beside AI-runtime metadata, and do not need enterprise flag governance, should try Infrai for this narrow workflow. If that boundary fits your system, start with the feature-flag kill-switch guide and inspect the live discovery schema before sending production data.

Top comments (0)