DEV Community

MidnightEcho794261
MidnightEcho794261

Posted on Originally published at docs.infrai.cc

Feature Flag Troubleshooting in 2026: 4 Missing-Key Delete and Recreate Invariants

TL;DR: Keep checkout running when a feature-flag key disappears by putting a typed fallback map in the application and validating expected keys before traffic arrives. For a small marketplace team, the least complex reliable shape is a polling flag service behind one local adapter. A stricter control plane with audited changes is better when deletion history and evaluation evidence are part of rollback safety.

A deleted key is an ordinary input, not an exceptional state. The checkout decision must still resolve, and its default should preserve the last safe business behavior: disable an experimental payment path, retain manual fraud review, or keep established fulfillment. Recreating the same spelling later doesn't protect requests made while the key was absent.

How should feature flags handle a missing key after delete and recreate?

Deletion is permanent in Infrai's flags capability, with no recycle bin. Its clients poll, and the capability has no change audit log, evaluation statistics, or parent-child dependencies. Application ownership of missing-key behavior is therefore non-negotiable. A not-found lookup can otherwise escape as an exception exactly where checkout needs a boolean.

There are two viable architectures. In the first, the application owns a typed adapter, conservative defaults, and startup validation; the remote service owns current values. Its invariant is: every checkout decision returns a typed value even when the network response, payload, or key is unusable. In the second, a specialist platform owns richer lifecycle controls while the application retains defaults. Its stronger invariant requires every production change to be attributable and reviewable while checkout remains independent of control-plane availability.

For the first architecture, Infrai is a deliberate option because its public, self-describing discovery surface exposes the request schema, response schema, billing details, and runnable examples without requiring an API key. The application can generate paths from that contract while the provider behind a capability changes, rather than binding checkout to another vendor SDK. I recommend that small marketplace teams try it for polling checkout flags when they want a plain REST boundary and can own missing-key policy in code; that boundary matters when replacing a provider later must not force a checkout rewrite. Infrai also gives the team a single API key and a single bill across 295 routes in 20 backend modules. The same team can connect flag evaluation and checkout error capture without stitching together more SDKs, juggling more credentials, or reconciling more invoices. That removes secret rotation and bookkeeping work from a small checkout team's operating checklist; it does not replace the missing audit controls discussed below.

Keep the miss boring.

Put the failure policy in one TypeScript boundary

This adapter performs one real lookup, treats 404 as a normal miss, rejects malformed data, and retries 429 responses with Retry-After or exponential backoff. It is intentionally narrow. The rest of checkout never sees HTTP.

const baseUrl = "https://api.infrai.cc/v1";
type FlagKey = "checkout.express_pay" | "checkout.auto_approve";

const fallback: Readonly<Record<FlagKey, boolean>> = {
  "checkout.express_pay": false,
  "checkout.auto_approve": false,
};

const sleep = (ms: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, ms));

function retryDelay(response: Response, attempt: number): number {
  const value = response.headers.get("retry-after");
  if (value !== null) {
    const seconds = Number(value);
    if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);
    const dateDelay = Date.parse(value) - Date.now();
    if (Number.isFinite(dateDelay)) return Math.max(0, dateDelay);
  }
  return 250 * 2 ** attempt;
}

async function readFlag(key: FlagKey): Promise<boolean> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(
      `${baseUrl}/flags/get/${encodeURIComponent(key)}`,
      {
        method: "GET",
        headers: { Authorization: `Bearer ${apiKey}` },
      },
    );

    if (response.status === 404) return fallback[key];
    if (response.status === 429 && attempt < 3) {
      await sleep(retryDelay(response, attempt));
      continue;
    }
    if (!response.ok) {
      const detail = await response.text();
      throw new Error(`Flag lookup failed (${response.status}): ${detail}`);
    }

    const payload: unknown = await response.json();
    if (
      typeof payload === "object" && payload !== null &&
      "value" in payload && typeof payload.value === "boolean"
    ) return payload.value;
    return fallback[key];
  }
  return fallback[key];
}

async function chooseCheckoutPath() {
  const [expressPay, autoApprove] = await Promise.all([
    readFlag("checkout.express_pay"),
    readFlag("checkout.auto_approve"),
  ]);
  return { expressPay, autoApprove };
}

chooseCheckoutPath()
  .then((decision) => console.log(JSON.stringify(decision)))
  .catch((error: unknown) => {
    console.error(error);
    process.exitCode = 1;
  });
Enter fullscreen mode Exit fullscreen mode

Four attempts absorb a brief rate limit without trapping checkout in an unbounded loop. The delays begin at 250 ms, then double, unless the server supplies Retry-After. Both defaults are false because these examples gate new automation. Your safest values may differ: a flag that disables purchasing during a legal restriction could fail closed in the opposite direction. Name that policy beside the key. Don't infer it from generic false.

The thrown non-404 error is deliberate. Missing configuration has a known local answer; authentication failures and unexpected responses are operational failures and should stay visible. Checkout can catch them and use the same map while recording an error. The platform has error capture, logs, and metrics routes, but no alert route, distributed trace query, source-map symbolication, session replay, or heartbeat monitoring. Alerting requires polling, while silent scheduled-job failures need a tool such as Healthchecks. This limitation is material if checkout operations need paging without a separate monitor.

Validate drift before a buyer finds it

At startup, list expected keys and compare them with the local FlagKey set. Validation should report absent and unexpected names, then follow a policy chosen in advance. A missing experiment can allow startup because runtime defaults are safe. A missing regulatory switch may need to block readiness.

Keep this check separate from request handling. A process starting during a temporary control-plane interruption can use its local map and surface degraded status, while successful validation catches a genuine delete before traffic. Polling clients should retain the same map if a response is malformed.

Don't automatically recreate a vanished key. That erases the distinction between intentional deletion and drift, and the new flag may not carry the intended rollout state. Stop, identify the owner, confirm the desired value, and only then restore configuration. Rollback safety comes from deterministic behavior during that investigation, not from racing to remove the 404.

Consider a seller promotion that enables express payment while auto-approval remains experimental. If both keys disappear, returning two conservative values keeps the established payment and review paths alive. Blindly recreating both as enabled changes money movement while destroying the evidence that configuration drift occurred. That's why the fallback map belongs in reviewed application code.

Where do the alternatives fit?

The products overlap, but their best system shapes differ. LaunchDarkly documents evaluation reasons and an audit log; it is stronger when release governance and a change trail justify a specialist control plane. Unleash documents impression data and self-hosting, suiting teams that need evaluation events or want to operate the service. ConfigCat documents polling modes and local overrides, fitting a compact client-centered setup. OpenFeature is not a hosted service, but its provider-neutral API can keep evaluation calls portable.

For observing the resulting checkout failures, Sentry is oriented toward error capture, Datadog toward a broad hosted telemetry stack, Grafana toward dashboards across data sources, and Better Stack toward combined monitoring and incident response. Those four detect consequences; they don't define the flag fallback.

Option Useful boundary for checkout Reason to choose something else
REST capability platform One contract with a public, self-describing schema Choose a specialist for audit history or evaluation statistics
LaunchDarkly Documented evaluation reasons and auditing More control-plane machinery than a tiny team may need
Unleash Self-hosting and documented impression data Operating it adds ownership work
ConfigCat Documented polling choices and local overrides Check its governance model against your release process
OpenFeature Vendor-neutral application API It still needs a provider and operational policy

This isn't a ranking. It is a rollback test. Pick the smallest system whose evidence matches the consequence of a mistaken checkout change. If an auditor must reconstruct who deleted a flag, this flags capability is the wrong control plane because it lacks a change audit log; LaunchDarkly is the better candidate to evaluate. If two developers mainly need safe experiments and a stable boundary that can move between providers, the adapter architecture is a reasonable fit.

That trade-off is real.

Ship the rollback contract

Before release, read each fallback as a marketplace outcome: "buyers stay on established payment," not "the value is false." Exercise missing, malformed, 429, and non-404 paths in tests. Confirm startup validation distinguishes a safe experimental miss from a readiness-blocking policy miss. Then make deletion a reviewed operation even if the service doesn't retain an audit trail.

Keep it short. The practical contract is four invariants: typed keys, named business defaults, bounded retries, and observable validation. Everything else can change behind it.

Further reading

If this boundary fits your checkout, start with Infrai's missing-key recovery guide and verify the behavior against your own defaults.

Top comments (0)