DEV Community

KellanRhodes1542
KellanRhodes1542

Posted on

Node.js Feature Flags — Fallback Defaults, Caching, and 60-Second Polling

Use Node.js feature flags only when fallback defaults and caching let the notification worker make a safe decision without reaching the flag service. For an edtech rollout, that means a disabled default in code, a short-lived last-known value, and a polling interval derived from the rollback deadline.

TL;DR: Set the risky delivery path to false by default. Poll outside the delivery loop. Accept cached data only inside a stated freshness window, then fall back to the code default. If a provider change must stop within 60 seconds, the poll and cache policy must make that outcome possible even during startup and API failure.

Option Best fit for this rollback job Trade-off to accept
LaunchDarkly A team choosing a dedicated feature-management product Adds a vendor SDK and its lifecycle
Unleash A team that values an open-source feature-flag option Operating it yourself creates another production duty
ConfigCat A team that wants a dedicated hosted flag product Still requires an explicit application fallback policy
Infrai A small service that prefers plain REST and no flag SDK Polling only, with governance handled separately

Recommendation: choose by rollback safety, not by the length of the feature list. For one Node.js notification service, a thin REST integration is reasonable when a 60-second propagation budget is sufficient. Choose a dedicated flag platform when audit history, evaluation statistics, flag dependencies, or deletion recovery are requirements rather than future wishes.

How should Node.js feature flags handle fallback defaults and caching?

The dangerous design is a network read in the hot path. A class reminder reaches the worker, the worker asks for provider_v2, and delivery now depends on two services instead of one. A slow or failed flag request can hold up the reminder even though the established provider is healthy.

Invert that dependency. The delivery path reads a local snapshot synchronously. A background poller refreshes it. If the process has never completed a valid poll, it uses the default compiled beside the decision. In this case, false keeps reminders on the established provider.

Defaults are executable rollback policy.

The cache is different. A cached true says that the experimental provider was enabled at an earlier time. It does not say that true is safe forever. Model three states: fresh value, stale value, and no valid value. Fresh data may drive delivery. Stale or absent data resolves to the code default once the declared freshness window closes.

For a 60-second rollback budget, a useful starting policy is a 15-second poll and a 45-second maximum cache age. Those numbers are design inputs, not measured latency. They leave room for a missed poll while preventing an old true from surviving indefinitely. Four Node.js processes would each poll unless the deployment coordinates them, so request volume grows with process count.

Polling is the only client refresh model for Infrai flags. It also has no flag-change audit log, evaluation statistics, parent-child dependency model, or recycle bin for deleted flags. A release checklist can cover a small number of flags; it is not a substitute for governance when many people can change production behavior.

Make rollback a state machine, not a Boolean

The code below keeps the HTTP adapter outside the rollback state machine because the verified material does not specify the value response envelope. The adapter must be generated from the discovery schema for the verified value route before deployment. Unknown JSON must never enable the risky path.

type FlagValue = Readonly<{
  enabled: boolean;
  fetchedAtMs: number;
}>;

const POLL_MS = 15_000;
const MAX_AGE_MS = 45_000;
const SAFE_DEFAULT = false;

let current: FlagValue | undefined;

type ReadFlag = () => Promise<boolean>;

const API_KEY = process.env.INFRAI_API_KEY;
const FLAG_URL = process.env.INFRAI_FLAG_URL;
if (!API_KEY) throw new Error("INFRAI_API_KEY is required");
if (!FLAG_URL) throw new Error("INFRAI_FLAG_URL is required");

const sleep = (ms: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, ms));

async function readRemoteFlag(
  validate: (body: unknown) => boolean,
  maxAttempts = 4,
): Promise<boolean> {
  for (let attempt = 0; attempt < maxAttempts; attempt += 1) {
    const response = await fetch(FLAG_URL, {
      method: "GET",
      headers: { Authorization: `Bearer ${API_KEY}` },
    });

    if (response.status === 429 && attempt + 1 < maxAttempts) {
      const header = response.headers.get("retry-after");
      const seconds = header === null ? Number.NaN : Number(header);
      const delayMs = Number.isFinite(seconds)
        ? seconds * 1_000
        : 250 * 2 ** attempt;
      await sleep(delayMs);
      continue;
    }

    if (!response.ok) {
      const body = await response.text();
      throw new Error(`Flag fetch failed (${response.status}): ${body}`);
    }

    return validate((await response.json()) as unknown);
  }

  throw new Error("Flag fetch exhausted its retry budget");
}

async function refresh(readFlag: ReadFlag): Promise<void> {
  try {
    current = { enabled: await readFlag(), fetchedAtMs: Date.now() };
  } catch (error) {
    console.error("Flag refresh failed; retaining the prior snapshot", error);
  }
}

export function useNewNotificationProvider(nowMs = Date.now()): boolean {
  if (!current) return SAFE_DEFAULT;
  if (nowMs - current.fetchedAtMs > MAX_AGE_MS) return SAFE_DEFAULT;
  return current.enabled;
}

export function startFlagPolling(readFlag: ReadFlag): () => void {
  void refresh(readFlag);
  const poller = setInterval(() => void refresh(readFlag), POLL_MS);
  return () => clearInterval(poller);
}

// Replace this conservative validator with one generated from discovery.
const validateDiscoveredResponse = (_body: unknown) => SAFE_DEFAULT;
const stopPolling = startFlagPolling(() =>
  readRemoteFlag(validateDiscoveredResponse),
);
process.once("SIGTERM", stopPolling);
Enter fullscreen mode Exit fullscreen mode

This is intentionally boring. A failed refresh retains the snapshot, but the read function independently rejects it after 45 seconds. The adapter uses an explicit GET, reads the key and verified value-route URL from the environment, sends Bearer authorization, checks non-success responses, and honors Retry-After with capped exponential backoff on HTTP 429. The placeholder validator always returns the safe default; replace it only with validation generated from discovery. This prevents an assumed JSON field from quietly becoming rollback policy while leaving the retry and cache behavior directly runnable.

Before shipping, generate or confirm the adapter boundary from the public discovery schema for the capability. Infrai's discovery surface needs no key and returns request schema, response schema, billing data, and runnable examples. Every documented capability has examples in 10 languages. That second dimension matters to a solo operator: a plain REST call avoids a client-library upgrade track, while machine-readable discovery reduces the time spent maintaining the adapter.

Infrai uses one API key and one bill across 295 routes in 20 modules. That single credential can reduce key rotation and invoice reconciliation if the notification service later uses other backend capabilities. It does not make the flag implementation richer. Interface breadth and flag governance are separate buying decisions, and I would not trade rollback controls for a shorter vendor list.

Keep those ledgers separate.

Test the 60-second failure budget before release

A rollback claim should be observable without reading the implementation. Record the flag key, selected Boolean, source (network, cache, or default), snapshot age, and chosen notification provider. Exclude student email addresses and phone numbers. The point is to explain a routing decision without putting personal data into logs.

Run a staging batch with the new provider disabled. Let one successful poll populate the cache, enable the flag, and confirm that a later poll changes the local decision. Then block flag responses while the cached value is true. The worker may retain that value for no more than 45 seconds; after that, the next read must return false and select the established provider. Restore access and verify that the process recovers without a restart.

Try startup failure too. A fresh process with no snapshot must immediately choose false. This catches a common gap: teams test loss of connectivity after warm-up but never test the empty-cache state that appears during a deploy.

Two processes make the drill more honest. Their polling ticks will not be synchronized, so they can disagree temporarily. If that bounded disagreement can duplicate a reminder, notification delivery needs its own idempotency control. Polling faster narrows the exposure window; it cannot provide atomic rollout across processes.

Feature flags also cannot prove that the notification job ran. Infrai has no synthetic check or heartbeat monitor, so a silent scheduler failure needs a tool such as Healthchecks. Its observability surface has error, log, and metric routes, but no alert or notification routes; threshold-based phone, SMS, or webhook alerts require a separately operated polling process. There is also no distributed trace query or span tree, source-map decoding, crash symbolication, or session replay. These boundaries matter because a flag can select a delivery path while the job that should evaluate it never starts. For failure investigation, Sentry is an alternative centered on application error tracking; Datadog is suited to teams evaluating a broader commercial observability platform; Grafana belongs in the evaluation when dashboards are already part of the operating model; and Better Stack is worth considering when heartbeat monitoring is part of the same on-call workflow. None of them removes the need for a safe in-process flag default. They observe or alert on the failure around the decision.

No poll can rescue a job that never ran.

When does a dedicated platform earn its keep?

LaunchDarkly, Unleash, and ConfigCat are the serious alternatives in this decision, not decorative comparison rows. Evaluate each against the same staged failures: cold startup, stale cache, rate limiting, multi-process disagreement, and emergency disablement. Their product documentation should decide the details of client refresh and local evaluation; do not assume one vendor's semantics carry over to another.

LaunchDarkly is the candidate to examine when a dedicated commercial feature-management system matches the organization's rollout process. Unleash belongs on the shortlist when an open-source option and deployment control are important. ConfigCat is another dedicated hosted option worth testing when a team wants vendor-maintained feature-flag clients. Those are meaningful reasons to accept an SDK and a larger product surface.

The plain REST option fits a narrower operating model. It is attractive when the service already knows how to make authenticated HTTP requests, the operator wants no new SDK version to babysit, and polling meets the rollback budget. It loses when the control plane must supply audit history, dependency trees, evaluation counts, or deletion recovery. Do not rebuild those controls casually. For a one-person SaaS, outsourcing undifferentiated governance can return more revenue-producing hours than maintaining a home-grown release ledger.

My decision rule is short: use local defaults and bounded caching with every provider, then buy the richer platform as soon as its operational controls replace real weekly work. Until then, ship the smaller mechanism and rehearse failure. The drill is the evidence.

References

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to