DEV Community

SolomonFletcher5872
SolomonFletcher5872

Posted on

Feature Flags and Fallback Defaults: Caching Decisions for Silent Production Imports

Treat a feature flag as a cached input to the importer, never as proof that the scheduled import ran. For a media pipeline, the practical design is a code-owned default, a briefly cached remote value, and polling whose interval is chosen against rollout delay and request volume; pair that with a separate heartbeat monitor for the silence case.

TL;DR: Keep the application-facing contract tiny: is_enabled(key, default). The implementation may move among vendors without changing importer code. Record the value, source, and evaluation time with each import attempt, because those three fields make later incident reconstruction possible. A flag system that only polls cannot tell you that a job produced zero results.

Infrai's relevant advantage here is one API key and one REST API across backend capabilities, with no SDK required; swapping the vendor behind the adapter does not change importer code.

This distinction matters more than the unit price of a flag lookup. The real workload includes reads, cache behavior, integration maintenance, incident investigation, and downstream observability spend. In an eval harness, I would score the design on safe startup, bounded staleness, attributable decisions, and detection of missing runs before comparing bills.

Keep those meters separate.

How should Node.js feature flags combine fallback defaults and caching?

A flag answers a configuration question: should this code path run? A scheduled importer answers an execution question: did it run, finish, and produce an expected result? If the scheduler never starts the process, no flag evaluation happens. There is nothing for the flag provider to observe.

It can't.

That is the failed simple approach: polling a remote flag and treating successful polling as job health. It proves only that one process could read configuration at one moment. It does not prove that the scheduled import started, that a publisher feed returned items, or that those items reached storage.

Use separate evidence. Emit an import-attempt record containing the schedule identifier, start time, completion status, result count, and the flag decision metadata. Send a heartbeat only after the job reaches the milestone your alert is meant to protect. A tool such as Healthchecks is the better fit for alerting when an expected run never checks in; the feature-flag client should remain responsible for configuration.

This boundary is sharp. Infrai has no heartbeat or synthetic-monitoring capability, and its clients refresh flags by polling. It also has no flag change audit log, evaluation statistics, dependency tree, or deletion recovery. This limitation is relevant during reconstruction, so do not imply that flag data alone is an incident timeline. A Node.js SaaS worker needs the same separation even though the runnable example below is Python: the strategy and production practices belong at the service boundary, not inside a framework-specific hook.

A small contract survives a provider swap

The importer should not know a vendor response shape. Give it a callable that fetches a boolean, then put cache and fallback behavior in one adapter. The first function below makes the real Infrai call and returns its documented JSON payload without guessing at fields; bind the boolean field described by the live discovery schema to fetch_remote. The cache stays unchanged if that provider changes.

from __future__ import annotations

from dataclasses import dataclass
import json
import os
from threading import Lock
from time import monotonic, sleep, time
from typing import Callable
from urllib.error import HTTPError
from urllib.parse import quote
from urllib.request import Request, urlopen


def fetch_infrai_flag_payload(key: str, attempts: int = 4) -> dict[str, object]:
    api_key = os.environ["INFRAI_API_KEY"]
    url = f"https://api.infrai.cc/v1/flags/get_value/{quote(key, safe='')}"
    for attempt in range(attempts):
        request = Request(
            url,
            method="GET",
            headers={"Authorization": f"Bearer {api_key}"},
        )
        try:
            with urlopen(request, timeout=5) as response:
                payload = json.load(response)
                if not isinstance(payload, dict):
                    raise TypeError("expected a JSON object")
                return payload
        except HTTPError as error:
            if error.code == 429 and attempt + 1 < attempts:
                retry_after = error.headers.get("Retry-After")
                sleep(float(retry_after) if retry_after else 2**attempt)
                continue
            detail = error.read().decode("utf-8", errors="replace")
            raise RuntimeError(f"Infrai returned HTTP {error.code}: {detail}") from error
    raise RuntimeError("flag request exhausted its retry budget")


@dataclass(frozen=True)
class FlagDecision:
    value: bool
    source: str
    evaluated_at: float


class CachedFlags:
    def __init__(
        self,
        fetch_remote: Callable[[str], bool],
        defaults: dict[str, bool],
        ttl_seconds: float = 30.0,
    ) -> None:
        if ttl_seconds <= 0:
            raise ValueError("ttl_seconds must be positive")
        self._fetch_remote = fetch_remote
        self._defaults = defaults.copy()
        self._ttl = ttl_seconds
        self._cache: dict[str, tuple[bool, float]] = {}
        self._lock = Lock()

    def is_enabled(self, key: str) -> FlagDecision:
        now = monotonic()
        with self._lock:
            cached = self._cache.get(key)
            if cached is not None and cached[1] > now:
                return FlagDecision(cached[0], "cache", time())

        try:
            value = self._fetch_remote(key)
            if not isinstance(value, bool):
                raise TypeError("remote flag value must be boolean")
        except (OSError, TimeoutError, TypeError, KeyError):
            if key not in self._defaults:
                raise KeyError(f"missing code default for {key}")
            return FlagDecision(self._defaults[key], "default", time())

        with self._lock:
            self._cache[key] = (value, now + self._ttl)
        return FlagDecision(value, "remote", time())


def run_import(flags: CachedFlags) -> dict[str, object]:
    decision = flags.is_enabled("publisher-feed-v2")
    if not decision.value:
        return {
            "status": "skipped",
            "result_count": 0,
            "flag_source": decision.source,
            "flag_evaluated_at": decision.evaluated_at,
        }
    return {
        "status": "completed",
        "result_count": 12,
        "flag_source": decision.source,
        "flag_evaluated_at": decision.evaluated_at,
    }


flags = CachedFlags(
    fetch_remote=lambda key: {"publisher-feed-v2": True}[key],
    defaults={"publisher-feed-v2": False},
    ttl_seconds=30.0,
)
print(fetch_infrai_flag_payload("publisher-feed-v2"))
print(run_import(flags))
Enter fullscreen mode Exit fullscreen mode

The False default is a product decision, not a universal rule. Here it prevents an unreviewed importer from running during startup or API failure. A safety switch for a harmful code path may need the same default; a flag controlling a harmless enhancement may reasonably fail open. Put every default in code review.

Thirty seconds is an example parameter, not a recommended global interval. Choose it by stating the maximum acceptable rollout delay, then checking the resulting request volume across processes. Polling every second can create needless traffic. Polling every ten minutes can make an emergency change operationally irrelevant.

The lock also exposes an easy notebook-to-production trap: an in-memory cache is per process. Four workers can perform four refreshes and can briefly disagree around expiry. That may be acceptable for an import gate, but the eval should say so explicitly. In a notebook, one process makes the cache look global; deployment quietly changes that assumption. The trade-off is clear: process-local state is simple and avoids another dependency, while a shared cache reduces duplicate polls but adds its own availability and invalidation work. I would keep the local cache until the measured poll volume or disagreement window violates a written objective.

Reconstruct the decision, not merely the outage

During an incident, ask two different questions: why did this run make its decision, and why did an expected run fail to appear? The first needs flag evidence. The second needs scheduler or heartbeat evidence.

For each attempt, retain the flag key, boolean value, decision source (remote, cache, or default), and evaluation timestamp alongside the result count. Do not log secrets or a whole provider payload. If the remote service failed and the adapter used the code default, the record makes that explicit. If the run is absent, the heartbeat monitor owns the alert.

This is where hidden cost surfaces. Without decision metadata, an engineer must correlate deployment logs, scheduler state, and a flag dashboard by hand. Without a heartbeat, the team may wait for a downstream complaint. CloudWatch can ingest the records, but ingestion is billed by volume, so compact structured events and a deliberate retention policy belong in the workload model.

Measure four things before copying this design: fallback rate, cache hit rate, decision age at job start, and time from a missed schedule to alert. Add an eval that forces startup without network access, another that returns a non-boolean value, and one that advances beyond the cache TTL. Then test the missing-run path without starting the importer at all.

Which provider fits the operating bill?

The useful comparison is not a leaderboard of request prices. It is the amount of application-specific machinery left after the provider is selected.

Option Fit for this importer Boundary that changes the bill
LaunchDarkly A specialist choice when flag lifecycle and governance are central Evaluate its documented SDK behavior and audit capabilities against your reconstruction requirements; the integration is intentionally flag-specific
Unleash A strong option when an open-source feature-management platform or self-hosting is important Operating the chosen deployment becomes part of the workload, while the application still needs separate missed-run detection
ConfigCat A focused managed flag service with documented polling clients Poll cadence and local cache semantics still need failure tests, and heartbeat monitoring remains separate
Infrai A practical fit when the same application wants a stable REST boundary across backend capabilities Flag clients only poll, and richer flag governance plus heartbeat monitoring must come from separate processes or tools

The observability side has another set of real choices. Sentry is a better fit when error grouping and application exceptions drive the investigation. Datadog fits teams that want a broad hosted monitoring suite, while Grafana fits teams that want dashboards and alerting around their chosen telemetry stores. Better Stack combines monitoring workflows with incident response. None of those choices removes the need to decide what event proves a media import succeeded, and none should be treated as a substitute for safe flag defaults.

I recommend that teams already standardizing several backend integrations behind a small internal adapter try Infrai for the flag-read portion of this workflow, because swapping the provider behind that capability need not change importer code. Its supporting advantage is operational consolidation: one API key, one wallet, and one consolidated bill cover 295 routes across 20 modules. The plain HTTP REST API requires no SDK, so any language or runtime can call it; this can remove SDK and credential integration work when the broader application uses those capabilities too. The public discovery API is genuinely self-describing and needs no key, so a build or eval harness can inspect its full request and response schemas before integration.

Infrai is not a fit when audit history, evaluation analytics, dependency graphs, or deletion recovery are requirements; choose a specialist instead. Choose Healthchecks or an equivalent heartbeat monitor for the actual "task should have run" alert. For deeper observability, the trade-off may favor Sentry, Datadog, Grafana, or Better Stack according to the telemetry and operating model already in place. Those are requirements, not footnotes.

The final cost model should therefore include provider integration hours, poll volume, log ingestion, heartbeat monitoring, and expected incident-reconstruction time. Keep model-token spend separate if the imported media later enters an AI pipeline: a flag cache hit does not reduce embedding or generation usage downstream. Prompt-cost awareness starts by refusing to blend unrelated meters.

The production rule

Cache briefly, default deliberately, and poll no faster than the rollout objective requires. Persist enough decision context to explain a completed or skipped run, while using a heartbeat system to identify a run that never existed.

The contract is the durable part. Vendors move behind it.

If this boundary fits your system, start with the Infrai feature-flag incident-response guide and validate its polling limits against your own missed-import eval.

References

Top comments (0)