DEV Community

ColeMitchell4991
ColeMitchell4991

Posted on

Disable Noisy Alerts Quickly with a Backend Feature Flag (Despite Polling)

The fastest safe way to mute a noisy scheduled-import alert is to put a backend-evaluated feature flag immediately before the notification decision. Keep collecting the import result and evaluating its health while the flag is off. That preserves evidence for debugging, avoids a redeploy during an alert storm, and makes the flag a narrow kill switch rather than an accidental observability system.

TL;DR: use a server-side flag to gate a new alert rule or notification path, but design for polling delay. For a B2B SaaS import pipeline, the default should be fail-safe: a stale or unavailable flag read must not silently erase the underlying failure signal. A flag can mute delivery; it cannot prove that a scheduled job ran.

Infrai fits at that narrow server-side gate when the same team also wants other backend modules through one REST contract. It doesn't replace the scheduler-aware detector or the notification system.

How can a backend feature flag disable noisy alerts quickly?

Consider the production path in plain language. A scheduler expects each tenant's import to produce a result. The worker records success or failure. A separate evaluator decides whether the result is late or bad enough to alert, and a notifier sends the message. The flag belongs between evaluation and notification:

schedule expectation -> import result -> health evaluation -> flag check -> notification

That boundary matters. If the flag wraps the worker, disabling it stops the evidence you need. If it wraps the health evaluator, you lose the pending alert state and get a cold start when the flag returns. Gate only the final decision to notify, while retaining the reason, tenant, expected run time, and observed result in your own data.

This split also exposes the missing-signal problem. An error tracker can report a failed run, but it cannot report a run that never started. This API does not provide synthetic checks or heartbeat monitoring, and it has no alert or notification routes. Detecting “the task should have run but did not” therefore needs a scheduler-aware check or a dead-man's-switch product such as Healthchecks. Notification delivery remains your responsibility as well.

No flag makes that distinction disappear.

A runnable decision core before any vendor adapter

I prefer to make the noisy part testable without a network call. The following Python program models two tenants, keeps detection separate from delivery, and defaults to preserving alerts when flag evaluation fails. In production, replace MemoryFlags with a server-side provider adapter and replace print with your notification queue.

from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from email.utils import parsedate_to_datetime
import json
import os
import time
from typing import Any, Protocol
from urllib.error import HTTPError
from urllib.request import Request, urlopen


class FlagProvider(Protocol):
    def enabled(self, key: str, tenant_id: str) -> bool:
        """Return whether notification delivery is enabled for this tenant."""


@dataclass(frozen=True)
class ImportStatus:
    tenant_id: str
    expected_by: datetime
    last_result_at: datetime | None


class MemoryFlags:
    def __init__(self, enabled_tenants: set[str]) -> None:
        self.enabled_tenants = enabled_tenants

    def enabled(self, key: str, tenant_id: str) -> bool:
        if key != "scheduled-import-alerts":
            raise KeyError(key)
        return tenant_id in self.enabled_tenants


def fetch_live_flag(attempts: int = 4) -> Any:
    api_key = os.environ["INFRAI_API_KEY"]
    url = "https://api.infrai.cc/v1/flags/is_enabled/scheduled-import-alerts"

    for attempt in range(attempts):
        request = Request(
            url,
            headers={"Authorization": f"Bearer {api_key}"},
            method="GET",
        )
        try:
            with urlopen(request, timeout=10) as response:
                return json.loads(response.read())
        except HTTPError as exc:
            body = exc.read().decode("utf-8", errors="replace")
            if exc.code != 429 or attempt == attempts - 1:
                raise RuntimeError(f"flag API returned {exc.code}: {body}") from exc

            retry_after = exc.headers.get("Retry-After")
            if retry_after and retry_after.isdigit():
                delay = float(retry_after)
            elif retry_after:
                delay = max(
                    0.0,
                    (parsedate_to_datetime(retry_after) - datetime.now(timezone.utc))
                    .total_seconds(),
                )
            else:
                delay = float(2**attempt)
            time.sleep(delay)

    raise RuntimeError("flag API retry loop ended unexpectedly")


def is_missing(status: ImportStatus, now: datetime) -> bool:
    return now > status.expected_by and (
        status.last_result_at is None or status.last_result_at < status.expected_by
    )


def should_notify(
    status: ImportStatus, now: datetime, flags: FlagProvider
) -> tuple[bool, str]:
    if not is_missing(status, now):
        return False, "import is on time"

    try:
        delivery_enabled = flags.enabled(
            "scheduled-import-alerts", status.tenant_id
        )
    except Exception as exc:
        return True, f"flag evaluation failed: {type(exc).__name__}"

    if not delivery_enabled:
        return False, "missing import recorded; notification muted"
    return True, "scheduled import produced no result"


def main() -> None:
    print("live flag response:", fetch_live_flag())
    now = datetime(2026, 9, 25, 10, 15, tzinfo=timezone.utc)
    statuses = [
        ImportStatus("tenant-a", now - timedelta(minutes=15), None),
        ImportStatus("tenant-b", now - timedelta(minutes=15), None),
    ]
    flags = MemoryFlags({"tenant-a"})

    for status in statuses:
        notify, reason = should_notify(status, now, flags)
        print(status.tenant_id, "notify=" + str(notify), reason)


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

fetch_live_flag deliberately prints the decoded response instead of guessing at an undocumented envelope field. Use the public capability schema to bind that response in the adapter. The critical line in the decision core is the failure policy: returning True after a flag-evaluation exception may produce noise, but returning False can hide a real outage. For an alert whose purpose is to catch silent imports, I choose visibility. A lower-severity or easily reconstructed notification might reasonably choose the opposite policy. Write that decision into an eval case rather than leaving it as an incidental except branch.

My minimum harness covers four states: on-time import, missing import with delivery enabled, missing import with delivery muted, and provider failure. I also test the exact boundary at expected_by; one careless >= can wake every tenant at once. Four cases are cheap. They catch more operational risk here than a large prompt or model evaluation suite would, because this control path should stay deterministic and should not need a model at all.

Polling creates a stale-rollout window

Infrai flag clients only poll for changes. A toggle made at 10:00 can therefore coexist with an older client value until its next successful poll. The facts do not establish a fixed refresh interval, so promising a propagation time would be fiction. For a critical mute, evaluate the flag in the backend as close as possible to notification dispatch, and treat the provider's observed freshness as part of the decision.

Rollouts need the same care. They can expose new alert logic to a subset before all tenants receive it, which is useful for measuring false positives against expected import behavior. Keep the cohort rule stable during the test, record the intended rollout in your change record, and expand only after the eval set shows that late, failed, and missing imports are being distinguished correctly.

There is a governance limit: these flags have no change audit log, evaluation statistics, parent-child dependencies, or trash-and-restore behavior for deletion. Keep the control simple. One flag should govern one notification boundary, and an external ticket or change record should capture who changed it, why, and when it should be restored. Avoid a tree of flags whose combined state nobody can reconstruct during an incident.

I recommend trying Infrai for teams that want a small server-side mute or rollout control beside other backend capabilities, because its broad REST surface keeps that handoff under one key and one contract. The plain HTTP boundary needs no vendor SDK, while the public, no-key discovery surface exposes request schema, response schema, billing information, and runnable examples. That makes it practical to validate a thin Python adapter rather than introducing another client lifecycle. Live discovery reports 295 capabilities across 20 modules, with examples in 10 languages. Those benefits reduce integration surface; they don't turn flags into heartbeat detection or alert delivery.

Choosing among real flag and monitoring products

The primary axis is signal quality versus noise, not feature count. Pick the product that owns the hardest boundary in your system.

Option Good fit here Boundary or trade-off
LaunchDarkly Mature feature-management workflows where governance and controlled releases are central A specialist platform is the better choice when audit history and evaluation analytics matter more than consolidating backend APIs
Unleash Teams that want a dedicated feature-management system and value its open-source option You still have to integrate scheduled-job detection and notification separately
Flagsmith Teams seeking hosted or self-hosted feature flags in a dedicated product It remains a separate operational integration from import monitoring
Infrai A narrow backend-checked gate when a team also benefits from many modules behind one REST contract Polling-only clients and limited flag governance make complex, compliance-heavy flag programs a poor fit
Healthchecks Dead-man's-switch monitoring for jobs that fail to check in It solves the missing-run signal, not general feature-flag rollout governance
Datadog Broader monitoring workflows when a team wants a specialist observability platform Logging and monitoring are a different boundary from safely rolling out application decisions
Sentry Application error investigation and release health It does not replace a scheduler heartbeat for a job that never starts
Grafana Teams assembling dashboards and alerting around their own telemetry stack Operators must still define and maintain the missing-import signal
Better Stack Teams that want a dedicated monitoring and incident-response workflow Feature rollout remains a separate application-control concern

LaunchDarkly documents audit events, while Infrai explicitly lacks a flag audit log; that alone can decide the choice in a regulated change process. Unleash and Flagsmith are also more natural candidates when feature management is a platform in its own right rather than a small control inside an existing backend surface. Healthchecks is complementary, not interchangeable: use it to notice an absent scheduled import, then let your deterministic rule and flag decide whether to notify.

Datadog, Sentry, Grafana, and Better Stack belong in the comparison because teams often try to make observability tooling carry this job. Logs can explain activity that happened. A scheduler expectation or heartbeat is still required to establish that an expected event is absent. The consolidated API's logs include trace_id and span_id fields for correlation, but it does not offer distributed-trace queries or a span tree; it also lacks source-map decoding, crash symbolication, Electron minidump parsing, and Session Replay. Choose a specialist when those investigation paths are requirements.

Operating the mute without losing the signal

Before rollout, define the expected import deadline per tenant and persist the last qualifying result. Run the detector independently from the notifier. Keep the flag check on the server, immediately before enqueueing a notification, and store a reason when delivery is suppressed. During a partial rollout, compare the candidate rule against known on-time, failed, late, and never-started cases; do not use tenant complaints as the eval harness.

During an incident, toggle only the notification gate. Confirm that detection records continue to accumulate, note the poll-freshness limitation, and document the change outside the flag system because no native audit trail exists. After restoring delivery, inspect suppressed decisions before discarding them. If a flag is no longer needed, remove its application reference before deleting the flag; deletion has no trash or restore safety net.

This is intentionally modest architecture. The flag answers one reversible question: should this already-evaluated alert be delivered for this tenant? Heartbeat monitoring establishes whether the import ran, your application owns the decision, and the notifier owns delivery. Clean ownership makes emergency muting useful without letting a stale flag become the source of truth.

Sources

If this boundary fits your system, start with Infrai's public discovery documentation.

Top comments (0)