Short answer: treat the feature flag as permission to deliver an alert, not permission to detect or record the failure. A polling Node.js process may hold an old value until its next successful refresh, so the rollback-safe design puts a second, authoritative suppression check at the notification boundary. Keep evaluating scheduled marketplace imports, persist every state transition, and make notification sends idempotent.
That rule gives us three invariants: detection continues while delivery is muted; a stale worker cannot bypass the final suppression decision; and changing the flag never destroys the evidence needed to resume safely. The trade-off is one extra control-plane read on the rare path where an alert would actually leave the system. For email, SMS, or OTP-style delivery, that is the right place to spend latency: just before the irreversible side effect.
How can a feature flag disable noisy alerts quickly?
Consider a marketplace that expects each scheduled catalog import to produce at least one result. A scheduler records a run, an evaluator decides that the run has stopped producing results, and a notifier sends the message. The operational problem appears when the evaluator is correct but the notifications are too noisy. Operators need to mute delivery quickly without erasing the stalled-import signal.
The failure boundary matters. A flag cached inside each Node.js process is a hint during its polling interval, not an immediate global barrier. A process can miss a refresh, pause between refreshes, or evaluate work that was queued under an earlier value. None of those conditions justify another outbound message after an operator has requested suppression.
Detection stays on.
Keep the state machine independent of the delivery switch. For example, an import can move from healthy to stalled, remain stalled, and later return to healthy while delivery is disabled. Store those transitions with the import identifier and observation time. Do not rewrite stalled as healthy merely because messaging is muted.
Silence is a delivery state, not a system state.
This separation also keeps the monitoring model honest. The four golden signals described in the Google SRE book are latency, traffic, errors, and saturation; a stopped import is a symptom worth observing even when its paging policy is temporarily inappropriate. Muting the messenger must not blind the measurement path.
Decision record and failure boundaries
The decision is to use a cached flag for early shedding and an authoritative check immediately before the outbound side effect. The cached check prevents needless queue traffic in the common case. The final check supplies the rollback guarantee. If the authority cannot be reached, this alert class fails closed: it records a suppressed delivery attempt and sends nothing.
Delivery stops.
Failing closed is a deliberate compliance-aware choice. Marketplace import alerts are recoverable from stored evidence; an accidental duplicate SMS or email cannot be recalled. A different class, such as a life-safety alarm, would require a separate policy and should never inherit this decision casually.
| Option | Fast mute guarantee | Evidence preserved | Failure behavior | Operational cost |
|---|---|---|---|---|
| Polling cache only | No; bounded by refresh success and interval | Yes, if detection stays separate | A stale process may send | Lowest control-plane traffic |
| Authoritative check before send | Yes, at the delivery boundary | Yes | Can fail closed and record the reason | One read per candidate send |
| Stop the evaluator or scheduler | Yes, after work stops | No complete transition history | Recovery can create blind spots | Low immediate complexity, high rollback risk |
There is another edge case: work already in the notification queue. Suppose import catalog-42 becomes stalled, transition t-17 is persisted, and a notification item is queued while the cached flag still says enabled. An operator disables delivery before a consumer claims that item. Checking only at enqueue time leaves t-17 armed, even though the control-plane decision changed before the external side effect. The consumer therefore checks suppression after claiming the item and immediately before calling the transport. If it observes disabled, it records operator_muted against t-17 and acknowledges the queue item without sending. If the policy authority is unavailable, the same path records policy_unavailable. Put an idempotency key on the attempted delivery as well, because retries and multiple consumers are separate duplicate paths from stale flags. I choose the extra authoritative read because a stored suppression record can be reviewed and replayed; a duplicate external message cannot be withdrawn.
Queues complicate it.
Do not turn the flag service into the alert-state database. The flag answers a narrow policy question. The durable alert record answers what happened, when it was observed, whether delivery was attempted, and why it was suppressed. Mixing those responsibilities makes rollback depend on control-plane history that may not contain the underlying observations.
Critical path at the notification boundary
The following Python sketch describes the contract even if the production workers are Node.js. The names are intentionally generic. The example policy uses a 30-second local refresh and a 120-second stalled threshold; those are design inputs for this marketplace, not universal recommendations. The safety property does not depend on either value.
from dataclasses import dataclass
from datetime import datetime, timezone
@dataclass(frozen=True)
class AlertCandidate:
import_id: str
transition_id: str
observed_at: datetime
def evaluate_import(import_run, policy, alert_store):
if import_run.result_count == 0 and import_run.age_seconds >= 120:
candidate = AlertCandidate(
import_id=import_run.import_id,
transition_id=import_run.transition_id,
observed_at=datetime.now(timezone.utc),
)
alert_store.record_transition(candidate, state="stalled")
# The cache sheds work; it does not authorize the eventual send.
if policy.cached_delivery_enabled(refresh_seconds=30):
alert_store.enqueue_once(candidate.transition_id, candidate)
def deliver(candidate, authoritative_policy, alert_store, transport):
decision = authoritative_policy.read(
key="marketplace.import_stalled.delivery",
context={"import_id": candidate.import_id},
)
if decision.unavailable or not decision.enabled:
reason = "policy_unavailable" if decision.unavailable else "operator_muted"
alert_store.record_suppression(candidate.transition_id, reason=reason)
return
if alert_store.reserve_delivery_once(candidate.transition_id):
transport.send(candidate)
alert_store.record_delivery(candidate.transition_id)
The critical ordering is easy to miss. read occurs after queue claim and before transport.send; reserve_delivery_once closes the retry race. In a real implementation, the reservation needs a durable uniqueness constraint or an equivalent atomic operation. An in-memory set is not enough across Node.js processes.
A transport timeout leaves an ambiguous outcome: the provider may have accepted the message even if the client did not receive a response. The alert record should retain that ambiguity instead of blindly issuing another send. This is the same deliverability discipline used around OTP flows: distinguish “request failed locally” from “recipient definitely received nothing.”
Ambiguity survives retries.
Test the boundary with state transitions, not just flag values. Start with delivery enabled, enqueue a candidate, disable delivery before the consumer handles it, and verify that no transport call occurs while the stalled transition remains queryable. Then restore delivery and choose an explicit resumption policy: send only new transitions, or intentionally replay selected suppressed transitions. Do not let backlog replay happen by accident.
Why reject scheduler shutdown?
Stopping the scheduled evaluator looks attractive because it is immediate and easy to explain. It is rejected here because rollback safety is the primary decision axis. During the mute, the system would lose observations about which imports remained stalled or recovered, and restarting could confuse old failures with new transitions. That gap makes the quiet period operationally expensive even though it produces no messages.
Scheduler shutdown still has a valid use case. Use it when the evaluator itself is causing harm, such as overloading a dependency, and when the missing-observation window is explicitly accepted and recorded. It is an emergency brake for computation, not the normal mute control for a healthy detector with a noisy delivery policy.
A polling-only flag is also valid when delayed suppression is acceptable and the maximum staleness is part of the runbook. Internal dashboards often fit that profile. Outbound marketplace notifications do not, because every stale evaluation can cross an external boundary and create duplicate, confusing, or policy-sensitive contact.
Roll out the control without losing the rollback
Deploy the final suppression check before exposing the operator control. Older consumers that lack the check are the actual rollout hazard; a flag cannot stop code that never consults it. Track consumer-version coverage, drain or replace old workers, and only then treat the control as an authoritative mute.
Exercise four cases in staging: enabled delivery, disabled delivery, unavailable policy authority, and re-enablement with a suppressed backlog. The acceptance criteria should cover stored observations, queue behavior, transport call count, and idempotency. Logs are useful evidence, but they are not the durable state machine. Log ingestion and indexing can also be separate cost dimensions, so retaining concise structured events is better than emitting repetitive stack traces for every poll.
The operator runbook can stay short: disable delivery, confirm suppression records are increasing while transport sends stop, repair the alert rule, decide what to do with suppressed transitions, and re-enable. No scheduler restart is required. No evidence is discarded.
That is the decision. Preserve detection, authorize at the last responsible boundary, and resume from durable transitions. Polling can remain an efficiency mechanism without being mistaken for an instantaneous safety control.
Top comments (0)