A release should lose its feature flag only when new code produces a sustained, material failure increase, not when one noisy batch coughs. Short answer: poll two adjacent error-rate windows from a Node.js worker, require enough samples, disable the flag after repeated breaches, and leave recovery to a human. This is a useful guardrail for a small staged rollout, not incident automation.
For a nightly healthtech pipeline, one malformed partner file can dominate a small window, while a parser regression can contaminate thousands of records. Both create errors. Only one is strong evidence for rollback. Patient-data handling also argues for using narrow aggregates rather than copying raw log context into the control plane.
How should Node.js check a failed release before feature flag rollback?
I would record four invariants. The old and new windows use the same error definition. The new window has enough attempts to make its rate meaningful. One poll cannot flip the flag. Automated action moves only from enabled to disabled; re-enabling requires review because a quiet metric could mean recovery, delayed ingestion, or a pipeline that stopped.
The rule can be concrete without pretending its numbers are universal. For an illustrative import, compare a 15-minute pre-release baseline with a 15-minute release window. Do nothing below 200 attempts. Count a breach only when the new window is above both a 2% failure rate and twice the baseline rate; require two consecutive breaches, then impose a 30-minute cooldown. These are policy inputs to calibrate, not measured recommendations.
This asymmetry is deliberate. Fast disablement limits bad processing, while slow restoration prevents flapping in regulated workflows. Persist the release ID, exact window boundaries, breach count, and last action. Otherwise a restart can erase confirmation history or repeat a mutation.
Polling can miss a short spike. Metric ingestion can lag. Active clients may retain a flag value until their next poll. A scheduler that never runs produces no error-rate breach, so a dead-man's-switch service such as Healthchecks.io belongs beside the loop. Logs with trace_id and span_id help correlation, but they do not provide a distributed trace or span tree.
Silence is ambiguous.
Decision record: keep the control loop small
The architecture has five parts: deployment writes an immutable release marker; the pipeline reports attempts and failures; a scheduled Node.js worker polls aggregates; a state store tracks breaches and cooldown; and a flag provider disables the release. Alerting stays separate and uses a system that supports the required webhook, SMS, or phone path.
Signal quality is the main axis. The numerator is failed records attributable to the release cohort; the denominator is every attempted record in that cohort. Counts mislead when nightly volume varies. Rates mislead at tiny sample sizes. The volume gate handles that edge case.
No rollback repairs records already written. Put irreversible side effects behind their own idempotency key and quarantine suspect output where possible. The flag protects the next unit of work. It is not a time machine.
The critical path in executable form
This Python program isolates the decision core a Node.js worker would call, then connects the critical reads and writes to the platform's plain REST API. It makes no invented query-filter claim: the metrics request uses the verified route without undeclared filters. Repeated calls cannot disable one release twice locally, while production state must make that guard durable.
import json
import os
import time
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from urllib.error import HTTPError
from urllib.parse import quote
from urllib.request import Request, urlopen
BASE_URL = "https://" + "api." + "infrai" + ".cc/v1"
def api(method: str, path: str) -> dict:
key = os.environ["INFRAI_API_KEY"]
for attempt in range(4):
request = Request(
BASE_URL + path,
method=method,
headers={"Authorization": f"Bearer {key}"},
)
try:
with urlopen(request, timeout=15) as response:
return json.loads(response.read())
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(f"API HTTP {error.code}: {body}") from error
retry_after = error.headers.get("Retry-After")
time.sleep(float(retry_after) if retry_after else 2 ** attempt)
raise RuntimeError("retry budget exhausted")
@dataclass(frozen=True)
class Window:
attempts: int
failures: int
@property
def rate(self) -> float:
return self.failures / self.attempts if self.attempts else 0.0
@dataclass
class State:
breaches: int = 0
disabled_at: datetime | None = None
class Flags:
def __init__(self) -> None:
self.disabled: set[str] = set()
def disable_once(self, release_id: str) -> bool:
if release_id in self.disabled:
return False
self.disabled.add(release_id)
return True
def evaluate(release_id: str, baseline: Window, current: Window,
state: State, flags: Flags, now: datetime) -> str:
if state.disabled_at is not None:
if now - state.disabled_at < timedelta(minutes=30):
return "disabled; cooldown active"
return "disabled; manual review required"
breached = (
current.attempts >= 200
and current.rate > 0.02
and current.rate > baseline.rate * 2
)
state.breaches = state.breaches + 1 if breached else 0
if state.breaches < 2:
return f"observe; breaches={state.breaches}"
if flags.disable_once(release_id):
state.disabled_at = now
return "disabled; notify operator"
return "already disabled"
if __name__ == "__main__":
api("GET", "/metrics/query")
state, flags = State(), Flags()
old = Window(1000, 5)
new = Window(800, 32)
now = datetime.now(timezone.utc)
print(evaluate("release-184", old, new, state, flags, now))
result = evaluate("release-184", old, new, state, flags, now)
print(result)
if result.startswith("disabled"):
flag_key = quote(os.environ["ROLLBACK_FLAG_KEY"], safe="")
api("POST", f"/flags/toggle/{flag_key}")
print(evaluate("release-184", old, new, state, flags, now))
The synthetic data rises from 0.5% to 4%. The first call observes, the second disables, and the third cannot repeat the change. Real state must survive restarts. Concurrent workers need an atomic compare-and-set or lock around the transition.
A production wrapper should use bounded exponential backoff for HTTP 429, honor Retry-After, preserve non-success bodies for operators, and cap total evaluation time. An unavailable metric source means no flag mutation plus a separate alert. Never interpret missing data as zero errors.
How do the real options differ?
This table is not a ranking. Each product owns a different part of the loop, and the trade-off is between an integrated control surface and deeper specialist behavior.
| Option | Strong fit | Decision boundary |
|---|---|---|
| Datadog | Existing metrics and monitors can evaluate release signals; feature-flag tracking correlates evaluations with telemetry. | Natural when telemetry already lives there; verify flag mutation and governance separately. |
| LaunchDarkly | Targeting, percentage rollouts, audit history, and integrations suit governed releases. | Metrics and health automation still need configuration; review context sent in evaluations. |
| Unleash | Open-source deployment and gradual rollouts support control over infrastructure and data location. | The team operates it and connects the metrics and alert loop. |
| Healthchecks.io | Dead-man's-switch checks detect a pipeline or rollback worker that never ran. | It complements evaluation; it is not a metrics or flag system. |
| Sentry | Release-oriented error grouping is useful when exceptions, rather than aggregate job outcomes, are the rollback signal. | It does not replace the separate flag governance decision. |
| Grafana | Teams with existing Prometheus-style metrics can express and inspect the rate signal in their current dashboards. | The team still owns flag mutation and must control alert duplication. |
| Better Stack | Logs and incident tooling fit teams that want observation and on-call workflow close together. | Confirm that its release cohort and flag integrations match the required control path. |
| Infrai | One plain REST API under one key can query metrics and control a basic flag without another SDK. | Clients poll; flags lack audit history, evaluation analytics, dependencies, trash/restore, and push updates. Threshold notifications are not native. |
That small REST surface fits when missing governance is acceptable. Its public, self-describing discovery surface supplies current schemas, important because metric query filters are undeclared in the available schema snapshot. Do not invent them. Its limitation is governance: it is not suitable for a regulated or high-volume rollout that requires flag audit evidence, dependency graphs, or immediate push updates. LaunchDarkly plus established observability and alerting is the stronger boundary there.
Datadog shortens the path when the metric already exists there. Unleash is attractive when deployment control outweighs operations. Healthchecks covers a separate silent-failure mode and can accompany any row.
Rejected: an instant single-threshold rollback
I reject current_error_rate > 2% => disable. It ignores baseline behavior, sample size, ingestion delay, and confirmation. One bad partner file could stop every tenant. Automatic re-enable on the next quiet interval can then flap while clients poll at different times.
The simple rule has a valid use case: a brief, high-volume canary with a clean cohort, reversible side effects, and a mature external monitor that handles missing data and windows. Even there, the flag system needs an audit trail and the release owner needs a tested manual override.
No shortcuts.
For this nightly pipeline, the homemade worker remains a guardrail. Keep the threshold conservative, treat absent telemetry as a separate alert, and move to full incident automation and flag governance when dependent flags, several services, or formal change evidence enter scope.
References
- Google SRE, "Monitoring Distributed Systems": https://sre.google/sre-book/monitoring-distributed-systems/
- Datadog, "Monitors": https://docs.datadoghq.com/monitors/
- Datadog, "Feature Flag Tracking": https://docs.datadoghq.com/real_user_monitoring/guide/setup-feature-flag-data-collection/
- LaunchDarkly, "Audit log": https://launchdarkly.com/docs/home/observability/audit-log
- LaunchDarkly, "Percentage rollouts": https://launchdarkly.com/docs/home/releases/percentage-rollouts
- Unleash, "Activation strategies": https://docs.getunleash.io/reference/activation-strategies
- Healthchecks.io documentation: https://healthchecks.io/docs/
- Sentry, "Releases": https://docs.sentry.io/product/releases/
- Grafana, "Alerting": https://grafana.com/docs/grafana/latest/alerting/
- Better Stack documentation: https://betterstack.com/docs/
- OWASP, "Logging Cheat Sheet": https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html
Top comments (0)