The central trade-off is signal quality versus noise: a feature flag should stop an unwanted notification path without making a real delivery failure disappear. TL;DR: keep flag administration separate from notification delivery, record every mutation, evaluate a small immutable snapshot in the send path, and classify suppressed_by_flag separately from delivery_failed. Operators then get four plain actions—list, set, toggle, and delete—while alerts continue to represent failures that require intervention.
This matters in property management because one service may send rent reminders, maintenance updates, lease notices, and one-time passcodes. My rule from email, SMS, and OTP work is to keep the tool's 4 actions legible and the delivery outcomes more precise. A burst of expected suppressions can bury a smaller set of rejected or expired messages if both land in the same counter. Silence is not success.
How Should an Admin Dashboard Handle Feature Flag CRUD?
Start with the invariant, not the dashboard. A flag decides whether the service may attempt a class of notification. It does not decide whether an attempted message was delivered, and it must not rewrite the outcome after the fact.
Use narrow keys such as maintenance_sms_enabled and rent_reminder_email_enabled. Avoid a single notifications_enabled switch: its blast radius is difficult to reason about, and an operator cannot tell which obligation or tenant workflow a change affects. The flag value should be boolean, while its record carries operational context: revision, update time, actor, and reason.
The resulting state machine is small:
| Condition | Delivery action | Observable outcome | Alert treatment |
|---|---|---|---|
| Flag enabled, provider accepts | Attempt send | accepted |
No failure alert |
| Flag enabled, provider rejects | Attempt send | delivery_failed |
Count as failure |
| Flag disabled | Do not send | suppressed_by_flag |
Track, but do not count as delivery failure |
| Flag missing | Follow declared default |
default_applied plus final outcome |
Alert on unexpected key |
That third row is the boundary that keeps the telemetry honest. Suppression volume is still useful. A sudden increase can identify a broad operational change, but it belongs on a change dashboard rather than in the delivery-failure numerator.
Step 1: Build a Small, Audited Control Store
For a simple internal tool, one transactional database table is enough. The example below uses Python's standard SQLite interface so the behavior is visible without a framework or commercial dependency. The same constraints belong in whichever transactional store the service already operates.
Deletion needs special care. Removing a flag should restore its documented default; it should not erase the mutation trail. The audit row stores the flag key, action, actor, reason, time, and resulting revision. Do not place tenant names, phone numbers, email addresses, message bodies, or OTP values in this table. Those fields add no value to a flag decision and complicate erasure handling.
import sqlite3
from datetime import datetime, timezone
SCHEMA = """
CREATE TABLE IF NOT EXISTS feature_flags (
key TEXT PRIMARY KEY,
enabled INTEGER NOT NULL CHECK (enabled IN (0, 1)),
revision INTEGER NOT NULL,
updated_at TEXT NOT NULL,
updated_by TEXT NOT NULL,
reason TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS flag_audit (
id INTEGER PRIMARY KEY AUTOINCREMENT,
flag_key TEXT NOT NULL,
action TEXT NOT NULL CHECK (action IN ('set', 'toggle', 'delete')),
enabled INTEGER,
revision INTEGER NOT NULL,
changed_at TEXT NOT NULL,
changed_by TEXT NOT NULL,
reason TEXT NOT NULL
);
"""
class FlagStore:
def __init__(self, connection):
self.db = connection
self.db.row_factory = sqlite3.Row
self.db.executescript(SCHEMA)
def list_flags(self):
rows = self.db.execute(
"SELECT key, enabled, revision, updated_at, updated_by, reason "
"FROM feature_flags ORDER BY key"
).fetchall()
return [dict(row) for row in rows]
def set_flag(self, key, enabled, actor, reason):
self._validate(key, actor, reason)
now = datetime.now(timezone.utc).isoformat()
with self.db:
current = self.db.execute(
"SELECT revision FROM feature_flags WHERE key = ?", (key,)
).fetchone()
revision = 1 if current is None else current["revision"] + 1
self.db.execute(
"""INSERT INTO feature_flags
(key, enabled, revision, updated_at, updated_by, reason)
VALUES (?, ?, ?, ?, ?, ?)
ON CONFLICT(key) DO UPDATE SET
enabled = excluded.enabled,
revision = excluded.revision,
updated_at = excluded.updated_at,
updated_by = excluded.updated_by,
reason = excluded.reason""",
(key, int(enabled), revision, now, actor, reason),
)
self._audit(key, "set", enabled, revision, now, actor, reason)
return revision
def toggle(self, key, expected_revision, actor, reason):
self._validate(key, actor, reason)
now = datetime.now(timezone.utc).isoformat()
with self.db:
current = self.db.execute(
"SELECT enabled, revision FROM feature_flags WHERE key = ?", (key,)
).fetchone()
if current is None:
raise KeyError(key)
if current["revision"] != expected_revision:
raise ValueError("flag changed; refresh before toggling")
enabled = not bool(current["enabled"])
revision = current["revision"] + 1
changed = self.db.execute(
"""UPDATE feature_flags
SET enabled = ?, revision = ?, updated_at = ?,
updated_by = ?, reason = ?
WHERE key = ? AND revision = ?""",
(int(enabled), revision, now, actor, reason, key, expected_revision),
)
if changed.rowcount != 1:
raise ValueError("concurrent flag change")
self._audit(key, "toggle", enabled, revision, now, actor, reason)
return enabled, revision
def delete(self, key, expected_revision, actor, reason):
self._validate(key, actor, reason)
now = datetime.now(timezone.utc).isoformat()
with self.db:
deleted = self.db.execute(
"DELETE FROM feature_flags WHERE key = ? AND revision = ?",
(key, expected_revision),
)
if deleted.rowcount != 1:
raise ValueError("missing flag or stale revision")
self._audit(key, "delete", None, expected_revision + 1,
now, actor, reason)
def _audit(self, key, action, enabled, revision, now, actor, reason):
self.db.execute(
"""INSERT INTO flag_audit
(flag_key, action, enabled, revision, changed_at, changed_by, reason)
VALUES (?, ?, ?, ?, ?, ?, ?)""",
(key, action, None if enabled is None else int(enabled),
revision, now, actor, reason),
)
@staticmethod
def _validate(key, actor, reason):
if not key or not actor or not reason.strip():
raise ValueError("key, actor, and reason are required")
if __name__ == "__main__":
store = FlagStore(sqlite3.connect(":memory:"))
store.set_flag("maintenance_sms_enabled", True,
"ops@example.invalid", "initial controlled rollout")
first = store.list_flags()[0]
store.toggle(first["key"], first["revision"],
"ops@example.invalid", "pause during template review")
print(store.list_flags())
The expected revision is deliberate. Two open browser tabs must not silently overwrite each other. A stale toggle should return a conflict to the UI, which can reload the current record and ask the operator to decide again. Requiring a reason also creates useful context without pretending that an audit log is authorization.
Authentication and authorization sit in front of these methods. Give read access more broadly than mutation access, use the authenticated principal as actor, and reject actor values supplied by the browser. The dashboard is a control surface, not the source of identity.
Step 2: Publish a Snapshot to the Send Path
The notification worker should not query the control table for every message. Publish a validated snapshot after a successful mutation, then let workers replace their local snapshot atomically. This keeps the delivery path predictable and prevents a slow dashboard dependency from becoming a delivery outage.
Defaults are policy. Define them in code, test them, and expose when a default was used. For a new or misspelled key, quiet fallback is dangerous because the service looks healthy while behaving differently from the operator's intent.
from dataclasses import dataclass
from types import MappingProxyType
DEFAULTS = MappingProxyType({
"maintenance_sms_enabled": False,
"rent_reminder_email_enabled": False,
"lease_notice_email_enabled": False,
"resident_otp_sms_enabled": True,
})
@dataclass(frozen=True)
class Decision:
enabled: bool
source: str
revision: int | None
def evaluate(flag_key, snapshot):
record = snapshot.get(flag_key)
if record is not None:
return Decision(bool(record["enabled"]), "snapshot", record["revision"])
if flag_key not in DEFAULTS:
raise KeyError(f"unknown feature flag: {flag_key}")
return Decision(DEFAULTS[flag_key], "default", None)
def plan_delivery(notification, snapshot, emit):
decision = evaluate(notification["flag_key"], snapshot)
emit("flag_evaluated", {
"flag_key": notification["flag_key"],
"enabled": decision.enabled,
"source": decision.source,
"revision": decision.revision,
"channel": notification["channel"],
"message_class": notification["message_class"],
})
if not decision.enabled:
emit("notification_outcome", {
"outcome": "suppressed_by_flag",
"channel": notification["channel"],
"message_class": notification["message_class"],
})
return None
return notification
Notice what is absent from the event: recipient address, unit number, message body, and OTP. Aggregate labels should describe the channel and message class, not the resident. This reduces sensitive data exposure and prevents high-cardinality identifiers from turning an operational chart into a costly search index.
The resident_otp_sms_enabled default above is intentionally different from the campaign-like flows. Losing access is a different risk from delaying a reminder. This is a policy example, not a universal answer; the organization responsible for the workflow must set the default and document the consequence.
Step 3: Measure Attempts, Suppressions, and Failures Separately
One ratio does most of the work:
delivery failure rate = delivery_failed / delivery_attempted
Do not include suppressed_by_flag in either side. Also do not report a suppression as an accepted delivery. Keep a second count by flag key and revision so an operator can align a traffic change with the exact control-plane mutation.
I use three diagnostic questions when a chart moves: Did the eligible volume change? Did the attempt volume change? Did the failure rate among attempts change? That sequence catches a common trap. If a flag cuts attempts in half while the remaining provider failures stay constant, an all-events denominator makes the failure rate look better even though delivery quality did not improve.
This is where the deliverability details earn their place. Separate permanent rejection from retryable failure, and keep rate-limit responses out of a generic “unknown” bucket. A retry can later succeed; a malformed destination will not become valid through persistence. The dashboard need not expose every transport detail, but the underlying outcome taxonomy must preserve enough information for a worker to apply the correct retry policy.
Keep alerts focused. Alert on failure rate and on inability to refresh the snapshot. Show suppressions and flag changes as annotations or companion panels. A high suppression count after an approved pause is expected; a high failure rate after an enable operation deserves attention.
Step 4: Test and Roll Out the Boundary
Test the invariant before polishing the dashboard. One test should prove that a disabled flag emits suppressed_by_flag and never calls the transport. Another should prove that an enabled flag passes the notification onward. Add tests for an unknown key, a declared default, stale revisions, deletion, and audit persistence.
def test_disabled_notification_is_suppressed():
events = []
snapshot = {
"maintenance_sms_enabled": {"enabled": False, "revision": 7}
}
notification = {
"flag_key": "maintenance_sms_enabled",
"channel": "sms",
"message_class": "maintenance_update",
}
result = plan_delivery(notification, snapshot,
lambda name, data: events.append((name, data)))
assert result is None
assert events[-1][1]["outcome"] == "suppressed_by_flag"
assert all(event[1].get("outcome") != "delivery_failed" for event in events)
def test_stale_toggle_is_rejected():
connection = sqlite3.connect(":memory:")
store = FlagStore(connection)
store.set_flag("maintenance_sms_enabled", True,
"ops@example.invalid", "begin rollout")
try:
store.toggle("maintenance_sms_enabled", 0,
"ops@example.invalid", "stale browser action")
except ValueError as error:
assert "revision" in str(error) or "changed" in str(error)
else:
raise AssertionError("stale toggle was accepted")
Then deploy in observation mode: evaluate flags and emit decisions while leaving existing send behavior unchanged. Compare eligible, attempted, and failed counts. Once the classifications agree with the existing pipeline, enable enforcement for one low-risk message class, watch one full operational cycle, and expand by message class rather than by every property at once.
Keep rollback boring. The previous snapshot should remain available, mutations should be traceable by revision, and restoring it should create a new audit event instead of deleting history. For records containing personal data, define retention and erasure procedures with counsel; GDPR Article 17 establishes a right to erasure and also lists circumstances in which it does not apply. The safer architecture still avoids collecting personal data in flag records at all.
The finished internal tool is intentionally modest. Its value comes from preserving meaning: a disabled workflow is visible as a policy decision, while a failed delivery remains a failure. That separation gives property operations a quiet dashboard without buying quiet at the cost of missing resident-impacting problems.
Sources
References:
- GDPR Article 17, “Right to erasure”: https://gdpr-info.eu/art-17-gdpr/
Top comments (0)