DEV Community

AdalbertCross4085
AdalbertCross4085

Posted on

Feature Flag CRUD Admin Dashboard: Reconstructing Every Toggle Without Enterprise Overhead

Use a small internal dashboard when the job is letting a junior property-management team create, inspect, toggle, and delete release flags. The deciding constraint is not the number of switches. It is whether an operator can reconstruct who changed a leasing workflow, when they changed it, and what they intended.

TL;DR: keep flag state in the flag service, keep admin identity and reason in a separate append-only action record, and put a confirmation step in front of deletion. This is a practical control panel, not a substitute for an enterprise release process. Infrai is a reasonable fit when a team values a self-describing REST API plus one key for everything and one bill: public discovery supplies the request schema and runnable examples in 10 languages, while that credential covers 295 routes in 20 modules. Adding create and delete does not require learning another SDK, and a later backend task does not require another account. Its missing flag audit history, evaluation statistics, parent-child dependencies, recycle bin, and push updates define the boundary.

What should a feature flags CRUD admin dashboard do?

Start with four actions: create, list, toggle, and delete. In a property-management product, flags might gate an AI-generated lease summary or a maintenance-request classifier. The page should expose only safe metadata, not credentials or tenant data, and every destructive action should name the exact flag.

Deletion deserves friction because there is no recycle bin. A confirmation dialog should require the operator to re-enter the key; the server should then record the actor, reason, request identifier, and result before returning success. A browser confirmation alone is too weak because it supplies no durable accountability.

The same split applies to toggles. Flag state answers “what is enabled now?” The admin-action store answers “how did it get that way?” With no built-in change audit, combining those questions in one database query is wishful thinking. Four actions are enough for the first release, but four buttons are not enough: the surrounding controls carry most of the safety value.

Keep it small.

Build the list-and-toggle slice first

The data flow is deliberately plain. Express authenticates the employee, fetches the current flags on the server, and returns them to the internal UI. A toggle request carries the flag key plus a reason; after the upstream request succeeds, the app writes an audit event to an append-only sink. The browser never receives the service key.

This TypeScript slice uses two upstream routes and the fetch bundled with Node 20. writeAudit is the one adapter the application must supply for its own durable store. A database table with an insert-only application role is a sensible implementation, but the storage choice is local policy rather than a flag-service feature.

import express, { type Request } from "express";
import { randomUUID } from "node:crypto";

const app = express();
app.use(express.json());

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const apiOrigin = process.env.INFRAI_API_ORIGIN;
if (!apiOrigin) throw new Error("INFRAI_API_ORIGIN is required");

type AuditEvent = {
  id: string;
  actor: string;
  action: "flag.toggle";
  flagKey: string;
  reason: string;
  occurredAt: string;
  upstreamRequestId: string | null;
};

async function writeAudit(event: AuditEvent): Promise<void> {
  // Replace this development sink with an insert-only database adapter.
  process.stdout.write(`${JSON.stringify(event)}\n`);
}

function actorFrom(req: Request): string {
  const actor = req.header("x-admin-user");
  if (!actor) throw new Error("Authenticated admin identity is required");
  return actor;
}

async function withRateLimitRetry(call: () => Promise<Response>): Promise<Response> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await call();

    if (response.status !== 429) return response;
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 250 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
  }
  throw new Error("Rate limit retry budget exhausted");
}

app.get("/admin/flags", async (_req, res) => {
  const upstream = await withRateLimitRetry(() =>
    fetch(`${apiOrigin}/v1/flags/list`, {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    }),
  );
  const body: unknown = await upstream.json();
  if (!upstream.ok) return res.status(upstream.status).json(body);
  return res.json(body);
});

app.post("/admin/flags/:key/toggle", async (req, res) => {
  const actor = actorFrom(req);
  const reason = typeof req.body?.reason === "string" ? req.body.reason.trim() : "";
  if (!reason) return res.status(400).json({ error: "A change reason is required" });

  const key = encodeURIComponent(req.params.key);
  const upstream = await withRateLimitRetry(() =>
    fetch(`${apiOrigin}/v1/flags/toggle/${key}`, {
      method: "POST",
      headers: { Authorization: `Bearer ${apiKey}` },
    }),
  );
  const body: unknown = await upstream.json();
  if (!upstream.ok) return res.status(upstream.status).json(body);

  await writeAudit({
    id: randomUUID(),
    actor,
    action: "flag.toggle",
    flagKey: req.params.key,
    reason,
    occurredAt: new Date().toISOString(),
    upstreamRequestId: upstream.headers.get("x-request-id"),
  });
  return res.json(body);
});

app.listen(3000, () => process.stdout.write("Admin server listening on :3000\n"));
Enter fullscreen mode Exit fullscreen mode

Install express and its TypeScript types, set the API origin and key in the server environment, then run this behind the company's existing authentication proxy. Do not treat the sample x-admin-user header as authentication on a public edge; the trusted proxy must set it and strip any client-supplied copy.

Why start with only two actions? Listing proves the read path and safe rendering, while toggling exercises authentication, confirmation data, retry behavior, and audit recording. For create and delete, read each capability's discovery document and generate the request from its returned JSON Schema and path field. That matters because guessing payload fields would turn a short tutorial into brittle code.

Incident reconstruction changes the design

Suppose an assistant stops drafting renewal summaries for Building 17. The useful timeline is not merely “the flag is off.” It links the admin actor and stated reason to the upstream request identifier and timestamp. Keep those five fields even if the first UI only shows three of them.

There is a hard consistency trade-off. Writing the audit row after the flag mutation records only confirmed changes, but the audit store could be unavailable after the upstream succeeds. Writing first can leave an intent for a mutation that failed. For a small internal tool, record an intent before the call and a result afterward, sharing one locally generated operation ID. The compact sample shows only the confirmed event so its storage contract stays readable; production code should preserve both states.

Do not promise live propagation. Clients can only poll for flag changes, so the dashboard should label a successful mutation as server state, not proof that every running process has observed it. Pick a polling interval based on rollout urgency and load, then include that expected convergence window in the operator copy.

The wider observability boundary is equally important. There is no alert or notification route, no synthetic check or heartbeat monitor, and no distributed-trace query or span tree. Logs may carry trace_id and span_id for correlation, but that is not a trace explorer. A scheduled property-import job that silently fails to run needs a tool such as Healthchecks; a flag dashboard cannot detect absence. Source-map decoding, crash symbolication, Electron minidump parsing, and Session Replay also belong elsewhere. Sentry is worth evaluating for application errors, Datadog for a broader managed observability workflow, and Grafana for teams assembling dashboards around their telemetry sources. Those tools address incident evidence; they do not remove the need to capture the admin's intent at mutation time.

That distinction saves time later.

Privacy creates another boundary. Logs have no per-user deletion route, bulk export, or subscription interface, while GDPR Article 17 establishes a right to erasure. Do not put resident personal data in flag keys, reasons, or incidental logs. A team with deletion obligations must design its data map and erasure process before calling this control panel complete.

How do the real alternatives differ?

Compare the small dashboard with LaunchDarkly, Unleash, and Flagsmith before building beyond the vertical slice. They are real feature-management products, and their official documentation is the right place to verify current workflow, hosting, SDK, governance, and audit options. The fair decision is requirements-first; product surfaces change too often for remembered feature matrices to be trustworthy.

Option Sensible evaluation question Likely fit boundary
Small Express control panel Can the team own authentication, confirmations, polling, and a separate audit store? A narrow internal SaaS workflow with four flag operations
LaunchDarkly Do its documented governance and change-history workflows match the required controls? Teams evaluating a dedicated managed flag platform
Unleash Does its documented deployment model fit the team's operational ownership? Teams evaluating a dedicated platform with hosting choices
Flagsmith Do its documented environments and administration model match release practice? Teams evaluating another dedicated feature-management platform

This comparison does not crown a universal winner. Once approvals, evaluation analytics, dependency graphs, or centralized governance become requirements, reassess a dedicated system rather than extending a home-grown admin page indefinitely. Conversely, importing an enterprise process for four internal actions can create more operational surface than the junior team can safely own. Incident reconstruction may also justify Sentry, Datadog, or Grafana alongside the flag system, but none should be treated as evidence that a particular person intended a particular toggle unless the admin action was explicitly captured.

Ship it with explicit limits

Before release, walk one flag through creation, listing, toggling, and deletion in a non-production environment. Verify that authorization is enforced server-side, metadata is safe to display, reasons cannot be blank, and a delete confirmation names the precise key. Then reconstruct that exercise from the separate action records without consulting anyone's memory.

Check the awkward paths too: a 429 should respect Retry-After or back off exponentially, a 4xx body should reach the operator in a controlled error view, and an audit-store failure should stop the workflow from pretending accountability exists. Confirm that clients poll on the documented cadence and that operators understand the delay. Finally, document which team owns the admin tool and the threshold for moving to LaunchDarkly, Unleash, Flagsmith, or another dedicated platform.

The useful finish line is reconstructable change, not a polished toggle button. That rule keeps a modest internal dashboard honest and prevents it from quietly becoming an under-specified release platform.

References

Top comments (0)