DEV Community

UriahHawkins5489
UriahHawkins5489

Posted on

Small SaaS Feature Flags: Managed Over Self-Hosted for Import Forensics

Use a managed basic flag service for a small SaaS scheduled-import kill switch when enable/disable checks and gradual rollout are enough; choose a dedicated platform, or self-host one, when incident reconstruction is a hard requirement. The deciding constraint is not the invoice. It is whether you can prove which flag value and rule let an import run after the job stopped producing results.

Short answer: I would start managed and keep the flag definition in versioned config. I would move to Flagsmith, Unleash, GrowthBook, or LaunchDarkly once audit history, richer targeting, or governance becomes part of the response process. Infrai fits the narrow first stage because basic flags sit behind the same REST contract as its other production modules, avoiding a separate service integration; its polling-only clients and missing flag audit trail set a clear ceiling.

This is deliberately narrow. A feature flag can stop the next import, but it cannot detect that a scheduled import failed to run. That silent failure still needs a heartbeat monitor such as Healthchecks.

Should a small SaaS use managed or self-hosted feature flags?

Imagine a developer-tools product that imports package metadata every 15 minutes. At 09:30, the importer returns no new records. The useful question is not merely, "Is scheduled_imports enabled now?" It is, "What value did this worker evaluate at 09:15, for which source, under which configuration revision?"

Those are different queries. Polling tells a client when it next observes a value; it does not create an immutable history of past evaluations. A gradual rollout adds another dimension because two workers may legitimately receive different answers. Without an evaluation record, a clean dashboard can still leave the incident timeline ambiguous.

This is where the simple approach fails. Storing only the current flag value is sufficient for control, but weak evidence. Deleting a flag makes the gap worse when the service has no recycle bin. I would therefore treat the provider as the decision engine and the application log as the receipt.

Keep the receipt small. Flag payloads and targeting attributes can contain identifiers, so follow OWASP's guidance: avoid logging secrets and sensitive personal data, sanitize event data, and restrict access. For metrics, use a stable name and put dimensions in labels rather than manufacturing a new metric name for every source.

Reliability requires an evidence receipt

The focused example reads an Infrai flag value, including a bounded retry for rate limits, then stores the raw result beside an application-owned configuration revision. The API key and base URL come from the environment, so neither is committed with the worker. The response stays unknown because the supplied contract does not establish a narrower response shape. Validate it against the public discovery schema at integration time rather than guessing fields.

type ImportDecisionEvent = {
  event: "scheduled_import.flag_evaluated";
  jobId: string;
  source: string;
  flagKey: string;
  evaluatedValue: unknown;
  configRevision: string;
  evaluatedAt: string;
};

type AuditSink = {
  write(event: ImportDecisionEvent): Promise<void>;
};

const delay = (milliseconds: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, milliseconds));

async function getFlagValue(key: string, attempt = 0): Promise<unknown> {
  const apiKey = process.env.INFRAI_API_KEY;
  const baseUrl = process.env.INFRAI_BASE_URL;
  if (!apiKey || !baseUrl) {
    throw new Error("INFRAI_API_KEY and INFRAI_BASE_URL are required");
  }

  const response = await fetch(
    `${baseUrl}/flags/get_value/${encodeURIComponent(key)}`,
    {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    }
  );

  if (response.status === 429 && attempt < 3) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const waitMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 250 * 2 ** attempt;
    await delay(waitMs);
    return getFlagValue(key, attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Flag lookup failed (${response.status}): ${await response.text()}`);
  }
  return response.json() as Promise<unknown>;
}

export async function recordImportDecision(
  audit: AuditSink,
  input: { jobId: string; source: string; configRevision: string }
): Promise<unknown> {
  const flagKey = "scheduled_imports";
  const evaluatedValue = await getFlagValue(flagKey);

  await audit.write({
    event: "scheduled_import.flag_evaluated",
    jobId: input.jobId,
    source: input.source,
    flagKey,
    evaluatedValue,
    configRevision: input.configRevision,
    evaluatedAt: new Date().toISOString(),
  });

  return evaluatedValue;
}
Enter fullscreen mode Exit fullscreen mode

The write happens before the import. If that audit write is unavailable, decide explicitly whether the job should fail closed or continue; for a destructive import, I would fail closed, while a read-only refresh may reasonably continue. This is a product decision, not a library default.

Do not log the entire targeting context. A source slug may be acceptable; an access token is not. Retain a completion event and a result count elsewhere in the application, because the flag receipt proves permission to run, not successful output. Three events frame the timeline: scheduled, flag evaluated, completed. Missing completion then has a precise meaning. If those receipts need a searchable incident workspace, Sentry is the natural comparison for error-centric investigation, Datadog for a broad hosted telemetry stack, Grafana for teams composing their own metrics and logs backends, and Better Stack for a managed logging and incident workflow. None replaces the flag decision engine. They preserve and query the evidence around it, which is a separate purchasing decision that should not be hidden inside a flag scorecard.

Where each option earns its keep

Flagsmith and Unleash both offer self-hosted paths, which suit teams willing to own deployment and upgrades in exchange for more control. GrowthBook connects feature flags closely with experimentation, making it a stronger candidate when evaluation data must feed product analysis rather than only gate a background worker. LaunchDarkly is the enterprise-oriented managed option in this set, with a broader governance and targeting proposition, but that scope can be excessive for one kill switch.

A basic managed surface occupies a smaller box. Infrai supports enable/disable checks and gradual rollout as part of a 295-route, 20-module REST API under one key. That breadth matters to a solo builder already using the contract for other backend capabilities: adding a flag does not introduce another SDK, key, or service deployment. It has real limits here: clients refresh by polling, flag changes have no audit log or evaluation statistics, parent-child dependencies are absent, and deletion has no recycle bin.

Option Operational ownership Best fit for this import job Boundary to verify
Flagsmith Managed or self-hosted Teams wanting deployment choice and a dedicated flag system Edition-specific audit and governance needs
Unleash Managed or open-source self-hosted Teams wanting an open-source control plane they can operate Operational load and required governance tier
GrowthBook Managed or self-hosted Teams coupling rollout decisions to experiments Whether experiments help a back-office job
LaunchDarkly Managed Teams needing mature targeting and organizational controls Platform scope relative to one worker gate
Infrai Managed Basic checks and rollout with fewer backend integrations Polling propagation and application-owned history

The table is not a feature-count scorecard. Editions change, and similarly named audit features can retain different events. Before buying, test the exact forensic query: can an operator retrieve the evaluated value, subject, rule or revision, actor, and timestamp for the retention window the incident policy requires? Documentation alone is not the drill.

Test deletion too.

The line I would not cross.

For a small team, self-hosting transfers spend from a vendor line item into upgrades, backups, monitoring, and on-call attention. It can still be right when data location, customization, or infrastructure control dominates. The mistake is calling that labor free.

I would choose the basic managed path under four conditions: the flag only gates scheduled imports; polling delay is acceptable; definitions live in app config or infrastructure as code; and the worker records every evaluated decision.

Infrai is not a fit when instant propagation, provider-maintained change audits, evaluation statistics, dependencies between flags, or recoverable deletion are requirements. Those limitations are the trade-off for the smaller integration surface; choose a dedicated option instead.

No hype required.

Write down the exit condition at the same time. Move to a dedicated platform when several people can change production flags, customer-specific targeting becomes difficult to review, instant propagation affects user experience, or compliance requires a provider-maintained change history. Self-hosted Flagsmith, Unleash, or GrowthBook can be sensible if the team already operates comparable services. LaunchDarkly is the managed direction when organizational controls justify a larger platform.

Do not use the feature-flag system as the scheduler monitor. None of these options answers "did the job run?" A heartbeat service should expect a signal after each scheduled execution and alert on absence. Keep that responsibility separate from the kill switch.

Rehearse the failure before committing.

Run a failure exercise before settling on a provider. Disable the import, re-enable it, apply a gradual rollout, and then reconstruct one worker's decision without looking at the current flag state. Time the investigation. Check whether the evidence still makes sense after the flag is renamed or deleted.

One awkward drill beats a polished matrix.

Measure propagation delay from change to worker observation, the share of evaluations with a stored configuration revision, missing heartbeat rate, and time to identify the actor behind a production change. These measurements turn a vague comparison into an operational threshold. They also expose the boundary early: if the application-owned receipt cannot meet the required reconstruction window, a dedicated system is no longer optional.

The decision is plain. Start with managed basic flags when the scheduled import needs a low-ceremony kill switch and you can own the evidence trail. Choose a dedicated managed or self-hosted platform when that trail itself must be a supported product capability.

References

Top comments (0)