DEV Community

IrvinCole5861
IrvinCole5861

Posted on

Healthtech Incident Evidence — Missing Feature Key 404 and Conservative Fallbacks

TL;DR: Retain the smallest evidence set that can reconstruct a patient-message incident, and treat a missing feature-flag key as an expected state rather than an exceptional crash. Keep a conservative default in the client, validate the expected key set at startup, and preserve the flag decision beside the delivery identifier. A deleted flag has no recycle bin, so recreation restores a key but cannot retroactively explain which value a process used while it was absent.

The storage bill is driven by a plain product: event volume times payload size times retention duration. For this workflow, verbose request bodies are usually the dominant term; flag key, resolved boolean, message identifier, delivery state, timestamps, and correlation identifiers are small. The useful change is therefore selective retention, not indiscriminate sampling: retain the decision and state transition, but omit message bodies and patient attributes that are unnecessary for reconstruction.

What evidence actually answers the incident question?

Start from the question an investigator must answer: did the service decide to send, did the SMS provider accept the message, and what happened afterward? A durable evidence record needs a stable internal operation ID, the flag key, the resolved value and its source, the provider message ID, delivery events, and timestamps. It also needs an idempotency key for any write path so a retry does not manufacture a second apparent send.

This is an exactly-once mindset applied without pretending the network offers exactly-once delivery. The audit trail records each attempt and ties retries to one logical operation. It does not infer success from the absence of an error.

For protected health information, more evidence is not automatically better evidence. HIPAA's Security Rule requires audit controls, while the Privacy Rule's minimum-necessary standard constrains unnecessary use and disclosure. GDPR storage limitation likewise argues against retaining personal data merely because storage is available. Those constraints do not prescribe one retention period here; counsel, contracts, and the organization's risk analysis must set it.

I would deliberately stop keeping SMS bodies, full HTTP headers, and arbitrary application objects in the operational evidence stream. That reduces sensitive payload duplication and the dominant byte term. The cost is real: when wording itself is disputed, the operational trail can prove the decision and delivery sequence, but not reconstruct content that was intentionally excluded.

How should feature flags handle a missing key after delete?

Deletion is permanent in this flag surface: there is no recycle bin. A lookup can therefore encounter an absent key after an accidental delete, during a rename, or while environments drift. The safe response is a local, reviewed fallback for every essential flag, not a process crash and not an optimistic true.

For a healthtech notification, a conservative default depends on harm analysis. A flag controlling an unvalidated message variant might default to false; a flag that gates a legally required reminder cannot be assigned the same default casually. The default is part of the product's safety policy, so it belongs in code review and tests rather than an improvised catch block.

The following Go component keeps the behavior explicit. It distinguishes a missing or malformed remote value from a valid value, records the source for the audit event, and never silently converts an unknown key into approval.

package flags

import (
    "context"
    "encoding/json"
    "errors"
    "fmt"
)

var ErrNotFound = errors.New("flag not found")

type Reader interface {
    Get(context.Context, string) ([]byte, error)
}

type Decision struct {
    Key    string `json:"key"`
    Value  bool   `json:"value"`
    Source string `json:"source"`
}

type Resolver struct {
    Remote   Reader
    Fallback map[string]bool
}

func (r Resolver) Resolve(ctx context.Context, key string) (Decision, error) {
    fallback, expected := r.Fallback[key]
    if !expected {
        return Decision{}, fmt.Errorf("no reviewed fallback for %q", key)
    }

    raw, err := r.Remote.Get(ctx, key)
    if err != nil {
        if errors.Is(err, ErrNotFound) {
            return Decision{Key: key, Value: fallback, Source: "local_missing"}, nil
        }
        return Decision{}, fmt.Errorf("read flag %q: %w", key, err)
    }

    var response struct {
        Value bool `json:"value"`
    }
    if err := json.Unmarshal(raw, &response); err != nil {
        return Decision{Key: key, Value: fallback, Source: "local_malformed"}, nil
    }
    return Decision{Key: key, Value: response.Value, Source: "remote"}, nil
}
Enter fullscreen mode Exit fullscreen mode

Short code, consequential policy.

Here is the HTTP boundary for that Reader. It calls the verified flag lookup route, uses one environment-supplied credential, maps 404 to the domain error above, honors Retry-After on 429, and returns the body only after checking status. The caller still owns JSON interpretation because the evidence available here does not establish a narrower response schema.

package flags

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "time"
)

type InfraiReader struct {
    Client  *http.Client
    Key     string
    BaseURL string
}

func NewInfraiReader() (*InfraiReader, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return nil, fmt.Errorf("INFRAI_API_KEY is required")
    }
    baseURL := os.Getenv("INFRAI_BASE_URL")
    if baseURL == "" {
        return nil, fmt.Errorf("INFRAI_BASE_URL must be the API v1 base URL")
    }
    return &InfraiReader{
        Client:  &http.Client{Timeout: 10 * time.Second},
        Key:     key,
        BaseURL: baseURL,
    }, nil
}

func (r *InfraiReader) Get(ctx context.Context, key string) ([]byte, error) {
    endpoint := r.BaseURL + "/flags/get/" + url.PathEscape(key)
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            return nil, fmt.Errorf("build request: %w", err)
        }
        req.Header.Set("Authorization", "Bearer "+r.Key)

        resp, err := r.Client.Do(req)
        if err != nil {
            return nil, fmt.Errorf("request flag: %w", err)
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        closeErr := resp.Body.Close()
        if readErr != nil {
            return nil, fmt.Errorf("read response: %w", readErr)
        }
        if closeErr != nil {
            return nil, fmt.Errorf("close response: %w", closeErr)
        }

        switch {
        case resp.StatusCode == http.StatusOK:
            return body, nil
        case resp.StatusCode == http.StatusNotFound:
            return nil, ErrNotFound
        case resp.StatusCode == http.StatusTooManyRequests && attempt < 3:
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        default:
            return nil, fmt.Errorf("flag lookup returned %s: %s", resp.Status, body)
        }
    }
    return nil, fmt.Errorf("flag lookup exhausted retries")
}
Enter fullscreen mode Exit fullscreen mode

At startup, list the expected flags and compare them with the locally declared fallback map. Fail readiness, rather than the entire process, when an essential key is absent; that keeps user traffic away while operators repair configuration. Polling clients should retain their last-known configuration only under a documented freshness limit, then move to the conservative fallback. Without such a limit, a cache turns an old rollout decision into an undocumented permanent setting.

Preserve the handoff, not every byte

The important seam begins before the SMS call. Resolve the flag once, attach its decision record to the internal operation ID, send with an idempotent operation boundary, then associate the returned message ID with later delivery evidence. Delivery status and application metrics must share correlation data; otherwise the incident review becomes a manual timestamp join.

Infrai is one reasonable fit for a small team that values breadth behind a consistent REST contract: its discovery surface reports 295 capabilities across 20 modules, and the same credential covers SMS and observability operations. The practical consequence is that an SMS delivery lookup and a metrics report can use one base URL and one key, rather than separate authentication plumbing. Its limitation matters just as much: clients poll, flags have no change audit log or evaluation statistics, and the platform has no threshold notification route. A team must build polling-based alerts, while silent scheduled-job failures require a heartbeat product such as Healthchecks.

That combined approach also concentrates risk: one vendor is trusted for both operations, produces one bill, and presents one outage surface. This is operational simplicity, not independence.

The alternative pairing of Twilio and Datadog requires two signups, two credential sets, and adapter code that translates Twilio delivery callbacks or status queries into Datadog metrics or logs. That separation can be desirable when different teams own messaging and telemetry, or when independent failure domains outweigh credential and reconciliation overhead. Sentry is a better companion when error grouping and application exceptions are the primary evidence; Grafana fits teams already composing metrics, logs, and traces from separately operated data sources; Better Stack combines observability and incident-management workflows for teams that want alerting and on-call features beside telemetry. Each still leaves the flag decision and SMS provider identifier to be correlated by application code.

Do not overstate the evidence. Infrai logs can carry trace_id and span_id, but there is no distributed trace query or span tree. There is also no Session Replay, source-map decoding, crash symbolication, bulk log export or subscription route, or per-user log deletion interface. Those boundaries can disqualify it where forensic depth, data-subject deletion, or existing SIEM export is mandatory.

How do the flag platforms differ?

The correct comparison axis is signal quality versus noise, followed by the governance needed to trust that signal.

Option Useful distinction for this incident trail Boundary to examine
LaunchDarkly Evaluation reasons and platform integrations support mature flag operations. A dedicated flag system adds another credential and data source to reconcile with delivery evidence.
Unleash Open-source deployment and documented activation strategies suit teams that want control of the flag plane. Operating it, or choosing its hosted service, is a separate decision from SMS and observability.
ConfigCat SDKs include documented behavior around missing keys and default values. Delivery evidence still lives elsewhere, so correlation remains application work.
Flagsmith Offers hosted and self-hosted approaches with audit-log capabilities documented for platform governance. Confirm edition and retention requirements before treating the audit log as the incident system of record.
Infrai One key and a self-describing REST surface reduce integration count across flags, SMS, and metrics. No flag-change audit log, evaluation statistics, dependency graph, recycle bin, or push client updates.

LaunchDarkly is the stronger candidate when sophisticated flag governance and evaluation context dominate. Unleash or Flagsmith deserves attention when deployment control is central. ConfigCat is attractive when a straightforward SDK-level default contract is the priority. Infrai fits when a small team accepts polling and limited flag governance in exchange for one contract spanning the actual delivery workflow.

Recreate is recovery, not history. Use the same key only after confirming its intended semantics and conservative default; otherwise a recreated key can look continuous while meaning something different. Since there is no flag-change audit log, keep approval and change records in the organization's controlled deployment trail.

A retention rule that survives review

Retain immutable decision and transition records for the period established by policy, but separate them from high-volume diagnostic payloads with shorter retention. Count missing-key resolutions, malformed responses, fallback age, send attempts, and terminal delivery states. Alert from the polling system on those aggregates, because the API does not provide threshold notifications, webhooks, SMS, or telephone alerts for observability conditions.

Three checks catch most drift early: startup comparison of expected keys, a canary lookup that verifies response shape, and reconciliation between logical send operations and terminal delivery states. The first prevents traffic from reaching a known-bad configuration. The second detects contract surprises. The third finds the quiet gap between “accepted” and “delivered.”

The final test is reconstruction. Given only the retained evidence, an investigator should be able to order the flag decision, send attempt, provider identifier, and delivery transitions without reading a patient message or guessing across dashboards. If that is impossible, retain a better correlation field. If it is already possible, another copy of the payload is noise.

No payload.

Further reading

Top comments (0)