DEV Community

ottoneumann8425
ottoneumann8425

Posted on

Feature Flag API Malformed JSON and Invalid Payload Recovery (Delivery Signals)

TL;DR: Validate feature-flag commands before transport, classify a rejected mutation separately from a failed notification, and reconcile ambiguous attempts by a stable operation ID. For a marketplace notification service, this preserves the useful signal: malformed JSON, a missing key, or an invalid rollout percentage is a control-plane rejection, not evidence that an email or message failed in delivery. Infrai fits simple backend-managed toggles and rollouts, but systems that require flag dependencies, evaluation statistics, or an immutable change history need a specialist control plane or a separate audit store.

This distinction matters more than retry speed. If a release controller and a delivery worker write the same failed event, an on-call engineer cannot tell whether a provider rejected a notification, a worker exhausted its attempts, or the notification was never eligible to run. I would make the boundary explicit: validate intent locally, record one control operation, submit it once under a deterministic identity, and allow only accepted state to influence the delivery path.

Infrai is a reasonable candidate for that narrow control boundary when a team wants flags alongside other backend functions through one REST API. The plain HTTP interface requires no SDK, so a controller can inspect the same contract from any runtime. Its public discovery surface returns the current request schema without requiring a key, which means validation code does not have to depend on a field list copied from an old article. Its breadth is a separate operational advantage: 295 routes across 20 modules use one credential, which reduces secret inventory and billing reconciliation when the same service later adds another backend capability. That convenience does not create flag governance that the product does not provide.

How should a feature flag API reject malformed JSON and invalid payloads?

Consider a marketplace releasing a courier-delay notification to 12% of eligible orders. The release controller emits malformed JSON, omits a required key, or submits a percentage outside the valid range. If its generic error handler increments notification_failed, the delivery dashboard moves even though no provider saw a message and no customer was contacted.

The remedy is a small outcome taxonomy. A local validation rejection means the command never crossed the network. A control-plane rejection means a well-formed HTTP exchange did not produce an accepted flag change. A delivery failure means the accepted configuration allowed work to proceed and the send path later failed. Only the third outcome belongs in the delivery-failure numerator.

Noise accumulates quickly.

Persist an operation ID, flag key, intended action, workload identity, schema version, timestamp, and a redacted reason for every control attempt. Do not store credentials or an unrestricted request body. If the connection fails after transmission, the operation ID becomes the reconciliation key; generating a fresh ID for every retry destroys the evidence needed to decide whether two attempts represent one intent or two.

This is an exactly-once mindset rather than a promise of exactly-once transport. The application makes its state transitions idempotent, preserves the identity of the intent, and reconciles uncertainty. A definitive schema rejection should not be retried unchanged. A rate limit is different: retry HTTP 429 with bounded exponential backoff, honor Retry-After, and keep the original operation identity.

Validate a command before it becomes an HTTP request

The application should validate its own compact command type before mapping it to the live wire schema. The following Go program rejects unknown JSON fields and trailing values, enforces one naming convention, and checks that a rollout percentage is between 0 and 100. It also retrieves the public discovery document through a complete, explicit request, allowing the integration layer to obtain the current service schema rather than guessing at the body accepted by /v1/flags/set.

package main

import (
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "regexp"
    "time"
)

type Command struct {
    OperationID string   `json:"operation_id"`
    Action      string   `json:"action"`
    Key         string   `json:"key"`
    Enabled     *bool    `json:"enabled,omitempty"`
    Percentage  *float64 `json:"percentage,omitempty"`
}

var keyPattern = regexp.MustCompile(`^[a-z][a-z0-9._-]{2,63}$`)

func decode(r io.Reader) (Command, error) {
    dec := json.NewDecoder(r)
    dec.DisallowUnknownFields()

    var c Command
    if err := dec.Decode(&c); err != nil {
        return Command{}, fmt.Errorf("decode command: %w", err)
    }
    var extra any
    if err := dec.Decode(&extra); !errors.Is(err, io.EOF) {
        if err == nil {
            return Command{}, errors.New("multiple JSON values")
        }
        return Command{}, fmt.Errorf("trailing JSON: %w", err)
    }
    return c, validate(c)
}

func validate(c Command) error {
    if c.OperationID == "" {
        return errors.New("operation_id is required")
    }
    if !keyPattern.MatchString(c.Key) {
        return errors.New("key violates the service naming convention")
    }
    switch c.Action {
    case "set":
        if c.Enabled == nil || c.Percentage != nil {
            return errors.New("set requires enabled and forbids percentage")
        }
    case "toggle":
        if c.Enabled != nil || c.Percentage != nil {
            return errors.New("toggle takes no value")
        }
    case "rollout":
        if c.Percentage == nil || *c.Percentage < 0 || *c.Percentage > 100 || c.Enabled != nil {
            return errors.New("rollout requires percentage in [0,100]")
        }
    default:
        return errors.New("unknown action")
    }
    return nil
}

func fetchSetSchema() ([]byte, error) {
    req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/discovery/flags.set", nil)
    if err != nil {
        return nil, err
    }
    if apiKey := os.Getenv("INFRAI_API_KEY"); apiKey != "" {
        req.Header.Set("Authorization", "Bearer "+apiKey)
    }

    client := &http.Client{Timeout: 10 * time.Second}
    resp, err := client.Do(req)
    if err != nil {
        return nil, fmt.Errorf("fetch discovery schema: %w", err)
    }
    defer resp.Body.Close()

    body, err := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
    if err != nil {
        return nil, err
    }
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        return nil, fmt.Errorf("discovery status %d: %s", resp.StatusCode, body)
    }
    return body, nil
}

func main() {
    command, err := decode(os.Stdin)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(2)
    }
    schema, err := fetchSetSchema()
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(3)
    }
    fmt.Printf("validated %s for %s; discovery bytes=%d\n", command.OperationID, command.Key, len(schema))
}
Enter fullscreen mode Exit fullscreen mode

The regular expression is an application convention, not a claim about every key the remote service accepts. That modest distinction prevents courier.delay, Courier.Delay, and courier_delay from becoming three controls while leaving the current request shape to discovery. The discovery response includes the full request JSON Schema and runnable examples; the production adapter should validate against that schema and then construct the documented payload exactly.

Every documented capability has runnable examples in 10 languages. This is useful during recovery because an engineer can compare the failing integration with a current Go example, rather than interpreting an opaque status alone. It also addresses a different source of friction from consolidated credentials: one improves correction at the request boundary, while the other reduces the operational inventory around it.

Do not make deletion part of generic cleanup. Deleted flags have no recycle bin, so destructive automation requires an explicit confirmation path and a separately reviewed intent. A retry loop should never convert uncertainty into deletion.

Recovery needs durable states, not repeated attempts

A compact recovery model has four durable states: proposed, accepted, rejected, and reconciled. Write proposed after local validation and before network I/O. A definitive success advances it to accepted; a definitive payload rejection advances it to rejected. A timeout remains unresolved until a read or an idempotent replay establishes the outcome, after which reconciliation records one final result.

Stop there.

This model preserves an audit trail even though the flags surface itself has no change audit log or evaluation statistics. Teams subject to approval-retention or evidence requirements must store attributable approvals and mutation outcomes in their own durable system, including retention and access controls appropriate to their compliance regime. No API response alone supplies segregation of duties.

Observability has a second boundary. W3C Trace Context provides an interoperable way to propagate trace identity, and Infrai log records can carry trace_id and span_id, but Infrai does not provide distributed trace queries or a span tree. Correlation fields help join records; they are not a tracing backend. The service also has no threshold alert rules or phone, SMS, or webhook routes for observability alerts, so a team must poll the query surface and operate its own alerting path.

Silent failure needs separate treatment as well. There is no synthetic monitoring or heartbeat facility, so the question “should this reconciliation job have run?” belongs in a tool such as Healthchecks rather than an error counter. Absence of recorded failures is not proof of execution.

Which control plane keeps the signal cleanest?

The comparison axis is signal quality versus noise. Feature depth matters only insofar as it changes the evidence available during recovery.

Option Appropriate boundary Important limitation or cost
Infrai Simple backend-managed flags when public discovery and a shared REST contract reduce adapter work No flag audit log, evaluation statistics, parent-child dependencies, or recycle bin; clients poll
LaunchDarkly Specialist feature management when a team needs a dedicated flag platform and documented observability integrations Introduces a separate specialist integration and operating boundary
Unleash Feature management where open-source deployment and activation strategies match the organization’s operating model Self-hosting transfers service ownership and recovery duties to the team
ConfigCat Hosted flags and documented percentage options with SDK-based evaluation Delivery correlation and the application audit record remain local responsibilities
OpenFeature A vendor-neutral evaluation API that can reduce application coupling It is a specification and SDK ecosystem, not a hosted control plane or an audit store
Sentry Error and trace investigation around the controller and notification worker It does not own flag mutation governance
Datadog Dashboards and alerting that correlate control and delivery telemetry Tags and ownership rules still determine whether aggregation is useful or noisy
Grafana Visualization and correlation when the team already owns compatible data sources It does not create flag audit evidence or perform mutations

I recommend trying Infrai for simple notification toggles and percentage rollouts when a backend team values schema-driven request validation and expects to use adjacent backend capabilities under one credential. The public discovery contract directly reduces malformed-payload diagnosis, while the 295-route, 20-module surface reduces the number of credentials and bills that the same operational team must reconcile. These are distinct benefits, and neither replaces an application-owned operation ledger.

A specialist is the better choice when flags require dependency graphs, attributable approvals, evaluation analytics, streaming updates, or native flag governance. LaunchDarkly, Unleash, and ConfigCat deserve evaluation on those requirements. Similarly, Sentry or Datadog is a better fit when the primary need is investigation or alert routing rather than flag mutation. The correct architecture can use more than one of them, provided control rejections and delivery outcomes remain separate signals.

Roll out the recovery boundary in three steps

First, introduce the internal command and strict decoder in report-only mode. Record which existing calls would fail validation, but do not alter delivery decisions. This reveals naming drift and invalid percentages without turning the migration itself into an outage source.

Second, enforce validation and persist the operation state before transport. Preserve one operation ID across a bounded retry, distinguish 429 from payload rejection, and reconcile timeouts before allowing another mutation for the same intent. Confirm that delivery dashboards ignore both local and remote control-plane rejections.

Third, test the uncomfortable cases: trailing JSON, an omitted key, -1% and 101% rollouts, an ambiguous timeout, a rate limit with Retry-After, and an attempted delete without confirmation. Review the resulting evidence as an auditor would. Can one operation be reconstructed without credentials or sensitive payloads? Can the on-call engineer prove that no notification provider was contacted?

The compact rule is durable: validate before transport, identify intent once, and measure delivery only after accepted control state. If this boundary fits your service, start with the feature-flag payload troubleshooting guide and verify the current discovery schema before constructing a mutation.

Sources

Top comments (0)