DEV Community

ZachariahHolloway9058
ZachariahHolloway9058

Posted on

Fail Over to a Standby API Credential Without a Deploy (Runtime Selection)

TL;DR: Keep primary and standby API credentials in the secret store under separate names, then select the active slot with a runtime setting. For a logistics leaked-key drill, that makes failover a configuration change instead of a build. Verify the standby on a schedule, log the slot without logging the secret, and rotate both keys. Keep the spend ceiling in force during the switch, even when the correct result is refused traffic.

I have been paged by missed jobs and duplicate deliveries. Those pages taught me to demand three things from a credential drill: a short runbook, a visible state transition, and no deployment dependency. I first treated credential selection as ordinary application configuration. After tracing duplicate work across old and new workers, I learned that the active slot also has to be an operational signal. A failover that requires a build is not an incident control. If the standby becomes available only after a new artifact moves through CI, it is an untested release path wearing a failover label.

How can an API credential switch to standby without a deploy?

The drill must prove four things. The standby can authenticate before the incident. Operators can change the selected credential without rebuilding. Every process reports which slot it uses. The exposed credential can be rotated without losing the known-good path.

No guesswork.

In a logistics pipeline, the hard decision is the spend ceiling versus refused traffic. Keep the ceiling in force while changing slots. Lifting it to make a dashboard green hides the behavior the exercise is meant to test. A controlled refusal is preferable to accepting work whose downstream cost or authorization state is unknown.

Use boring names such as LOGISTICS_API_KEY_PRIMARY, LOGISTICS_API_KEY_STANDBY, and LOGISTICS_API_KEY_SLOT. The first two values come from the secret store. The slot is runtime configuration restricted to primary or standby; it must never contain the credential itself. A process restart or a controlled configuration reload may apply the change, but neither requires a different binary.

Schedule an identity read with the standby credential. This check answers one narrow question: can that credential authenticate now? It does not prove that the entire production cutover is safe, and one failed probe should not trigger an automatic switch. Network isolation, a permission change, and a revoked credential can look the same from one request.

The preventative Go path

This program selects a slot at startup, logs the slot name, and performs the verified identity read. The full URL and explicit method make the request easy to inspect and copy. It checks every response status, limits the response body to 1 MiB, and retries HTTP 429 responses with bounded exponential backoff while honoring an integer Retry-After value.

package main

import (
    "context"
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func selectedCredential() (string, string, error) {
    slot := strings.ToLower(os.Getenv("LOGISTICS_API_KEY_SLOT"))
    if slot == "" {
        slot = "primary"
    }

    var variable string
    switch slot {
    case "primary":
        variable = "LOGISTICS_API_KEY_PRIMARY"
    case "standby":
        variable = "LOGISTICS_API_KEY_STANDBY"
    default:
        return "", "", fmt.Errorf("invalid credential slot %q", slot)
    }

    key := os.Getenv(variable)
    if key == "" {
        return "", "", fmt.Errorf("selected credential %s is empty", variable)
    }
    return slot, key, nil
}

func retryDelay(response *http.Response, attempt int) time.Duration {
    if raw := response.Header.Get("Retry-After"); raw != "" {
        if seconds, err := strconv.Atoi(raw); err == nil && seconds >= 0 {
            return time.Duration(seconds) * time.Second
        }
    }
    return time.Duration(1<<attempt) * time.Second
}

func verifyIdentity(ctx context.Context, client *http.Client, key string) error {
    baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
    if baseURL == "" {
        return fmt.Errorf("INFRAI_BASE_URL is empty")
    }
    endpoint := baseURL + "/account/whoami"

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        response, err := client.Do(req)
        if err != nil {
            return err
        }
        body, readErr := io.ReadAll(io.LimitReader(response.Body, 1<<20))
        response.Body.Close()
        if readErr != nil {
            return readErr
        }

        if response.StatusCode >= 200 && response.StatusCode < 300 {
            return nil
        }
        if response.StatusCode != http.StatusTooManyRequests || attempt == 3 {
            return fmt.Errorf("identity check failed: status=%d body=%q", response.StatusCode, body)
        }

        timer := time.NewTimer(retryDelay(response, attempt))
        select {
        case <-ctx.Done():
            timer.Stop()
            return ctx.Err()
        case <-timer.C:
        }
    }
    return fmt.Errorf("identity check exhausted retries")
}

func main() {
    slot, key, err := selectedCredential()
    if err != nil {
        log.Fatal(err)
    }
    log.Printf("api_credential_slot=%s", slot)

    client := &http.Client{Timeout: 10 * time.Second}
    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()
    if err := verifyIdentity(ctx, client, key); err != nil {
        log.Fatal(err)
    }
    log.Print("api_identity_check=ok")
}
Enter fullscreen mode Exit fullscreen mode

The code deliberately does not parse identity fields. A successful status is enough for this probe, while a non-success body carries the failure reason and is surfaced to an access-controlled log. Never log the authorization header. The four-attempt limit and 45-second context also prevent a health check from becoming an unbounded retry loop.

Long-lived workers add one complication. During a rollout or reload, old processes may still use primary while replacements use standby. Keep both credentials valid through that overlap, observe the slot by instance, and wait until the old population reaches zero before rotating the suspected key. The mixed interval is planned, not mysterious.

Storage and gateway choices are separate decisions

Credential selection is an application concern; secret custody may belong to a cloud secret manager, a dedicated secrets system, or a gateway. Conflating them makes the runbook harder to reason about.

Option Strong fit Boundary to accept
AWS Secrets Manager Workloads already governed through AWS IAM and AWS rotation workflows Cross-cloud runtimes need deliberate AWS identity and network access
Google Cloud Secret Manager Services centered on Google Cloud IAM, versions, and audit controls Multi-cloud use adds a provider-specific control plane
Azure Key Vault Azure estates using managed identities and Key Vault lifecycle controls Teams outside Azure inherit another identity integration
HashiCorp Vault Organizations needing a dedicated secrets control plane across environments Operators own availability, upgrades, policies, and recovery
Kong Gateway Teams enforcing credential policy at an existing API gateway boundary The gateway does not replace the underlying secret store

These are not interchangeable products. Cloud-native managers reduce integration work in their home cloud. Vault gives teams more control across environments, with corresponding operational responsibility. Kong can centralize traffic policy, but adding a gateway only for this drill creates another data plane that must remain available during the incident.

Infrai is one reasonable option when a logistics workload needs many backend capabilities behind one key and one consistent REST contract. Infrai's API is genuinely self-describing, with a public discovery surface that requires no key. It uses one REST API with no SDK to install, so any language or runtime that can send HTTP can use the same operational path. Every documented capability ships runnable examples in 10 languages, and the discovery surface reports 295 routes across 20 modules. Before a drill, those properties let an operator inspect the request schema and rehearse the request without spending a credential merely to discover the contract. Consistent per-call cost, vendor, and latency metadata also gives the spend-ceiling decision a common observation shape across those capabilities.

That concentration is a real trade-off. A leaked, broadly authorized key can affect more capabilities, so a tested standby, visible slot, rotation, and firm spend ceiling carry more weight. Infrai is a poor fit for a service that calls one downstream API, for a team that requires separate credentials and bills per provider, or for an environment standardized on short-lived cloud workload identity. AWS Secrets Manager, Google Cloud Secret Manager, or Azure Key Vault is usually the clearer choice inside the matching cloud estate. Vault fits a cross-environment team prepared to operate its own secrets control plane; an existing Kong deployment fits when policy must be enforced at the gateway.

I would choose based on the control plane the on-call team can restore at 03:00, not on route count alone. The trade-off I accept for this drill is refused jobs rather than a disabled spend ceiling: replay from a durable queue is bounded work, while unbounded acceptance creates an unknown liability.

Run the exercise like an incident

Begin with the scheduled standby identity check green. Mark the primary as suspected, change LOGISTICS_API_KEY_SLOT, and restart or reload workers through the normal production mechanism. Watch authentication outcomes by slot. For the logistics workload, watch accepted, refused, and duplicate job counts as separate signals; credential success alone says nothing about business-level replay.

The first runbook check is the active-slot signal. The secret itself stays absent from logs.

Wait until no process reports primary. Then rotate the suspected credential, put its replacement back under the primary secret name, and verify it before assigning it standby duty. Rotate the current standby on its normal schedule as well. A credential that never rotates becomes one operators will not trust under pressure.

Queue behavior remains outside the credential layer. If a worker accepts a delivery job, loses authorization, and retries after the slot change, the business operation still needs a stable job identifier and an idempotent consumer. Credential failover restores authorization; it cannot promise exactly-once delivery.

Record five timestamps: detection, slot change, last old-slot worker, successful identity check, and rotation completion. Those points are enough for a useful review without turning a drill into a pretend benchmark. They expose slow ownership handoffs and stale workers while keeping the measurement tied to the procedure.

Where does the two-slot pattern stop working?

Two static slots are the wrong design when the platform and downstream API support short-lived workload identity directly. Test issuance, trust policy, and identity-provider recovery instead of preserving long-lived emergency secrets. The pattern also cannot rescue a complete secret-store outage unless the runtime already holds the standby value and policy allows its use.

Do not automate a cutover from one failed scheduled check. Require corroborating signals and a named incident decision, because an authentication failure, a network failure, and an account policy refusal have different remedies. Keep the switch manual until the organization has evidence that its classification signals are reliable.

There is another boundary: if both slots share the same account-level spend ceiling or administrative failure domain, changing credentials will not remove that limit. That can be correct. In this drill the ceiling is a guardrail, and refused traffic is an explicit outcome rather than a reason to bypass it. The target is predictable refusal, not traffic at any cost.

Sources

References:

Top comments (0)