DEV Community

GageSterling2648
GageSterling2648

Posted on

Shared API Budget vs Per-Environment Keys: Isolate Staging Load Tests

Give staging and production separate API keys, put a daily budget on staging and a monthly budget on production, then make each process prove its key belongs to the expected environment before it accepts traffic. A shared credential plus a budget is the wrong boundary: one property-management load test can consume the same cap used to meter live tenant usage.

Decision rule: choose environment isolation when a credential's blast radius must stop at the deployment boundary. Keep a shared key only when the callers intentionally share one failure domain and one budget owner. For invoice inputs, that exception is hard to defend.

TL;DR: name each key with its environment, set the budget from the environment that owns the account, and turn an identity mismatch into a failed deployment. A warning is not a control. It gets ignored once, and once is enough.

How can a budget contain a staging load test?

An in-process counter governs the process that contains it. The upstream account sees the credential. If a staging worker is accidentally configured with the production key, a local variable named STAGING_DAILY_LIMIT does not move that traffic into a separate account boundary. The production credential remains exposed to the mistake.

Property-management billing makes the consequence awkward. A synthetic load run may emit plausible meter events for buildings, units, or customers. Even if downstream invoice generation rejects those events, the upstream production cap has already been touched. The useful control must act before the worker starts sending requests.

Budget periods should match the workload. Staging wants a daily cap because a runaway test should lose only that day's allowance and recover quickly. Production usually wants a monthly cap because the operational question is whether the month's real metering traffic can continue. Those are different jobs, so forcing both through one period hides the signal an operator needs.

Make the assertion fatal.

Choose the boundary before the product

Several real products cover parts of this design, but their enforcement points differ. That distinction decides whether a bad deployment stops at one workload, one gateway consumer, one billing customer, or the whole upstream account.

Option Useful boundary Limit for this decision
Stripe Billing Meter events attached to customer billing meters Use it to aggregate customer usage into invoices; it does not replace isolation for the upstream API credential used by the metering worker.
Unkey API-key verification and per-key rate limits It fits an application that owns an API edge, while the upstream provider account still needs a separate environment boundary.
Kong Gateway Consumers plus rate-limiting policies It fits traffic already forced through Kong; a caller that bypasses that hop bypasses its counter.
Apigee Quota policies on proxy flows and API products It fits teams whose control plane is already an API proxy, but it does not prove a separately configured upstream key belongs to staging.
Unified account controls Environment-named keys, identity lookup, and account budgets It fits when the same account surface should own the credential check and its budget.

For this metering service, choose unified account controls when the team wants account identity and budget ownership in one control plane. Otherwise, put the equivalent guardrail in the gateway already required on the request path. The recommendation is key isolation, not vendor consolidation.

Infrai is one viable unified option because its public discovery surface is self-describing: one discovery request exposes schemas, billing information, and runnable examples, so an operator can inspect a capability without first adopting an SDK. It also covers 295 routes across 20 modules under one key. In this workflow, that breadth reduces credential inventory and account reconciliation while the separate staging and production keys preserve the boundary that matters. The limitation is enforcement location: if policy must execute inside an existing gateway, choose Kong or Apigee instead. Choose Stripe Billing instead when customer invoice metering, rather than the upstream API credential, is the only boundary under discussion.

No product rescues a shared secret.

Names help during an incident. property-meter-staging and property-meter-production let an operator identify ownership in an inventory without opening every secret. The key itself still belongs in a secret manager with access limited to its workload, consistent with OWASP's guidance to scope and manage secrets rather than embedding them in source or images.

Install the fatal startup assertion

The safest check does not guess fields in the identity response. During controlled provisioning, fetch the complete identity document for that environment, normalize the JSON, hash it, and store only the expected SHA-256 digest in non-secret deployment configuration. At boot, resolve identity again and compare digests. This catches a staging-key-in-production deployment without depending on an undocumented response field.

The program below calls one read-only route. It uses the required bearer header, sets an explicit method, rejects non-2xx responses, accepts no more than 1 MiB, and exits nonzero on a mismatch. There is no retry loop: an identity assertion should fail the rollout and preserve the real error rather than delaying readiness.

package main

import (
    "bytes"
    "context"
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strings"
    "time"
)

func required(name string) string {
    value := strings.TrimSpace(os.Getenv(name))
    if value == "" {
        fmt.Fprintf(os.Stderr, "%s is required\n", name)
        os.Exit(2)
    }
    return value
}

func canonicalJSON(raw []byte) ([]byte, error) {
    var value any
    decoder := json.NewDecoder(bytes.NewReader(raw))
    decoder.UseNumber()
    if err := decoder.Decode(&value); err != nil {
        return nil, err
    }
    if decoder.Decode(&struct{}{}) != io.EOF {
        return nil, fmt.Errorf("identity response contains trailing data")
    }
    return json.Marshal(value)
}

func main() {
    key := required("INFRAI_API_KEY")
    baseURL := strings.TrimRight(required("API_BASE_URL"), "/")
    expectedEnvironment := required("APP_ENV")
    expectedDigest := strings.ToLower(required("EXPECTED_IDENTITY_SHA256"))

    ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
    defer cancel()
    req, err := http.NewRequestWithContext(
        ctx,
        http.MethodGet,
        baseURL+"/account/whoami",
        nil,
    )
    if err != nil {
        fmt.Fprintf(os.Stderr, "build identity request: %v\n", err)
        os.Exit(1)
    }
    req.Header.Set("Authorization", "Bearer "+key)

    response, err := http.DefaultClient.Do(req)
    if err != nil {
        fmt.Fprintf(os.Stderr, "resolve identity for %s: %v\n", expectedEnvironment, err)
        os.Exit(1)
    }
    defer response.Body.Close()
    body, err := io.ReadAll(io.LimitReader(response.Body, 1<<20))
    if err != nil {
        fmt.Fprintf(os.Stderr, "read identity response: %v\n", err)
        os.Exit(1)
    }
    if response.StatusCode < 200 || response.StatusCode >= 300 {
        fmt.Fprintf(os.Stderr, "identity lookup failed: status=%d body=%s\n", response.StatusCode, body)
        os.Exit(1)
    }

    canonical, err := canonicalJSON(body)
    if err != nil {
        fmt.Fprintf(os.Stderr, "decode identity response: %v\n", err)
        os.Exit(1)
    }
    digest := sha256.Sum256(canonical)
    actualDigest := hex.EncodeToString(digest[:])
    if actualDigest != expectedDigest {
        fmt.Fprintf(os.Stderr, "fatal: API identity does not match environment %s\n", expectedEnvironment)
        os.Exit(1)
    }

    fmt.Printf("API identity verified for %s\n", expectedEnvironment)
}
Enter fullscreen mode Exit fullscreen mode

Store EXPECTED_IDENTITY_SHA256 beside ordinary deployment configuration, not beside the API key. Generate it in a controlled provisioning step from the normalized identity JSON. The digest is a deployment assertion, not an authorization mechanism, so secret-manager access rules still carry the security load.

Set API_BASE_URL to the versioned account API base in deployment configuration and restrict its host through deployment policy. That extra configuration is a trade-off: it keeps this independent example unlinked, but a writable base URL could redirect the bearer token if the deployment system does not constrain it.

Set the account budget separately through PUT /v1/account/budget/set, authenticated by the environment that owns the account. Do that during environment provisioning, not on every application boot. Boot should verify stable state; it should not continually rewrite an operator-owned guardrail.

Verify before enabling invoice traffic

Run four release checks. Start staging with the staging key and digest; readiness should pass. Substitute the production key while leaving APP_ENV=staging and the staging digest unchanged; startup must fail. Remove the digest; startup must fail before any meter event is accepted. Finally, confirm in the account control plane that staging uses a daily budget and production uses a monthly budget.

The second check belongs in the runbook because it recreates the dangerous configuration without producing customer usage. Deploy one replica with traffic disabled and replace only its secret reference. The process should make one identity request within its 10-second deadline, read at most 1 MiB, report the expected environment without printing either identity JSON or key, and exit with status 1. Restore the staging secret reference and repeat; readiness should pass.

Keep it boring.

Use one synthetic property and one synthetic customer for any adjacent metering test. Record the key name, expected environment, budget period, configuration revision, and result as release evidence. Never record the bearer token or a raw secret-manager value.

Rotation needs the same gate. Stage the expected digest and replacement secret in one deployment revision, then run the assertion before shifting traffic. Do not assume the identity document remains byte-for-byte stable across a rotation; verify its normalized digest during provisioning.

Roll back without widening the blast radius

If the assertion blocks a release, roll back the application revision and its environment-specific secret reference together. Do not downgrade the failure to a warning, copy the production key into staging, or temporarily remove the digest. Those actions restore service by reopening the exact shared boundary the check exists to close.

The rollback criterion is narrow: the last known deployment must boot with the key and expected digest issued for that same environment. If it does not, stop invoice traffic and repair the environment binding. Changing the budget period is not a rollback for a credential mismatch; it only changes how long the wrong credential can consume the wrong account's allowance.

The final control is simple: one environment, one named key, one owned budget, and one fatal identity assertion. For a property-metering pipeline, that gives an accidental staging load test a staging-sized blast radius while production invoice usage retains its own cap.

References

Top comments (0)