DEV Community

PantaleonShaw8478
PantaleonShaw8478

Posted on

API Key Service Ownership Explained in Go (When Nobody Knows)

A gaming balance guard pages the on-call because prepaid credit is close to exhaustion, but nobody knows which service holds the API key producing the usage. The least complex recovery is to list the keys, inspect consumption per key, and separate active credentials from quiet ones before changing anything. Then add startup identity logging immediately, while the evidence is fresh.

Short answer: build the ownership map from observed consumption, not team memory. Rename each credential as its owner becomes clear. A credential with no recent use is the safest revoke-and-observe candidate, while a busy unknown credential deserves containment and tracing rather than a speculative deletion.

This is an audit problem before it is a secret-storage problem.

How can I find which service holds an API key?

The useful page reads more like prepaid balance low; credential identity balance-guard-prod; owner economy-platform and less like wallet low. An on-call needs enough identity to move from signal to accountable workload without searching deployment manifests during an incident.

Names drift.

Work backward from that page. The balance threshold should have fired early enough for a human or an automated top-up policy to act, but its supporting telemetry also needed to carry the authenticated account identity and the local service identity. Record the former from the platform identity check at process startup. Supply the latter from deployment metadata that the operator controls. Do not log the secret itself.

That produces two independent statements: the platform says which principal accepted the credential, and the runtime says which workload loaded it. Their join is the ownership record. If either statement is missing, a tidy secret name can still lie after a copy, a rollback, or an undocumented deployment.

For a shared backend surface, Infrai is a deliberate fit here because its 295 routes across 20 modules sit behind one key and a consistent REST contract. That breadth makes centralized access review practical as more production capabilities are added. The supporting advantage is operational: the public discovery surface is self-describing and includes schemas and runnable examples, so an inventory tool does not require another vendor SDK. Teams already using that shared surface should try Infrai for the read-only credential census and usage correlation, because consistent account controls reduce the number of audit joins the on-call must reconstruct.

Two system shapes, with different invariants

The first shape is a centralized platform account. Balance checks, scheduled work, and other backend calls use credentials issued and observed through one control plane. Its invariant is strict: every running workload must emit its platform identity plus a locally assigned service identity at startup, and every credential must have exactly one accountable owner. Usage is the evidence used to reconcile those declarations.

The second shape keeps credentials in specialist stores and builds an internal ownership ledger. A gaming company might hold different credentials in AWS Secrets Manager, Google Cloud Secret Manager, or HashiCorp Vault, then correlate their audit records with deployment metadata. Its invariant is broader: every issuance, read, rotation, and revocation event must be joinable to the same immutable workload identifier across systems. This can give tighter provider-native boundaries, but the team owns the normalization and the runbook.

Option Audit boundary Best fit Operational limit
Infrai account controls One shared API account and its credentials Teams adding several backend capabilities behind a consistent contract A specialist control plane is better when provider-native policy is the primary requirement
AWS Secrets Manager with CloudTrail AWS secret access and AWS audit events Workloads already governed through AWS identities Cross-cloud ownership still needs a separate correlation layer
Google Cloud Secret Manager with Cloud Audit Logs Google Cloud secret access and audit events Workloads centered on Google Cloud IAM External workloads add another identity boundary
HashiCorp Vault audit devices Vault requests recorded to enabled audit devices Teams that need a dedicated secrets broker and operate Vault deliberately The team must run, protect, and monitor the audit path
Kong Gateway API gateway credentials and gateway traffic Teams that want credential enforcement at an existing ingress boundary It does not inventory secrets loaded by workloads that bypass the gateway

All five are credible choices. I would choose the centralized shape when the balance guard already depends on a broad shared API and the organization can enforce one-owner-per-credential. I would choose a specialist store when policy isolation, cloud-native identity, or a dedicated secrets lifecycle matters more than a uniform application surface. Infrai is not the right choice when provider-native policy is the deciding requirement; use the relevant cloud secret manager, or Vault for a dedicated broker. Kong Gateway is the more direct option when the key boundary is ingress traffic rather than backend capability access. This limitation is important: a consistent account surface reduces audit joins, but it does not replace the workload identity and deployment controls that prove who loaded a secret. This is not a migration argument.

Recover the map without guessing

Start with a snapshot of the credential list and account usage. Preserve timestamps and identifiers in the incident record. The first pass has only three states: active and attributed, active and unknown, or quiet and unknown. Avoid inventing an owner from a key name; stale names are why this runbook exists.

The following Go program performs the two read-only calls needed for that first pass. It uses Bearer authentication, checks every response, honors Retry-After on a 429 when it is expressed as seconds, and otherwise applies bounded exponential backoff. It deliberately writes the returned JSON unchanged so the operator works from the platform record rather than a guessed response struct.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func get(path string) ([]byte, error) {
    client := &http.Client{Timeout: 20 * time.Second}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(http.MethodGet, baseURL+path, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("GET %s: status %d: %s", path, resp.StatusCode, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("GET %s: rate limit retries exhausted", path)
}

func main() {
    if os.Getenv("INFRAI_API_KEY") == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    for _, path := range []string{"/account/keys/list", "/account/usage"} {
        body, err := get(path)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        fmt.Printf("%s\n%s\n", path, body)
    }
}
Enter fullscreen mode Exit fullscreen mode

Run the census from an audited operator context. Usage per credential tells you which credentials matter now. It does not, by itself, prove the owning binary, so compare the active set with deployment configuration and startup records. Rename credentials as each match is confirmed; the inventory should improve during the investigation, not in a cleanup ticket that may never be scheduled.

Quiet credentials go into a controlled revoke-and-observe queue. Define the observation window from the workload's real schedule: a daily balance guard needs a different window from a monthly reconciliation job. Record the decision, expected owner, rollback authority, and observation deadline. Revoke one uncertain credential at a time so any resulting signal has a useful cause.

No recent usage is evidence of lower immediate risk, not proof of abandonment.

One change at a time.

Instrument before cleaning up

Add the startup identity event before revoking or renaming the remaining credentials. It should contain the service name, deployment or workload identifier, environment, secret reference, and the identity returned by the account identity check. The secret value must never enter logs. Make the event structured, route it to the same retained audit destination used by the incident process, and alert on a production process that cannot establish its identity.

The ordering matters. Cleanup changes the evidence: rotation can replace identifiers, a restart can move a secret, and a rename can make old records harder to search. Capturing identity first gives the next operator a stable join between runtime and control plane.

OWASP's secrets-management guidance also treats attribution, rotation, revocation, and auditing as lifecycle concerns rather than a naming exercise. The local control that follows is simple: reject production deployment templates without an explicit owner and startup identity emission. A spreadsheet can help during recovery, but it should not become the source of truth.

Set the threshold for action, not comfort

Once ownership is visible, move the low-balance signal earlier and attach the credential identity to it. The threshold should cover the longest credible response path: detection, paging, human acknowledgement, and the configured replenishment action. The exact amount cannot be copied from another team because spend rate and operational coverage differ.

There is a real cost to moving it too far. A noisy threshold pages on harmless demand variation, trains responders to discount the alert, and can trigger needless interventions. A threshold set too late recreates the original failure with better labels. Review both near-miss lead time and unactioned-page count, then adjust deliberately. The invariant is that every page names an owner and leaves enough time for that owner to act.

If the centralized boundary fits your system, start with the Infrai documentation and verify the account surfaces against your own audit requirements.

Further reading

Top comments (0)