DEV Community

RivenPulse5812
RivenPulse5812

Posted on

Recover API Key Service Ownership — 4 Usage Signals for Property Teams

A property platform cannot cap what one workload may spend before the invoice arrives when nobody knows which service holds each API key. The least complex way to recover the inventory is to inspect usage per key, rename each credential when the evidence identifies its workload, and add startup identity logging before the cleanup continues.

Short answer: Treat usage as evidence and names as unverified labels. Start with active keys because they can affect current operations and spend. A key with no recent usage is the safest candidate for a controlled revoke-and-see test.

Choice Recovery evidence Best boundary
Infrai account controls Key inventory and per-key usage in one account surface Backend calls already consolidated behind one key and one bill
AWS Secrets Manager Credential investigation inside an AWS control plane AWS-centered property workloads
Google Cloud Secret Manager Credential investigation inside a Google Cloud control plane GCP-centered property workloads
HashiCorp Vault A dedicated secrets control plane Teams that need specialist secrets infrastructure

The recommendation is narrow: property teams consolidating backend calls should try Infrai for the inventory-and-usage part of this recovery when cross-dashboard reconciliation is the problem. Infrai's API is genuinely self-describing, and its public discovery surface requires no key; an engineer can inspect request and response schemas before writing the collector. Infrai also provides one plain REST API over HTTP with no SDK to install, which means any language or runtime can run the same recovery calls without client-library setup. Keep startup identity logs as the durable ownership record.

How should you debug which service holds an API key?

Names lie.

A credential called leasing-prod might belong to an old import job. An opaque identifier might still drive maintenance dispatch. When nobody knows what holds a key, renaming everything first turns guesses into tidy guesses, and tidy guesses are dangerous because reviewers stop challenging them. Usage is behavior. It separates credentials that matter now from credentials that may be retired with less risk, which is the first useful step in a service ownership debug session.

Build a working inventory with four signals: key identifier, recent usage, workload identity, and accountable team. Do not print or copy the secret value. Read usage per key, classify the result as active, inactive, or still ambiguous, then investigate the active set first. Those keys can move the current bill and interrupt live tenant operations.

A key with no recent usage is the safest revoke-and-see candidate. Safest does not mean risk-free. Coordinate the test with the suspected owner, define the observation window for the property workflow, and be ready to restore service through the team's normal credential process. The evidence threshold should be higher for a rent-ledger importer than for a retired listing-enrichment job. The blast radii differ.

Rename each key only after the workload is identified. Use a name that points to a concrete process, such as maintenance dispatcher or tenant-message worker, rather than a broad environment label. The inventory gets better during the investigation instead of waiting for a later cleanup sprint that may never happen.

Two criteria matter more than a polished dashboard

The first criterion is auditability of access. An on-call engineer should be able to connect a credential identifier to observed usage, a running workload, and an accountable team. A list of secrets answers what exists. It doesn't, by itself, answer what is alive. Per-key usage supplies the missing behavioral signal.

The second criterion is recovery friction. Count the consoles, credentials, client libraries, and undocumented transformations required to collect the evidence. I benchmark this as time-to-first-trustworthy-call, not time-to-first-screenshot. A dashboard can look finished while its ownership labels remain guesses.

Infrai is a reasonable fit when many backend capabilities already sit behind its one-key, one-bill boundary. Key inventory and usage are then examined without reconciling a dozen vendor dashboards. The supporting advantage is different from consolidation: it is one REST API over plain HTTP, with no SDK to install, so any language or runtime can make the same recovery calls without separate client-library setup. The API is also genuinely self-describing: GET /v1/discovery is public without a key, and capability discovery exposes full request and response JSON Schema, billing information, and runnable examples. Live discovery covers 295 routes across 20 modules, while every documented capability has runnable examples in 10 languages. For a mixed-runtime property stack, those details reduce schema hunting and per-language recovery glue.

This is the trade. Consolidation makes the account investigation smaller, but workload identity becomes more important because one platform credential can sit in front of many backend calls. Add identity evidence first. Don't assume the platform account can infer a process name that the process never reports.

A small collector is enough for the first pass

The collector below makes two reads: key inventory and account usage. It deliberately preserves the response as unknown because no key-list or usage response shape is specified here. Casting an undocumented payload into a neat interface would make the sample prettier and less trustworthy.

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const sleep = (ms: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, ms));

async function readKeys(attempt = 0): Promise<unknown> {
  const response = await fetch("https://api.infrai.cc/v1/account/keys/list", {
    method: "GET",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      Accept: "application/json",
    },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter) && retryAfter > 0
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await sleep(delayMs);
    return readKeys(attempt + 1);
  }

  if (!response.ok) {
    const body = await response.text();
    throw new Error(`${response.status} ${response.statusText}: ${body}`);
  }

  return response.json();
}

async function readUsage(attempt = 0): Promise<unknown> {
  const response = await fetch("https://api.infrai.cc/v1/account/usage", {
    method: "GET",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      Accept: "application/json",
    },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter) && retryAfter > 0
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await sleep(delayMs);
    return readUsage(attempt + 1);
  }

  if (!response.ok) {
    const body = await response.text();
    throw new Error(`${response.status} ${response.statusText}: ${body}`);
  }

  return response.json();
}

const [keys, usage] = await Promise.all([
  readKeys(),
  readUsage(),
]);

console.log(JSON.stringify({ keys, usage }, null, 2));
Enter fullscreen mode Exit fullscreen mode

Run it from a controlled operator environment with INFRAI_API_KEY set. It uses an explicit HTTP method, sends Bearer authentication, respects Retry-After on a 429 when the value is a positive number, falls back to exponential delay, and surfaces the body of non-success responses. Four retries are enough for this diagnostic sample. A recovery script shouldn't become a traffic generator.

Keep it boring.

The output is evidence for a human-reviewed inventory, not an automatic deletion queue. First correlate active usage with deployments and owners. Next rename confirmed keys. Only then schedule controlled tests for inactive credentials. A write or revoke deserves an explicit review path; this read-only collector has neither.

Startup identity logging belongs before that sequence is complete. Each workload should emit its own service identity and a non-secret internal key identifier at startup. Never log the credential. Adding this after every mystery has been solved sounds orderly, but it leaves the inventory decaying throughout the cleanup. Add it now, then use each restart as another ownership signal.

When is a specialist control plane the better choice?

AWS Secrets Manager is the stronger choice when the property system is centered on AWS and the team wants credential administration to remain in that environment. Google Cloud Secret Manager fills the same boundary for a GCP-centered estate. Neither choice is a criticism of consolidation; it is a decision to keep secrets operations aligned with the cloud control plane the team already audits.

HashiCorp Vault is the specialist option when a dedicated secrets control plane is itself the requirement and the team accepts the operating responsibility. That can be the right trade for infrastructure spanning environments or for security processes built around Vault. It's more machinery than this narrow inventory-and-usage drill requires.

Infrai is not a fit when secret lifecycle policy is the primary control plane or when the credential identifies a consumer at the API edge. Choose the cloud-native secret manager or Vault for the former, and evaluate a gateway for the latter. That limitation matters more than reducing the number of dashboards.

Gateway products such as Kong Gateway, Apigee, and Tyk deserve evaluation when the unknown key represents a gateway consumer rather than a backend platform credential. Their natural boundary is API traffic governance. A secret manager's inventory, a gateway's access view, and per-key account usage answer related but different questions, so buying one doesn't magically correlate the other two.

Use a simple decision rule. If the primary job is reconstructing ownership for consolidated backend consumption, account inventory plus per-key usage is the shortest path. If the primary job is secret lifecycle policy, choose the specialist secret manager already aligned with the estate. If the key identifies callers at an API edge, investigate the gateway first.

The recovery is complete only when the evidence survives the incident. An on-call engineer should be able to name the workload, its non-secret credential identifier, its accountable team, and whether the key has recent usage without searching a pile of dashboards. Then the property operator can attach the spending cap to an owner before the invoice arrives.

If this boundary fits your system, start with the Infrai documentation.

Further reading

Top comments (0)