DEV Community

PerNilsson3147
PerNilsson3147

Posted on

One API Key, Many Capabilities: 4 Provisioning Changes Explained

A single provisioning call can make many capabilities available under one API credential, but it does not make them one undifferentiated service. The useful change is contractual: onboarding becomes one state transition, while authorization, usage attribution, limits, and credential lifecycle remain separate controls. For a developer tool that meters each customer for an invoice, I would accept the design only if those four contracts can be tested independently.

TL;DR: create the tenant and its policy once, return a secret once, then attach every capability call to the same stable tenant identity. Record usage by tenant, capability, request identity, and quantity. Enforce a spend ceiling before expensive work starts, and decide explicitly whether uncertainty should refuse traffic or permit bounded overage. The short provisioning path is valuable. The hidden shared state is where the risk lives.

What actually changes with a single call?

The simple mental model is that one call creates one key, and the key opens several capabilities. That skips the important part. Provisioning must commit a tenant identity, an allowed capability set, a budget policy, and a credential binding as one observable outcome. A caller should never have to guess whether the key exists while its usage ledger or policy does not.

What changes for onboarding automation is the unit of work. Instead of coordinating a credential call, several entitlement calls, and a metering setup call, automation submits one desired account specification. A successful response means that specification is ready to use. A failed response means it is not ready. Internal components may still perform several writes, but those steps sit behind a stable provisioning state.

One call is not one permission.

The credential should identify the calling account, while policy decides what that account may do. Keeping those concepts apart lets an operator disable one capability without rotating every integration. It also avoids encoding mutable plan details into a secret that may remain in a deployment system for months. OWASP's secrets guidance treats lifecycle, rotation, revocation, expiration, and auditing as operational concerns; a shorter signup flow removes none of them.

The four contracts I would test

First is atomic onboarding. Repeating the same logical request must not create two billable tenants, two active credentials, or two budget records. Give the request an idempotency identifier, persist its result, and return the same account outcome when automation retries after losing a response. The storage transaction depends on the system, but the visible rule should be boring.

Second is authorization. Each request resolves the credential to a tenant, then evaluates the requested capability against current policy. A valid key is necessary, not sufficient. Revocation and rotation also need deterministic tests: after a replacement is activated, the old secret follows the documented overlap or rejection rule rather than working indefinitely.

Third is attribution. An invoice meter needs more than a total counter. Every accepted usage record should carry a stable tenant ID, capability, event ID, quantity, and occurrence time. The event ID makes retry handling testable. The capability dimension explains an invoice without reconstructing intent from endpoint names.

Fourth is enforcement. A spend ceiling is a decision boundary, not a dashboard decoration. Before work begins, the service reserves an estimated amount against the tenant's remaining allowance. After completion, it settles that reservation against actual metered usage. Without a reservation, concurrent requests can each see room under the ceiling and collectively exceed it.

Here is a focused TypeScript sketch. The numbers are sample units for a test fixture, not prices or benchmark results.

type Capability = "completion" | "embedding" | "rerank";

type ProvisionInput = {
  requestId: string;
  customerRef: string;
  capabilities: Capability[];
  monthlyCeilingUnits: number;
  onLimit: "reject" | "allow-bounded-overage";
};

type ProvisionResult = {
  tenantId: string;
  apiKey: string; // Return once, then store through a secrets manager.
  policyVersion: number;
};

const fixture: ProvisionInput = {
  requestId: "onboard-acme-0042",
  customerRef: "customer-1048",
  capabilities: ["completion", "embedding"],
  monthlyCeilingUnits: 50_000,
  onLimit: "reject",
};
Enter fullscreen mode Exit fullscreen mode

The result deliberately returns no mutable quota balance. A balance observed during provisioning is stale before the first workload request arrives. Policy versioning is more useful because later authorization and usage records can name their precise decision context.

That distinction pays for itself during a billing dispute.

Choose the failure policy before production

A hard ceiling and uninterrupted service conflict at the boundary. There is no configuration that guarantees both when usage arrives concurrently and metering has delay. The product decision is which error is acceptable for each tenant: refuse legitimate traffic near the limit, or accept a bounded amount beyond it.

For a prepaid trial or an account with strict contractual limits, I would choose rejection when a reservation cannot be obtained. That makes the ceiling meaningful, though customers can see refused requests while reconciliation catches up. For an established account whose developer workflow must continue, bounded overage can be reasonable if the bound is explicit, small in the system's own accounting units, and visible to operators. This trade-off belongs in tenant policy, with a version recorded beside each decision.

The awkward case is partial execution. If a capability starts costly work and the client disconnects, billing cannot be derived from the HTTP response alone. Define the billable event for each capability: accepted input, completed output, or measured resource consumption. The meter emits that event under a stable event ID. Retries may deliver it again, but ledger ingestion deduplicates it.

HTTP success is not a ledger.

Test the boundary, not just the happy path

A green test that provisions an account and calls two capabilities proves little. The stronger suite applies pressure where one-call onboarding concentrates risk. Run two provisioning requests with the same idempotency identifier concurrently. Retry after the server has committed but before the client receives the response. Rotate the secret while requests are in flight. Remove one capability and confirm the others still work. Submit the same usage event twice. Race enough reservations to cross the sample ceiling by one unit.

I would also inject a delayed meter. The request path should apply the documented reservation rule rather than trust a lagging aggregate. Then stop the ledger consumer, fill the permitted uncertainty window, and verify that policy either rejects new work or admits only bounded overage. This test connects a spend promise to behavior.

The one-unit race matters.

Observability follows the same boundaries. Track provisioning outcomes by request ID, authorization decisions by policy version, reservation-versus-settlement deltas, duplicate usage events, and refusals caused by the ceiling. Never put the raw credential in logs. Audit secret access and lifecycle operations so a leaked log archive cannot become an authentication database.

Keep invoice reconciliation outside the hot request path. Serving needs a fast decision and a durable event; billing can aggregate accepted events per tenant and capability, compare reservations with settlements, and flag unexplained differences. That split reduces latency pressure without weakening attribution.

Fast serving and slow reconciliation can coexist.

What should you measure before copying this design?

Start with retry ambiguity: how often does automation retry because it cannot tell whether provisioning completed? Then measure authorization latency, reservation contention, settlement delay, duplicate-event rate, and requests refused at the ceiling. These values reveal whether a single call simplified the customer's job or merely moved coordination somewhere less visible.

Measure secret operations too: age of active credentials, rotation completion time, use of revoked credentials, and access to secret storage. A credential shared across capabilities has wider impact when exposed, so lifecycle evidence matters more, not less.

Use one provisioning call when you can offer one atomic account outcome and preserve separate, inspectable contracts for policy, metering, enforcement, and secret lifecycle. Keep a multi-step workflow when those parts have different owners or cannot fail as one unit. Convenience at signup is useful only when the invoice and refusal decision remain explainable afterward.

Further reading

Top comments (0)