DEV Community

BenedictVance6863
BenedictVance6863

Posted on

Tenant Billing Permission Error — Missing Scoped Key Capability on One Code Path

Issue a separate scoped key for each property-management tenant, and diagnose an isolated permission failure by comparing the failing operation's required capability with the capabilities and tenant binding of the actual key sent on that path. Short answer: if routine reads work but key revocation fails, do not broaden the key immediately. First establish which credential the revocation worker used, which tenant it belonged to, and whether revocation was authorized for that credential. A shared key can make the request succeed while quietly destroying billing attribution.

A typical flow starts when a tenant administrator requests a new integration key. The control plane records the tenant, the requested capabilities, and a non-secret key identifier; a worker later uses an appropriately authorized credential to revoke that key. Usage records carry the tenant identifier established by authentication, not one copied from an untrusted request body. If the application has a Node.js route for reads and a separate background path for revocation, those paths may select different credentials even though they share an API client module. The symptom can look like a mysterious one-route permission error.

One route fails. Why?

Why does a scoped key report a permission error on one code path?

Trace the credential selection boundary before changing permissions. In the Node.js application, record a correlation ID, operation name, authenticated tenant ID, credential identifier or fingerprint, and outcome at the point where the outgoing request is built. Never log the secret itself. Compare a successful tenant read and the failing revoke attempt with the same tenant and deployment configuration. A capability-discovery response, if the service exposes one, helps confirm what the presented key can do; it does not prove which key the worker actually loaded until you tie it to that request's credential identifier. For example, a request queued with tenant property-group-17 might execute in a worker whose default credential belongs to the maintenance account. If the worker chooses that default before consulting the queued tenant, the successful read from the web route tells you nothing about its authorization decision. Trace the credential selection for both paths, then match the denial against the expected capability map.

For a notebook-to-production check, I would model the authorization decision as data and run the same cases in CI. The following Python example is deliberately local: it makes no claim about a particular service's endpoints or error schema. It keeps the tenant binding separate from the capability check, because passing one must not imply the other.

from dataclasses import dataclass


@dataclass(frozen=True)
class Credential:
    key_id: str
    tenant_id: str
    capabilities: frozenset[str]
    active: bool = True


def authorize(credential: Credential, tenant_id: str, required: str) -> str:
    if not credential.active:
        return "inactive_key"
    if credential.tenant_id != tenant_id:
        return "tenant_mismatch"
    if required not in credential.capabilities:
        return "missing_capability"
    return "allow"


tenant = "property-group-17"
reader = Credential("reader-42", tenant, frozenset({"usage:read"}))
manager = Credential(
    "manager-19", tenant, frozenset({"usage:read", "keys:revoke"})
)

assert authorize(reader, tenant, "usage:read") == "allow"
assert authorize(reader, tenant, "keys:revoke") == "missing_capability"
assert authorize(manager, tenant, "keys:revoke") == "allow"
assert authorize(manager, "other-property-group", "keys:revoke") == "tenant_mismatch"
assert authorize(Credential("old-8", tenant, manager.capabilities, False),
                 tenant, "keys:revoke") == "inactive_key"
Enter fullscreen mode Exit fullscreen mode

That five-case harness is a specification, not an implementation of remote authorization. Replace its capability strings with the service's documented operation-to-capability mapping and compare the outcome with the real API's response in an isolated test environment. For a Node.js failure, inspect the final outbound request construction too: a worker environment variable, per-tenant credential lookup, or retry path can select a different key from the one exercised in the notebook. Do not put live secrets in a test fixture.

The limitation of this local harness is important: it cannot verify the service's policy, the live credential selected by a worker, or the tenant attached to a billable event. It is useful for pinning down expected behavior during development; integration tests and usage reconciliation still carry the production evidence. That is a trade-off I would keep visible when moving a notebook check into CI.

How does one path lose tenant attribution?

The failure may be a correct denial. A read-only credential should not be able to revoke another credential, even for its own tenant. The tempting fix is to grant keys:revoke to every tenant-facing key. That expands the blast radius of a leaked integration key and confuses the distinction between an integration's usage and an administrative action. Prefer a narrowly authorized control-plane credential for lifecycle operations, with the actor and target tenant recorded separately. The service must enforce the target tenant binding; an application-side assertion alone is not an authorization boundary.

Keep those identities separate.

There is a second failure worth testing: the revoke request succeeds under a shared administrative key, but downstream usage gets attributed to that shared identity rather than the property group's identity. Whether this happens depends on the service's billing and audit contract, so verify it against actual usage records instead of assuming an API response proves attribution. Keep authenticated_tenant, target_tenant, actor, and key_id distinct in your event schema. A tenant ID supplied in a URL is a requested target, not evidence of who authenticated.

OAuth describes scopes as a way to limit access and defines an insufficient_scope error for protected resources. Those conventions are useful diagnostic vocabulary, but a service-specific API key is not automatically an OAuth access token; its capability names and denial response must come from its own contract. HTTP 403 indicates that a server understood a request but refused to fulfill it. It does not by itself tell you whether the problem was scope, tenant binding, a disabled key, or policy applied later in the request.

Put the check into the deployment path

Before releasing the worker, exercise issuance, one permitted read, a denied revoke with a read-only key, a permitted revoke with the lifecycle credential, and a read after revocation. Run the matrix across two tenants. Include a retry after a denied operation to check that the retry does not silently switch credentials; include a retry after a successful revoke to verify how the API handles repeated requests. Document that behavior rather than assuming idempotency. Keep the expected capability map alongside the integration tests so a deployment changing policy cannot pass on read-only coverage alone.

On-call diagnostics should emit structured denial categories from the application where known, while preserving the upstream status and request ID. Keep key material out of logs, traces, exception text, and notebook outputs. Rotate a credential if exposure is suspected; permission debugging is not a reason to paste it into a ticket. The OWASP secrets-management guidance covers limiting access to secrets and handling their lifecycle.

Cost matters here as attribution accuracy, not a headline unit price. For an AI-assisted maintenance workflow, aggregate token-consuming operations by the authenticated tenant and compare those records against the control-plane issue/revoke events. Evaluate the integration with a small, repeatable fixture before letting a prompt-driven agent request administrative actions: an eval that checks both the denial and the billing tenant catches a regression that a successful HTTP response misses. Keep the agent's execution credential narrower than the operator's lifecycle credential.

The operational checklist is short enough to keep in prose: confirm the failing operation's documented capability, fingerprint the credential actually sent, verify its active state and tenant binding, and correlate the denial with an audit event. Then test a properly authorized lifecycle credential without substituting it for tenant usage. Finally, check the billing record and the post-revocation behavior. The right outcome is an explainable denial or an authorized change with correct attribution, not merely a green response.

References

Top comments (0)