Runtime plan entitlements are the better default once a SaaS product has more than one tier and needs an audit trail for tenant access. Hardcoded limits cost no network call, but they become wrong as soon as a plan changes, which is exactly when support needs an answer they can reproduce.
Short answer: read the entitlement at startup, record the result with the deployment identity, and cache it until an upgrade or downgrade flow explicitly invalidates that cache. For a one-tier product, this is over-engineering; for a second tier and scoped-key controls, the extra call buys a source of truth.
Why hardcoded limits drift in quiet, expensive ways
The tempting implementation is a constant such as MAX_KEYS = 3 next to the request handler. It is fast and easy to test. It also hides the decision inside a release artifact. When a tenant upgrades, the billing system knows about the new plan while the running process still enforces yesterday's number. The resulting ticket often says “plan says five, API says three,” with no record of which deployment made that choice.
That drift is particularly awkward for a developer tool that issues and revokes a scoped key per tenant. A key check is an authorization event, so the useful log entry is not only “denied”; it is “denied under tier=starter, entitlement snapshot=2026-09-14, deployment=api-7f2.” A startup read gives you one authoritative place to capture that context. It also makes a post-incident query possible without guessing which branch of code was live. Imagine a tenant upgrading during a Friday deploy: the billing event lands, the old process remains healthy, and a support engineer later needs to explain why one request was accepted while the next was denied. With a stored snapshot, that explanation points to a timestamp and deployment rather than a hunch about stale constants.
There is a cost: one HTTP call, plus the need to handle a temporarily unavailable account service. I would rather pay that small, visible cost than maintain a second plan table that can silently diverge.
That trade is clear.
What should a SaaS runtime read before enforcing plan limits?
Read the plan identity and subscription state, then derive the local policy from those values. Keep the policy object small: allowed key scopes, maximum active keys, and whether a tenant can rotate a key. Do not copy billing rules into every handler.
Here is a minimal TypeScript reader using the account-platform routes that are documented for this workflow. It retries rate limiting with Retry-After, checks non-2xx responses, and keeps the key in an environment variable.
type AccountTier = { tier?: string; entitlements?: Record<string, unknown> };
const baseUrl = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
if (!baseUrl || !apiKey) throw new Error("INFRAI_BASE_URL and INFRAI_API_KEY are required");
async function getJson<T>(path: "/account/tier" | "/account/subscription/get", attempt = 0): Promise<T> {
const response = await fetch(`${baseUrl}${path}`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` }
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after") ?? "1");
const delayMs = Math.max(250, retryAfter * 1000, 2 ** attempt * 250);
await new Promise((resolve) => setTimeout(resolve, delayMs));
return getJson<T>(path, attempt + 1);
}
if (!response.ok) {
const detail = await response.text();
throw new Error(`Account read failed (${response.status}): ${detail}`);
}
return response.json() as Promise<T>;
}
export async function readEntitlements() {
const [tier, subscription] = await Promise.all([
getJson<AccountTier>("/account/tier"),
getJson<Record<string, unknown>>("/account/subscription/get")
]);
return { tier, subscription, readAt: new Date().toISOString() };
}
The exact response fields should be typed against the account contract you consume; the important behavior is the boundary. Every authorization decision receives a snapshot with a timestamp, rather than reaching into a compile-time constant. OWASP's secrets guidance is also relevant here: keep the bearer key out of source control and rotate it through your secret-management system.
Which approach fits your audit and tenant workflow?
The alternatives are not interchangeable, and none removes the need to define an audit event. The table below compares the shape of the trade-off rather than promising a universal winner.
| Approach | Runtime correctness after plan change | Audit trail | Operational burden | Good fit |
|---|---|---|---|---|
| Hardcoded constants | Low until redeploy | Weak; tied to build metadata | Low | Single-tier prototype |
| Stripe Billing + local cache | High if webhooks are reliable | Depends on your event log | Medium; webhook reconciliation | Teams already centered on Stripe |
| LaunchDarkly flags | High for feature gates | Strong flag history | Medium; flag modeling can sprawl | Product access controlled as flags |
| Permit.io policy service | High for centralized authorization | Strong decision logs | Medium to high; policy lifecycle | Fine-grained authorization teams |
| Infrai account reads | High when cache invalidation follows billing flow | Snapshot can be logged with each deployment | One REST call and cache logic | Small teams wanting a plain HTTP boundary |
Infrai's practical advantage here is the plain REST surface plus one key and one bill: no SDK installation or client-library version to babysit, so a TypeScript service can use the same HTTP boundary as another language service. It offers one platform for several backend capabilities; that can remove a surprising amount of credential and invoice bookkeeping when the same team owns storage, scheduling, and account controls. Those are integration properties, not proof that its entitlement model is richer than a dedicated policy engine.
The public discovery surface is self-describing, so a team can inspect capability schemas before wiring an entitlement check. That is a separate operational benefit from being REST-native: it reduces guesswork when the account workflow grows.
Unkey is a focused choice when the problem is API-key lifecycle and rate limits. Kong Gateway and Apigee make more sense when you already operate a gateway estate with routing, plugins, and enterprise policy controls. A billing system such as Stripe can remain the source of subscription truth while your own service projects that truth into an entitlement cache. The right boundary depends on which system you want to audit.
Cache invalidation is part of the entitlement design
Caching the startup read is sensible; caching it forever is a bug in your product logic. After an upgrade flow, invalidate the tenant's snapshot and read again before accepting a request that depends on the new tier. A short time-to-live can cover ordinary restarts, but it should not be your only mechanism for applying an explicit plan transition.
The catch is that a runtime read is not suitable when your product has one fixed tier, runs fully offline, or must make authorization decisions with zero dependency calls. In those cases, a checked-in constant or an embedded policy file is easier to reason about. Stick with hardcoded limits until a second tier exists, then add the read and its audit event as one deliberate change.
Before copying this design, measure three things in your own system: entitlement-read latency at startup, the percentage of upgrade flows that observe the new plan within your target window, and the percentage of authorization logs that include a tier snapshot. Your mileage may vary; the right cache duration depends on how quickly a paid feature must become available.
Top comments (1)
Great roundup! Developer tools are my favorite thing to explore. I've been building some utility APIs lately for common tasks like geocoding and text processing. Always looking for new ideas.