Adding AI features to a multi-tenant product creates an operational question that application-level token counters do not fully answer:
How much did each tenant actually spend, and what stops the next request when the budget is gone?
Your application can estimate usage, but retries, concurrent requests, streaming behavior, and process failures can make that estimate drift away from what the inference layer served.
The open-source ai-gateway-tenant-meter example moves the control point closer to the model request. It combines Telnyx AI Gateway with one durable Edge actor per tenant.
The gateway meters and enforces. The actor remembers and reports.
What the example builds
The TypeScript application exposes a small HTTP API:
-
POST /provisioncreates a tenant ledger, gateway token group, and token key -
GET /spend?tenantId=...returns current spend, model usage, guardrail findings, alerts, and budget state -
POST /adjustchanges the tenant budget with optimistic concurrency -
GET /healthconfirms the worker is available
Each provisioned tenant receives:
- a 30-day AI Gateway budget
- an allowlist of models
- prompt and response guardrails
- a one-time gateway token key
- one durable
SpendLedgeractor - private SQL tables for daily spend, alerts, and guardrail events
- an hourly rollup task
- SMS notifications at configurable budget thresholds
By default, the warning threshold is 80% and the hard threshold is 100%. The sample starts with DEMO_MODE=true, so alerts are logged instead of sent while you test the flow.
Provision one gateway scope per tenant
The worker routes a tenant ID to a stable actor identity. The actor then creates a gateway token group with its own budget and guardrail policy:
const group = await client.createTokenGroup(
tenantId,
allowedModels,
monthlyBudget,
);
const key = await client.createTokenKey(
`${tenantId}-primary`,
group.id,
);
The token group request includes the controls that should travel with that tenant:
const GUARDRAILS_POLICY = {
secrets: {
prompt: "block",
response: "block",
},
dlp: {
profiles: ["financial"],
prompt: "flag",
response: "flag",
},
streaming: "buffered",
};
The tenant's application receives a one-time ltg_sk_... key and uses the OpenAI-compatible inference endpoint:
curl https://llm.telnyx.com/v1/chat/completions \
-H "Authorization: Bearer ltg_sk_..." \
-H "Content-Type: application/json" \
-d '{
"model": "Kimi-K2.6",
"messages": [
{"role": "user", "content": "Summarize this support request."}
]
}'
Every request made with that key is associated with the tenant's token group. When the group exhausts its budget, the gateway returns 403 budget_exceeded rather than relying on an application process to notice the overage first.
Keep the operational history in a durable actor
Gateway enforcement answers whether the next inference request may run. The actor answers a different set of questions:
- Has the 80% alert already been sent?
- Which models generated the spend?
- What guardrail detectors fired?
- When does the current budget period reset?
- Who changed the budget, and from what value?
Each tenant maps to one SpendLedger actor. That actor runs an hourly durable task and writes normalized usage into private SQL.
private armRollup(): Promise<string> {
return this.every(
3600,
"rollup",
undefined,
{ id: "hourly-rollup" },
);
}
The SQL update is idempotent by tenant and day:
INSERT INTO spend_days (
tenant,
day,
spend,
input_tokens,
output_tokens,
blocked,
flagged
)
VALUES (?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(tenant, day) DO UPDATE SET
spend = excluded.spend,
input_tokens = excluded.input_tokens,
output_tokens = excluded.output_tokens,
blocked = excluded.blocked,
flagged = excluded.flagged;
If the runtime restarts, the actor's state and SQL history remain available. The next rollup can continue without rebuilding the tenant ledger from scratch.
Store guardrail findings without storing the matched text
The actor reads gateway guardrail events and stores detector metadata such as:
- stage
- outcome
- detector
- finding code
- count
It does not store the sensitive text that triggered the finding.
That gives an operations team evidence that a secrets or financial-data rule fired without turning the ledger itself into another repository of sensitive prompt data.
Send each budget alert once per period
During the rollup, the actor compares the current spend with the tenant's configured budget.
const pct = pctOf(monthToDate, current.monthlyBudget);
if (pct >= warnPct && !current.alerted80) {
await recordAlert("80", warningMessage);
}
if (pct >= hardPct && !current.alerted100) {
await recordAlert("100", hardLimitMessage);
}
The alerted80 and alerted100 flags live in durable actor state. That prevents a restart or repeated hourly rollup from sending the same notification again.
When the gateway reports a new resets_at value, the actor recognizes that a new budget period has started and clears the alert guards.
With DEMO_MODE=false, alerts are sent through the Telnyx messaging binding:
await this.env.TELNYX.messages.send({
to: this.env.ADMIN_SMS_TO,
from: this.env.ADMIN_SMS_FROM,
text: body,
});
Protect gateway mutations from retries and concurrent edits
The example centralizes the AI Gateway wire contract in a GatewayClient.
Every POST, PATCH, and DELETE includes an Idempotency-Key:
headers["Idempotency-Key"] = crypto.randomUUID();
Budget updates first read the token group's current version, then send that version through If-Match:
const group = await client.getTokenGroup(tokenGroupId);
const updated = await client.patchTokenGroup(
group.id,
group.version,
{ max_budget: newBudget },
);
If another writer changes the group first, the API returns 412 precondition_failed instead of silently overwriting the newer value.
Run the offline checks first
Clone the example and run its smoke test before using live credentials:
git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/ai-gateway-tenant-meter
npm install
cp .env.example .env
npm run smoke
npm run typecheck
The smoke test verifies the gateway request shape, data response envelopes, idempotency headers, If-Match handling, spend calculations, alert guards, and finding normalization without making live API requests.
For a deployed test, configure TELNYX_API_KEY, ADMIN_SMS_FROM, and ADMIN_SMS_TO, keep DEMO_MODE=true, and deploy:
npm run deploy
curl -s https://<your-deployment>/health
Then provision a tenant:
curl -X POST https://<your-deployment>/provision \
-H "Content-Type: application/json" \
-d '{
"tenantId": "demo-tenant",
"monthlyBudget": 100
}'
Store the returned token key securely. The tenant's application uses that credential for model requests through https://llm.telnyx.com/v1.
Where this pattern fits
This architecture is useful for platforms that operate AI features for multiple customers, teams, or environments:
- SaaS products offering embedded AI assistants
- managed-service providers running customer-specific agents
- internal AI platforms with departmental budgets
- agencies managing inference workloads for clients
- development platforms separating staging and production spend
The reusable idea is straightforward: put budget enforcement at the inference gateway, and give each tenant a durable operational ledger that can explain what happened.
Top comments (0)