DEV Community

Sonam Gupta
Sonam Gupta

Posted on

Build a Multi-Tenant AI Spend Meter with Telnyx AI Gateway

Adding AI features to a multi-tenant product creates an operational question that application-level token counters do not fully answer:

How much did each tenant actually spend, and what stops the next request when the budget is gone?

Your application can estimate usage, but retries, concurrent requests, streaming behavior, and process failures can make that estimate drift away from what the inference layer served.

The open-source ai-gateway-tenant-meter example moves the control point closer to the model request. It combines Telnyx AI Gateway with one durable Edge actor per tenant.

The gateway meters and enforces. The actor remembers and reports.

What the example builds

The TypeScript application exposes a small HTTP API:

  • POST /provision creates a tenant ledger, gateway token group, and token key
  • GET /spend?tenantId=... returns current spend, model usage, guardrail findings, alerts, and budget state
  • POST /adjust changes the tenant budget with optimistic concurrency
  • GET /health confirms the worker is available

Each provisioned tenant receives:

  • a 30-day AI Gateway budget
  • an allowlist of models
  • prompt and response guardrails
  • a one-time gateway token key
  • one durable SpendLedger actor
  • private SQL tables for daily spend, alerts, and guardrail events
  • an hourly rollup task
  • SMS notifications at configurable budget thresholds

By default, the warning threshold is 80% and the hard threshold is 100%. The sample starts with DEMO_MODE=true, so alerts are logged instead of sent while you test the flow.

Provision one gateway scope per tenant

The worker routes a tenant ID to a stable actor identity. The actor then creates a gateway token group with its own budget and guardrail policy:

const group = await client.createTokenGroup(
  tenantId,
  allowedModels,
  monthlyBudget,
);

const key = await client.createTokenKey(
  `${tenantId}-primary`,
  group.id,
);
Enter fullscreen mode Exit fullscreen mode

The token group request includes the controls that should travel with that tenant:

const GUARDRAILS_POLICY = {
  secrets: {
    prompt: "block",
    response: "block",
  },
  dlp: {
    profiles: ["financial"],
    prompt: "flag",
    response: "flag",
  },
  streaming: "buffered",
};
Enter fullscreen mode Exit fullscreen mode

The tenant's application receives a one-time ltg_sk_... key and uses the OpenAI-compatible inference endpoint:

curl https://llm.telnyx.com/v1/chat/completions \
  -H "Authorization: Bearer ltg_sk_..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Kimi-K2.6",
    "messages": [
      {"role": "user", "content": "Summarize this support request."}
    ]
  }'
Enter fullscreen mode Exit fullscreen mode

Every request made with that key is associated with the tenant's token group. When the group exhausts its budget, the gateway returns 403 budget_exceeded rather than relying on an application process to notice the overage first.

Keep the operational history in a durable actor

Gateway enforcement answers whether the next inference request may run. The actor answers a different set of questions:

  • Has the 80% alert already been sent?
  • Which models generated the spend?
  • What guardrail detectors fired?
  • When does the current budget period reset?
  • Who changed the budget, and from what value?

Each tenant maps to one SpendLedger actor. That actor runs an hourly durable task and writes normalized usage into private SQL.

private armRollup(): Promise<string> {
  return this.every(
    3600,
    "rollup",
    undefined,
    { id: "hourly-rollup" },
  );
}
Enter fullscreen mode Exit fullscreen mode

The SQL update is idempotent by tenant and day:

INSERT INTO spend_days (
  tenant,
  day,
  spend,
  input_tokens,
  output_tokens,
  blocked,
  flagged
)
VALUES (?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(tenant, day) DO UPDATE SET
  spend = excluded.spend,
  input_tokens = excluded.input_tokens,
  output_tokens = excluded.output_tokens,
  blocked = excluded.blocked,
  flagged = excluded.flagged;
Enter fullscreen mode Exit fullscreen mode

If the runtime restarts, the actor's state and SQL history remain available. The next rollup can continue without rebuilding the tenant ledger from scratch.

Store guardrail findings without storing the matched text

The actor reads gateway guardrail events and stores detector metadata such as:

  • stage
  • outcome
  • detector
  • finding code
  • count

It does not store the sensitive text that triggered the finding.

That gives an operations team evidence that a secrets or financial-data rule fired without turning the ledger itself into another repository of sensitive prompt data.

Send each budget alert once per period

During the rollup, the actor compares the current spend with the tenant's configured budget.

const pct = pctOf(monthToDate, current.monthlyBudget);

if (pct >= warnPct && !current.alerted80) {
  await recordAlert("80", warningMessage);
}

if (pct >= hardPct && !current.alerted100) {
  await recordAlert("100", hardLimitMessage);
}
Enter fullscreen mode Exit fullscreen mode

The alerted80 and alerted100 flags live in durable actor state. That prevents a restart or repeated hourly rollup from sending the same notification again.

When the gateway reports a new resets_at value, the actor recognizes that a new budget period has started and clears the alert guards.

With DEMO_MODE=false, alerts are sent through the Telnyx messaging binding:

await this.env.TELNYX.messages.send({
  to: this.env.ADMIN_SMS_TO,
  from: this.env.ADMIN_SMS_FROM,
  text: body,
});
Enter fullscreen mode Exit fullscreen mode

Protect gateway mutations from retries and concurrent edits

The example centralizes the AI Gateway wire contract in a GatewayClient.

Every POST, PATCH, and DELETE includes an Idempotency-Key:

headers["Idempotency-Key"] = crypto.randomUUID();
Enter fullscreen mode Exit fullscreen mode

Budget updates first read the token group's current version, then send that version through If-Match:

const group = await client.getTokenGroup(tokenGroupId);

const updated = await client.patchTokenGroup(
  group.id,
  group.version,
  { max_budget: newBudget },
);
Enter fullscreen mode Exit fullscreen mode

If another writer changes the group first, the API returns 412 precondition_failed instead of silently overwriting the newer value.

Run the offline checks first

Clone the example and run its smoke test before using live credentials:

git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/ai-gateway-tenant-meter

npm install
cp .env.example .env
npm run smoke
npm run typecheck
Enter fullscreen mode Exit fullscreen mode

The smoke test verifies the gateway request shape, data response envelopes, idempotency headers, If-Match handling, spend calculations, alert guards, and finding normalization without making live API requests.

For a deployed test, configure TELNYX_API_KEY, ADMIN_SMS_FROM, and ADMIN_SMS_TO, keep DEMO_MODE=true, and deploy:

npm run deploy
curl -s https://<your-deployment>/health
Enter fullscreen mode Exit fullscreen mode

Then provision a tenant:

curl -X POST https://<your-deployment>/provision \
  -H "Content-Type: application/json" \
  -d '{
    "tenantId": "demo-tenant",
    "monthlyBudget": 100
  }'
Enter fullscreen mode Exit fullscreen mode

Store the returned token key securely. The tenant's application uses that credential for model requests through https://llm.telnyx.com/v1.

Where this pattern fits

This architecture is useful for platforms that operate AI features for multiple customers, teams, or environments:

  • SaaS products offering embedded AI assistants
  • managed-service providers running customer-specific agents
  • internal AI platforms with departmental budgets
  • agencies managing inference workloads for clients
  • development platforms separating staging and production spend

The reusable idea is straightforward: put budget enforcement at the inference gateway, and give each tenant a durable operational ledger that can explain what happened.

Resources

Top comments (0)