Short answer: agent execution cost is optimised in this order - meter every call at the
gateway, attribute it to a tool even when the tool forgot to report tokens, count with the
tokeniser that matches the model, then expose the numbers as a daily series, a monthly history
and a quota ring. The question to answer first is attribution: you cannot cut a cost you
cannot attribute to a caller, a tool and a month.
Key takeaways
-
Recording must never break serving.
record_usagewraps the counter increment in a try/except and logs a warning on failure - metering is important, and it is never important enough to fail a request that already produced an answer. -
Infer, do not assume zero.
infer_token_usagesupplies per-tool estimates when a module omitstoken_used: a count result is itstokens, a compression is itsorigin_tokens, and text-bearing tools fall back to a characters/4 heuristic. -
Decide billability explicitly.
maybe_record_from_toolskips failed calls, skips an explicit skiplist, and only records a non-billable tool when tokens are actually known. -
Count with the model's own tokeniser.
resolve_budget_count_modelprefers the model named in the request body and falls back to the team's effective default, so the number and the bill refer to the same tokeniser. -
Two conversions, two purposes.
estimateCostprices a request before it runs;tokensToUsdconverts tokens at the internal savings rate - keep them separate, because the number you show and the number you bill are not the same number. -
Drive the dashboard from states.
resolveSavingsHeroStatereturns one of five named states, so "no savings yet" and "savings" are different screens rather than the same screen with a zero in it. - Do this next: add a per-tool token count to your agent's logs today, and compare it with your provider invoice at the end of the week - the gap is your first optimisation target.
The short version for whoever owns the number
Most teams that ask about agent cost are actually asking two different questions. The first is
"why is the bill this high", which is an attribution question: which tool, which caller, which
retry, which month. The second is "how do we make it lower", which is a design question about
context, retrieval and model choice. The answer to the first question is the input to the
second, and it is the half that is usually missing.
The architecture in this article is the metering half: one recording path that cannot break a
request, an inference layer for the tools that do not report tokens, a tokeniser decision, and
a set of read models - a daily series, a six-month history with its caps, a quota display and a
savings state machine. Everything downstream, including what a customer is charged, is derived
from those numbers.
Why attribution comes before optimisation
Three facts make the ordering non-negotiable. Token counts are model-specific, so a count made
with the wrong encoder is a different number, not an approximation
(tiktoken). Counters live in the request path and must
be shared, atomic and cheap, which is exactly what Redis counters are for
(INCRBY). And the only component that sees
every call, whatever client made it, is the gateway - the same place where MCP tool
definitions are declared (MCP tools).
Put those together and the design writes itself: record in the gateway, count with the right
encoder, store in a shared counter keyed by team and month, and expose read models that make
the spend legible. The sections below are the shipped implementation of each step.
The counting, caching and metering decisions behind these numbers are laid out step by step in
Build a Low-Cost AI Backend Architecture.
record_usage: the recording path must not throw
# backend/smartgate/core/record_usage.py — source lines 58–76 (record_usage)
async def record_usage(team_id: str, token_used: int) -> None:
"""Increment monthly usage counter for *team_id*."""
if not team_id or token_used <= 0:
return
try:
guard = BudgetGuardModule()
await guard.process(
ToolContext(team_id=team_id),
action="record",
team_id=team_id,
tokens=token_used,
)
except Exception as exc:
logger.warning(
"record_usage failed for team %s (%s tokens): %s",
team_id,
token_used,
exc,
)
The shape of this function is the lesson: a guard for empty input, a single call into the
budget guard's record action, and a try/except that logs a warning instead of propagating.
Metering runs after the work is done; if the counter store is briefly unavailable, the request
that already produced an answer should still return. The cost of that choice is a small
undercount during an incident - which is the right failure direction, and worth stating out
loud in your own runbook.
infer_token_usage: a per-tool estimate beats a zero
# backend/smartgate/core/record_usage.py — source lines 29–55 (infer_token_usage)
def infer_token_usage(tool: str, data: Any, params: dict | None = None) -> int:
"""Best-effort token estimate when modules omit result.meta.token_used."""
if not isinstance(data, dict):
return 0
if tool == "budget_count":
return int(data.get("tokens") or 0)
if tool == "compress":
return int(data.get("origin_tokens") or 0)
if tool == "fetch":
content = data.get("markdown") or data.get("content") or ""
if isinstance(content, str) and content:
return max(1, len(content) // 4)
if tool == "search":
results = data.get("results")
if isinstance(results, list):
chars = sum(
len(str(r.get("content") or r.get("snippet") or ""))
for r in results
if isinstance(r, dict)
)
if chars:
return max(1, chars // 4)
if tool == "dedup":
text = (params or {}).get("text") or data.get("text") or ""
if isinstance(text, str) and text:
return max(1, len(text) // 4)
return 0
Tools that compute their own consumption report token_used; the rest are estimated here, and
the per-tool rules are the content. A budget_count result is its own tokens field. A
compress result is priced on origin_tokens, because compression is billed on what it
processed. fetch and search estimate from returned text at roughly four characters per
token, and dedup from the text it was handed. Two properties make this honest: the estimate
is always at least one token when content exists (so a call is never recorded as free), and the
function returns zero for anything it cannot reason about rather than inventing a number.
maybe_record_from_tool: one place decides what is billable
# backend/smartgate/core/record_usage.py — source lines 79–97 (maybe_record_from_tool)
async def maybe_record_from_tool(
team_id: str,
tool: str,
*,
success: bool,
token_used: int = 0,
data: Any = None,
params: dict | None = None,
) -> None:
"""Record billable tool usage when successful and tokens are known or inferable."""
if not success or not team_id or tool in _SKIP_RECORD_TOOLS:
return
if tool not in _BILLABLE_TOOLS and token_used <= 0:
return
amount = token_used
if amount <= 0:
amount = infer_token_usage(tool, data, params)
if amount > 0:
await record_usage(team_id, amount)
Billability is a product decision, so it lives in one function with explicit rules. Failed
calls are not recorded. A skiplist of tools is excluded outright. A tool that is not on the
billable list is only recorded when its token count is known (token_used > 0) - which means a
new tool cannot start billing by accident just because it returned some data. When it does
record, it prefers the reported count and falls back to the inference above. Read this function
as the answer to "what exactly are we charging for", because that is what it is.
resolve_budget_count_model: the count must match the bill
# backend/smartgate/core/effective_settings.py — source lines 77–82 (resolve_budget_count_model)
def resolve_budget_count_model(request_state: Any, body_model: str | None) -> str:
if body_model:
return body_model
es = getattr(request_state, "effective_settings", None) or {}
defaults = es.get("defaults") or {}
return str(defaults.get("budget_count_model", "deepseek-chat"))
Two tokenisers can disagree by a few percent on the same text, which is enough to make a limit
and a usage report argue with each other. This resolver fixes the ambiguity: the model named in
the request body wins, and otherwise the team's effective settings decide, defaulting to the
configured budget model. If you self-host or integrate, this is the single place where "which
tokeniser counts" is answered - make sure it matches the model actually being called, not the
model in your marketing copy.
estimateCost: price a request before it runs
# app/api/smartgate/playground/run/route.ts — source lines 71–73 (estimateCost)
function estimateCost(tokens: number): number {
return tokens * COST_PER_TOKEN;
}
One multiplication against a per-token constant, used by the playground to show what a request
will cost. The value here is not the maths but the placement: an estimate available before the
call turns "how expensive is this prompt" from a review question into a loop - edit, see the
number, edit again. Keep the constant in configuration next to the model list, so a price
change is not a code change.
tokensToUsd: the conversion used for savings
# config/billing-savings-share.ts — source lines 46–48 (tokensToUsd)
function tokensToUsd(tokens: number): number {
return (tokens / 1_000_000) * INTERNAL_SAVINGS_USD_PER_M;
}
The second conversion converts tokens into USD at the internal savings rate per million. Two
things are worth noticing. First, it divides by a million before multiplying, so the rate reads
as a per-million price - the unit humans quote. Second, this is a savings conversion, not a
billing one: it exists to price avoided tokens consistently across the dashboards and the
monthly fee calculation. Keeping the estimate constant and the savings rate as named,
versioned values is what lets an old number be re-derived later.
buildDailyUsageSeries: a 30-day series with zero-filled days
# lib/analytics/daily-usage-series.ts — source lines 17–47 (buildDailyUsageSeries)
function buildDailyUsageSeries(
dailyUsage: DailyUsageRow[],
days = 30,
metric: DailyMetric = "tokens",
): { date: string; value: number }[] {
const byDay = new Map<string, number>();
for (const row of dailyUsage) {
const key = row.date.includes("T")
? row.date.slice(0, 10)
: row.date.slice(0, 10);
const amount =
metric === "calls" ? (row.calls ?? 0) : row.tokens;
byDay.set(key, (byDay.get(key) ?? 0) + amount);
}
const out: { date: string; value: number }[] = [];
const end = new Date();
end.setHours(0, 0, 0, 0);
for (let i = days - 1; i >= 0; i--) {
const d = new Date(end);
d.setDate(d.getDate() - i);
const key = localDateKey(d);
out.push({
date: key.slice(5),
value: byDay.get(key) ?? 0,
});
}
return out;
}
Charts fail in a specific way: missing days are drawn as gaps or as interpolated lines, and
both mislead. This builder aggregates rows into a map by date, then walks backwards over the
requested window (30 days by default) emitting one entry per day, defaulting to zero. The
metric switch (tokens or calls) is decided per row, so the same series serves both charts.
Fill the gaps in the data layer, not in the chart library.
buildBudgetHistory: six months, both caps
# lib/analytics/budget-history.ts — source lines 31–68 (buildBudgetHistory)
async function buildBudgetHistory(
teamId: string,
monthsCount = 6,
): Promise<BudgetHistoryPayload> {
const team = await prisma.team.findUniqueOrThrow({
where: { id: teamId },
select: {
monthlyTokenLimit: true,
l2BudgetUnit: true,
l2BudgetValue: true,
},
});
const l1Cap = Number(team.monthlyTokenLimit);
const l2Unit = team.l2BudgetUnit;
const l2Cap =
l2Unit === "tokens" && team.l2BudgetValue != null
? Number(team.l2BudgetValue)
: null;
const redis = getRedis();
const now = new Date();
const months: BudgetHistoryMonth[] = [];
for (let i = monthsCount - 1; i >= 0; i--) {
const d = shiftMonth(now, -i);
const yyyyMm = d.toISOString().slice(0, 7);
const usage =
(await redis.get<number>(usageMonthKey(teamId, d))) ?? 0;
months.push({
month: yyyyMm,
label: monthLabel(yyyyMm),
usage,
});
}
return { months, l1Cap, l2Cap, l2Unit };
}
The history model joins two things: the per-month usage read from the same monthly key the
meter writes, and the team's two caps - the L1 monthly token limit and, where the unit is
tokens, the L2 budget value. Reading the usage for the previous months is a straightforward
GET per month precisely because the key shape was designed to be read that way. If you build
this yourself, keep the caps in the payload: a history chart without the cap invites the reader
to guess what "high" means.
computeQuotaRingDisplay: one number, two stories
# lib/analytics/quota-ring-display.ts — source lines 9–19 (computeQuotaRingDisplay)
function computeQuotaRingDisplay(
used: number,
cap: number,
display: QuotaRingDisplay = "used",
) {
const usedPct = cap > 0 ? Math.min(100, Math.round((used / cap) * 100)) : 0;
const centerPct =
display === "used" ? usedPct : Math.max(0, 100 - usedPct);
const label = display === "used" ? "used" : "left";
return { usedPct, centerPct, label, strokeColor: strokeColor(usedPct) };
}
The quota ring shows either "used" or "left", and the function makes that explicit: the
percentage is computed once, the centre value flips depending on the display mode, and the
label flips with it, while the stroke colour always follows the used percentage. That is a
small piece of code doing an important job - "12% left" and "88% used" describe the same state
and produce very different behaviour from the person reading them.
resolveSavingsHeroState: five states beat one number
# lib/analytics/savings-display-state.ts — source lines 16–26 (resolveSavingsHeroState)
function resolveSavingsHeroState(input: SavingsHeroInput): SavingsHeroState {
const { totalCalls, autDisplayed, uMonth, operationalCalls } = input;
if (totalCalls === 0 && uMonth === 0) {
return operationalCalls > 0 ? "s2_operational" : "s1_cold";
}
if (autDisplayed > 0) return "s3_savings";
if (uMonth > 0 && autDisplayed === 0 && totalCalls > 0) {
return "s4_usage_no_savings";
}
return "s5_zero";
}
The hero area of a cost dashboard has to answer a question that is not "what is the number":
is anything happening at all? The state machine returns one of five values - cold (nothing
yet), operational (calls but no usage), savings (something was avoided), usage-without-savings
(you are spending and nothing is being saved), and zero. Naming the states means the UI can
say "you are spending, and no tokens have been avoided yet" instead of rendering a zero, which
is the single most useful thing a cost dashboard can tell a new user.
runMonthlySavingsShareCharges: batch with a receipt
# lib/billing/savings-share-charge.ts — source lines 59–92 (runMonthlySavingsShareCharges)
async function runMonthlySavingsShareCharges(yyyyMm: string): Promise<{
processed: number;
charged: number;
skipped: number;
}> {
const teams = await prisma.team.findMany({
where: {
plan: { in: ["PRO", "TEAMS"] },
externalSubscriptionId: { not: null },
subscriptionStatus: "active",
},
select: { id: true },
});
let charged = 0;
let skipped = 0;
for (const team of teams) {
const owner = await prisma.teamMember.findFirst({
where: { teamId: team.id, role: "OWNER" },
select: { userId: true },
});
if (!owner) {
skipped += 1;
continue;
}
const result = await chargeTeamSavingsShare(team.id, owner.userId, yyyyMm);
if (result.charged) charged += 1;
else skipped += 1;
}
return { processed: teams.length, charged, skipped };
}
The monthly run selects teams on a paid plan with an active external subscription, finds the
owner, and delegates each charge to the per-team function, counting charged and skipped.
The counts are the point: a batch that reports processed: 240, charged: 61, skipped: 179 is
auditable, and the 179 have reasons attached to them by the per-team call. Run it as a job,
store the result, and treat "skipped" as a category to review rather than noise.
enrich_compress_audit_params: make the saving auditable
# backend/smartgate/core/audit_params.py — source lines 4–15 (enrich_compress_audit_params)
def enrich_compress_audit_params(
params: dict,
result_data: object,
) -> dict:
"""Add raw_tokens / compressed_tokens for savings aggregation (spec §4.1)."""
if not isinstance(result_data, dict):
return params
origin = int(result_data.get("origin_tokens") or 0)
compressed = int(result_data.get("compressed_tokens") or 0)
if origin > 0:
params = {**params, "raw_tokens": origin, "compressed_tokens": compressed}
return params
The audit record is where a savings claim becomes evidence. This function adds the compression
pair - raw_tokens and compressed_tokens - to the audit parameters whenever the result
carries an original token count. It only writes when origin_tokens > 0, so a non-compression
call does not gain misleading fields. With both numbers in the audit row, the aggregate saving
is a sum over records rather than a statistic someone remembered, and a customer dispute can be
answered call by call.
How SmartGate compares
Cost tooling splits by where the number comes from.
| Where the number comes from | Attribution | What you pay | |
|---|---|---|---|
| Provider invoice | Monthly total per API key | Per key, at best | Nothing extra - and no per-tool or per-caller detail |
| Client-side instrumentation | Your own logging in each caller | As good as your tagging discipline | Engineering time, and everything breaks when a caller bypasses the SDK |
| FinOps platform | Exported usage and tags | Strong, after the fact | Subscription per seat or per spend |
| Gateway metering (SmartGate) | The request path: every call, every tool, per team and month | Per tool, per caller, per month, with audit rows | Free tier: 2M tokens/mo, all seven tools, no card; Pro from $18/mo, with a savings share only after measured savings pass a threshold |
The pricing shape is what makes the metering credible: the same token counts that drive your
dashboard drive the vendor's fee, so an inflated savings number would inflate the bill. Compare
that against your provider invoice on the pricing page.
How to get started
- Record at the gateway, never throw. One function, one counter, a warning on failure.
- Estimate the tools that stay silent. Per-tool rules, a characters/4 fallback, and zero only when nothing can be reasoned about.
- Decide billability out loud. A skiplist plus "billable tools may record without a known count, others may not" is a policy you can review.
- Fix the tokeniser. One resolver decides which model counts, and it follows the model actually being called.
- Publish the read models. Daily series with zero-filled days, six-month history with both caps, a quota ring that can show either side, and named dashboard states.
- Keep the audit trail. Raw and compressed tokens on every compression record, and a monthly batch that reports what it charged and what it skipped.
Start on the free tier and watch the usage page populate - 2M tokens a month, all seven
smart_* tools, no card: start free.
The tool surface is documented in the docs, and
contract-sized deployments start with the contact form.
FAQ
Why not let every tool report its own token usage?
Tools that can measure, do; the rest would report wrong numbers or none. A single inference
layer with per-tool rules keeps the estimate consistent, visible and testable, instead of
spread across a dozen call sites with different heuristics.
Do estimates make the usage numbers untrustworthy?
They make them estimates, which is why the reported count always wins when it exists. The
practical consequence is that usage is a slight undercount for tools that never self-report -
a stable bias you can compare against a provider invoice, rather than random noise.
How do I find out which tool is driving the bill?
Attribute by tool and by month. infer_token_usage exists so that every tool has a number, and
the monthly keys make the window explicit. Sort by tool for one month and you have your first
optimisation target.
What is the difference between estimateCost and tokensToUsd?
Purpose and rate. estimateCost prices a request against the model's per-token constant before
it runs, so it is a planning tool; tokensToUsd converts tokens at the internal savings rate
for reporting and savings-share maths. Mixing them produces a number that is neither a quote
nor a saving.
Should a failed request be recorded?
No - maybe_record_from_tool skips failures. Charging for a request that produced an error is
the fastest way to lose trust in the whole dashboard, and the skipped volume is a useful
reliability signal in its own right.
How do we show a customer that the gateway actually saves them money?
Show the arithmetic: raw and compressed tokens per call, aggregated avoided tokens, the host
model equivalent, and a price-table version. The audit row is the unit of evidence, and every
number above it is a sum you can re-derive.
Limitations and what this does not do
- Estimates have a bias, not an error bar. The characters/4 rule is a reasonable average for prose and a poor one for code or CJK text. Tools that matter should self-report.
- Usage metering is not a cost model. It counts tokens; it does not know your provider contract, discounts, cached-input pricing or regional rates. Treat converted USD values as internal, consistent and versioned - not as your invoice.
-
A fail-soft recorder undercounts during incidents.
record_usagelogs and returns. If your counters are unreachable for an hour, that hour is under-metered; the log line is where you find out. - The monthly batch charges; it does not reconcile. Failures are counted as skips with reasons and need a human to look at them before the month is closed.
-
Dashboard states encode product judgement. Five named states are one reasonable
decomposition; a different product will name them differently, and the names are the contract
with the UI, not a law.Sources
tiktoken — model-specific tokenisation used for counting: https://github.com/openai/tiktoken
Redis — INCRBY, the atomic counter behind usage keys: https://redis.io/docs/latest/commands/incrby/
Redis — EXPIRE, keeping counters self-cleaning: https://redis.io/docs/latest/commands/expire/
Model Context Protocol — tool definitions and descriptions: https://modelcontextprotocol.io/specification/2026-07-28/server/tools
Redis — key patterns and expiry practices: https://redis.io/docs/latest/develop/use/patterns/
SmartGate — documentation, pricing and sales contact: https://smartgate.network/docs · https://smartgate.network/pricing · https://smartgate.network/contact
Method note
The code in this article is not transcribed. Each block was cut directly out of the slice body
returned by the SmartGate slice API and re-asserted byte-for-byte as a substring of that body
before publication; the first line inside every fence records the file and the exact source
lines. Symbols were pinned by whole-name containment (rule A level 2) and confirmed by the
service's slot-proof endpoint before being written into the prose - 12 of 12 planned sections
pinned, no abstentions.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|---|---|---|---|---|---|
| 1 | record_usage record usage tokens for agent execution cost | record_usage |
backend/smartgate/core/record_usage.py |
58–76 | rule A L2 → slot-proof | 31014a2528fe |
| 2 | infer_token_usage infer token usage per tool for cost accounting | infer_token_usage |
backend/smartgate/core/record_usage.py |
29–55 | rule A L2 → slot-proof | c95fde0433db |
| 3 | maybe_record_from_tool maybe record from tool for agent cost tracking | maybe_record_from_tool |
backend/smartgate/core/record_usage.py |
79–97 | rule A L2 → slot-proof | 8a77a10cb0a5 |
| 4 | resolve_budget_count_model resolve budget count model for token counting | resolve_budget_count_model |
backend/smartgate/core/effective_settings.py |
77–82 | rule A L2 → slot-proof | 7194cb800a16 |
| 5 | estimateCost estimate cost per request in the playground | estimateCost |
app/api/smartgate/playground/run/route.ts |
71–73 | rule A L2 → slot-proof | 31e82a4b4236 |
| 6 | tokensToUsd tokens to usd conversion for agent execution cost | tokensToUsd |
config/billing-savings-share.ts |
46–48 | rule A L2 → slot-proof | eae4fd82916a |
| 7 | buildDailyUsageSeries build daily usage series for cost trending | buildDailyUsageSeries |
lib/analytics/daily-usage-series.ts |
17–47 | rule A L2 → slot-proof | 8a01ebf2b133 |
| 8 | buildBudgetHistory build budget history for agent spend | buildBudgetHistory |
lib/analytics/budget-history.ts |
31–68 | rule A L2 → slot-proof | 7e679ee19b3a |
| 9 | computeQuotaRingDisplay compute quota ring display for usage dashboard | computeQuotaRingDisplay |
lib/analytics/quota-ring-display.ts |
9–19 | rule A L2 → slot-proof | f102407a5e5e |
| 10 | resolveSavingsHeroState resolve savings hero state for the cost dashboard | resolveSavingsHeroState |
lib/analytics/savings-display-state.ts |
16–26 | rule A L2 → slot-proof | 45a96b447fe8 |
| 11 | runMonthlySavingsShareCharges run monthly savings share charges | runMonthlySavingsShareCharges |
lib/billing/savings-share-charge.ts |
59–92 | rule A L2 → slot-proof | 966f323c15ca |
| 12 | enrich_compress_audit_params enrich compress audit params with token savings | enrich_compress_audit_params |
backend/smartgate/core/audit_params.py |
4–15 | rule A L2 → slot-proof | 952380769807 |
Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before
publication. 12 of 12 sections pinned, 0 abstentions, 0 misses.
Top comments (0)