A hard spend ceiling is useful only if the meter can explain every accepted and refused request. For a multi-tenant B2B SaaS system, record one immutable usage event per billable operation, deduplicate it with a stable operation key, and derive both the live cap and the invoice evidence from that ledger. TL;DR: use provider timeseries to challenge or confirm a result, not as the sole source of truth; reconcile both sides in fixed windows and classify the delta before changing either counter.
| Control | Pick it when | Main failure mode | Dispute evidence |
|---|---|---|---|
| Pre-request reservation | Refusing traffic is preferable to exceeding the ceiling | Abandoned reservations suppress valid work | Reservation, release, and commit events |
| Post-response accounting | Small temporary overshoot is acceptable | Concurrent completions cross the ceiling | Accepted operation events plus completion time |
| Periodic timeseries polling | The upstream system is the billing authority | Polling gaps and window mismatch hide individual operations | Raw samples, query parameters, and retrieval time |
The decision axis is blunt: tighter spend control means more requests may be refused near the boundary. There is no counter design that removes that trade-off. There is, however, a design that makes the decision reviewable.
Start there.
How should API usage counters reconcile an invoice dispute?
Use a reservation counter when crossing the ceiling is worse than rejecting a valid request. Before dispatch, atomically reserve the request's maximum billable units. Commit the actual units on success, release the difference, and expire stale reservations through an explicit event. The cap then follows committed + reserved, not an eventually updated dashboard.
Pick post-response accounting when continuity matters more than a perfectly hard boundary. It is simpler, but two requests can both see remaining capacity and complete beyond it. That is expected behavior, not a mysterious billing drift. Document the tolerated overshoot in units and concurrency, then alert on that bound.
Timeseries polling belongs in reconciliation, not synchronous admission control. Samples can answer, "What did the platform count in this window?" They usually cannot prove which local attempt produced each unit unless both systems share an operation identifier and matching dimensions.
Save both.
The cap and the invoice must share the same definition of billable work. Write that definition down before comparing totals: tenant, workload, credential scope, unit, event timestamp, timezone, window boundaries, and whether refused, failed, or retried attempts count.
Make retries boring with one operation key
A retry is transport behavior. It must not silently become a second business operation. Generate an operationId before the first attempt and reuse it for every retry. Keep attemptId separate so observability still shows the retry storm.
Here is the useful diagram in words: request enters, reservation is written, attempts fan out, one terminal outcome wins, usage is committed once, and later timeseries samples are attached to the same reconciliation window.
type UsageEvent = {
tenantId: string;
workloadId: string;
operationId: string;
attemptId: string;
kind: "reserved" | "committed" | "released" | "refused";
units: number;
occurredAt: string;
};
interface UsageLedger {
appendOnce(event: UsageEvent, idempotencyKey: string): Promise<boolean>;
}
async function commitUsage(
ledger: UsageLedger,
event: Omit<UsageEvent, "kind">
): Promise<boolean> {
return ledger.appendOnce(
{ ...event, kind: "committed" },
`usage:commit:${event.tenantId}:${event.operationId}`
);
}
The idempotency key deliberately excludes attemptId. Attempts remain visible in logs and traces, while only the first commit changes billable usage. Make the uniqueness check and ledger append one atomic storage operation. A read followed by a write leaves a race.
Keep credentials out of those events. Store a stable credential fingerprint or internal credential ID for grouping, never the secret itself. Restrict access to reconciliation exports, rotate secrets through their lifecycle, and log access to secret-management operations; the OWASP Secrets Management Cheat Sheet provides the broader handling guidance.
Reconcile by window, then explain the delta
Start with immutable inputs. Save the platform's raw sample response, the exact query dimensions, the retrieval timestamp, and the local ledger watermark. Normalize both sides into half-open UTC windows such as [10:00, 10:05). Half-open boundaries ensure an event at exactly 10:05 belongs to one window.
Do not compare a dashboard screenshot with a mutable database total. Recompute local committed units from the captured ledger range. Then produce a row for every workload and window, including zeroes.
type Bucket = {
workloadId: string;
windowStart: string;
windowEnd: string;
localCommitted: number;
platformReported: number;
};
type Reconciliation = Bucket & {
delta: number;
status: "matched" | "investigate";
};
function reconcile(bucket: Bucket): Reconciliation {
const delta = bucket.platformReported - bucket.localCommitted;
return {
...bucket,
delta,
status: delta === 0 ? "matched" : "investigate",
};
}
Suppose an illustrative five-minute window contains 1,240 local committed units and 1,280 platform-reported units. The delta is +40; it is not yet "double counting." Freeze the two source snapshots before investigating, because rerunning a query after late samples arrive can change the evidence beneath the dispute. Next, test definitions. Are both values scoped to the same tenant, workload, and credential? Do they use event time, ingestion time, or completion time? Does one side round units per operation while the other rounds only the aggregate? Are refused or failed attempts billable on one side? Did a late sample revise an earlier bucket? Finally, split the local 1,240 by logical operation and by physical attempt. That view can distinguish 40 repeated attempts from 40 correctly unique operations that were attributed to the wrong workload. Keep the original +40 row even after finding the cause; add the classification and correction as new evidence so a reviewer can follow the decision without trusting a rewritten export.
Only after those checks should you join attempts by operationId. If 40 extra platform units line up with repeated attempts for already committed operations, the evidence supports a retry-accounting explanation. If they cluster at a boundary, inspect timestamp semantics. If they appear under another credential scope, fix attribution. This classification turns an argument over totals into a finite investigation.
Alert on the delta's behavior, not every nonzero point. A single late-arriving bucket may settle on the next retrieval, while a sustained same-sign delta suggests a definition or deduplication mismatch. Track reconciliation lag, unmatched units, duplicate commit rejections, expired reservations, and refused requests. Put tenant and workload in structured fields, but keep high-cardinality operation IDs in logs or traces rather than every metric label.
Test the evidence path before an invoice dispute
Replay the awkward cases. Send two concurrent attempts with one operationId; exactly one commit should change usage. Crash after the upstream operation succeeds but before the client receives the response, retry, and verify the ledger stays at one commit. Advance across a window boundary. Expire a reservation. Refuse a request at the ceiling and confirm it creates evidence without adding committed units.
Then run a shadow reconciliation during deployment. Do not let the new result enforce caps immediately. Compare old and new aggregates for complete windows, inspect classified deltas, and promote the new counter only after its event definition and boundary behavior agree with the intended policy.
The operational split should stay crisp: metrics show that drift exists, logs explain the affected dimensions, and traces connect attempts to one logical operation. An alert should link to the captured reconciliation artifact, not ask an on-call engineer to reconstruct a historical query from memory.
Limits to state plainly
A local ledger cannot prove how an external platform bills undocumented categories. A provider aggregate cannot reconstruct local intent. Clock skew, late data, sampling, aggregation, and differing failure semantics can all leave a residual delta even after retry deduplication.
Set the ceiling policy accordingly. Near the boundary, reserve conservatively if financial exposure dominates; accept bounded overshoot if refused customer traffic is the larger harm. The defensible outcome is not always a zero delta. It is a reproducible delta with preserved inputs, explicit definitions, and a named owner for resolution.
Top comments (0)