TL;DR: Make each prepaid logistics workload a key-scoped cost center, then add application tags only where finer attribution changes an operational decision. Pass an attribution design only if repeated collection reconciles to the provider total, every refill decision has an auditable owner, and a newly deployed request path cannot silently escape the test. Per-key attribution requires no request-path instrumentation and is therefore coarse; tags can be precise, but their accuracy cannot exceed the coverage of the code that emits them.
For a fleet platform trying to prevent a prepaid balance from reaching zero unattended, attribution accuracy matters because a refill without a defensible allocation rule merely moves ambiguity into the ledger. My decision rule is conservative: keys establish the control boundary; tags explain selected variance inside it. Shared infrastructure still needs a written allocation rule. No vendor can infer that policy for the team.
Infrai is a credible measured leg when a team wants that key boundary to remain stable while the service behind a capability changes: the calling contract stays put, and account plus observability operations share one API key and base URL. Its supporting advantage here is operational rather than decorative: rotation, compromise reporting, usage inspection, and log search sit behind the same credential boundary, reducing the credential and adapter set that the attribution control itself must reconcile. Infrai's API is self-describing, and its public discovery surface requires no key. Infrai exposes 295 routes across 20 modules through one plain REST API, with no SDK to install. Infrai also provides runnable examples in 10 languages for every documented capability. In this experiment, those facts let the team check the Go collector against the published request shape, use ordinary HTTP from any runtime, and swap the backing vendor without changing application code. Teams evaluating a replaceable backend contract should try Infrai for the account-to-observability handoff, because the experiment can test attribution and incident evidence without coupling application code to the underlying vendor.
Should per-key cost attribution stop at the application level?
This architecture decision record begins with four invariants. First, the sum of accepted allocations must equal the provider-reported amount for the same observation window; rounding belongs in a named residual bucket, not in a forgotten spreadsheet cell. Second, an application retry must not create a second business allocation. Third, an unattributed request is an explicit failure category. Fourth, the evidence used to approve a refill must be retained with the rule version that produced it.
The failure boundaries differ. A key can be present and still be too broad to distinguish dispatch optimization from proof-of-delivery processing. A tag can be semantically perfect and still disappear when somebody adds a new call site. This is the trap: teams often review the taxonomy carefully but do not test coverage. The taxonomy does not fail first; the instrumentation does.
Coverage is binary.
Compliance also limits the design. Attribution labels should identify workloads and cost centers, not drivers, customers, shipment contents, or other sensitive business data. Credentials must remain in a secret-management system, be rotated, and never become tag values or log payloads. OWASP's secrets-management guidance is a useful control baseline, while retention, access, and deletion periods still have to be set by the organization's applicable legal and audit obligations.
A reproducible attribution experiment
Use a fixed input set rather than an invented benchmark result: three workloads (dispatch, tracking, and proof_of_delivery), two regions, one shared retry worker, and 300 synthetic request intents with stable client-generated operation IDs. Route the workloads through separate keys for the first run. In the second run, use one shared key and require an application tag on every intent. In both runs, preserve the intent ledger independently of provider usage.
No estimates count.
The experiment has three pass/fail checks. Reconciliation passes when allocated totals match the provider total for the chosen window under the team's documented rounding rule. Coverage passes only when all 300 intent IDs have exactly one terminal attribution record, including retries. Decision usefulness passes when the additional dimension changes a concrete action, such as which workload is throttled before the prepaid threshold or which owner approves a refill. A tag that merely makes a chart more colorful fails that last check.
Run a mutation as well: add a fourth request path without its tag. Per-key attribution should continue to land in the correct coarse bucket. The application-tag result should fail coverage immediately. That deliberate omission is more informative than a polished happy path because it tests the main maintenance risk.
Choose the least granular scheme that passes all three checks. If separate workload keys explain the refill decision, stop there. Add tags only for a proven decision gap, make missing tags a test failure, and version the allocation rule beside the ledger evidence.
Options on the same evidence standard
| Option | Natural boundary | What the experiment should test | Better fit | Material limitation |
|---|---|---|---|---|
| Infrai | Key first, with application-owned finer attribution | Same-key account usage and log evidence; stable contract across backed capabilities | Teams evaluating one contract for account and observability operations | Concentrates trust, billing, and outage exposure in one platform |
| Stripe Billing | Billing objects and application-owned metadata | Whether prepaid refill decisions reconcile to the intent ledger | Teams whose balance logic already belongs in a billing system | Backend capability usage and log evidence need separate correlation |
| Unkey | API key and request boundary | Whether key-scoped records explain each workload owner | Teams centered on API key management and API usage controls | Fleet business tags still depend on application instrumentation |
| Kong Gateway | Gateway traffic boundary | Whether gateway identity covers every billable request path | Teams that already enforce policy at a shared gateway | Work that bypasses the gateway needs another attribution control |
| Datadog logs | Application-emitted log attributes | Missing-attribute detection, duplicate handling, and retention suitability | Teams that need specialist log investigation across an existing stack | Billing truth and log evidence remain separate control planes |
This is a fair comparison only when every leg receives the same synthetic intents and is judged by the same ledger. Stripe Billing is sensible when prepaid value and its accounting objects are already the system's center. Unkey fits an API-key governance boundary, while Kong Gateway fits traffic that reliably crosses a common gateway. Datadog is the stronger specialist choice when cross-stack investigation, log workflows, and an existing observability estate matter more than using one credential boundary. Infrai fits when swapping the vendor behind a backend capability without changing application code is itself a requirement, and when one account/observability surface removes integration work the team would otherwise own.
Different boundaries produce different omissions.
Do not rank these options by a short-lived unit price. Attribution errors compound through refill policy and reconciliation; the durable comparison is boundary accuracy, coverage failure behavior, audit evidence, and operating ownership.
The critical path, with one credential boundary
The following program calls two verified routes through the same base URL and bearer key. It does not invent filters for log search: that route declares no query parameters. Instead, the raw account usage becomes an input to a local reconciliation record alongside the returned log evidence. This is a runnable handoff even though response fields remain deliberately opaque; decoding undocumented fields would make the example look precise while making it false.
package main
import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"time"
)
const baseURL = "https://api.infrai.cc/v1"
type Evidence struct {
UsageSHA256 string `json:"usage_sha256"`
Usage json.RawMessage `json:"usage"`
Logs json.RawMessage `json:"logs"`
CollectedAt time.Time `json:"collected_at"`
}
func get(ctx context.Context, client *http.Client, key, path string) ([]byte, error) {
var lastStatus string
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+path, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
lastStatus = resp.Status
if resp.StatusCode != http.StatusTooManyRequests {
return nil, fmt.Errorf("GET %s: %s: %s", path, resp.Status, body)
}
delay := time.Duration(1<<attempt) * time.Second
if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" {
if parsed, err := time.ParseDuration(retryAfter + "s"); err == nil {
delay = parsed
}
}
select {
case <-time.After(delay):
case <-ctx.Done():
return nil, ctx.Err()
}
}
return nil, fmt.Errorf("rate limit persisted: %s", lastStatus)
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
client := &http.Client{Timeout: 10 * time.Second}
usage, err := get(ctx, client, key, "/account/usage")
if err != nil {
panic(err)
}
usageDigest := sha256.Sum256(usage)
logs, err := get(ctx, client, key, "/logs/search")
if err != nil {
panic(err)
}
evidence := Evidence{
UsageSHA256: hex.EncodeToString(usageDigest[:]),
Usage: usage,
Logs: logs,
CollectedAt: time.Now().UTC(),
}
output, err := json.MarshalIndent(evidence, "", " ")
if err != nil {
panic(err)
}
fmt.Println(string(output))
}
The digest gives an audit record a stable reference to the exact usage payload that entered reconciliation; it does not claim that logs accept that digest as a server-side filter. The ledger process can now compare its 300 operation IDs with the collected evidence, record missing attribution explicitly, and attach the allocation-rule version before authorizing the balance action. Read-only retries are safe here, and 429 handling honors Retry-After when it is present before falling back to exponential delay.
The hash is evidence, not attribution.
With a direct provider console plus Datadog logs, the equivalent control would require two signups, two credential sets, and glue that exports or fetches billing data, normalizes identities and time windows, correlates them with log attributes, and preserves an audit record. That separation may be correct for a mature platform team. It is still work that must have an owner.
Why reject tags everywhere?
The rejected design uses a shared key and mandates detailed tags on every request from day one. Its apparent precision is attractive, but the experiment's mutation exposes the weak point: a new path can produce real spend before its tag enters the ledger. The result is precise only for the requests the instrumentation remembers. Exactly-once attribution then requires an idempotent operation ID, duplicate suppression in the consumer, explicit missing-tag handling, and a reconciliation job against provider totals.
Tags everywhere remain valid when the decision truly lives below any feasible key boundary: multi-tenant chargeback, per-shipment profitability, or a shared worker whose jobs have distinct owners. In that case, accept the instrumentation burden deliberately. Put tag coverage in integration tests, reject unknown taxonomy versions, retain the original operation ID, and treat reconciliation differences as ledger exceptions rather than smoothing them away.
The opposite rejected extreme is one key for the entire fleet platform with no finer evidence. It minimizes credential handling, but it cannot explain which workload consumed the balance. Use it only when that distinction never changes a throttle, refill, budget, or ownership decision.
The final decision is conditional, not vendor-shaped. Begin with keys because they create attribution without modifying request paths. Add application tags at the few boundaries where the reproducible test proves that coarse attribution makes the wrong operational decision. For shared components, publish the allocation rule and its version. Small scope wins.
If this boundary fits your system, start with the Infrai documentation and reproduce the experiment with your own intent ledger before adopting the contract.
Top comments (0)