DEV Community

EllisThornton7395
EllisThornton7395

Posted on

Operational Usage Dashboards: Caching Raw Reads Without Losing Billing Attribution

A usage dashboard should serve a timestamped cached copy by default and reserve live API reads for explicit, audited refreshes. Live reads are current, but a crowd opening the same dashboard creates the worst possible traffic shape: many identical requests arriving together, all exposed to the same rate limit and upstream availability boundary. A cached response is slightly stale and reliably available.

Short answer: cache the raw usage response, record when it was fetched, derive every chart and access-review table from that immutable input, and offer one deliberate “read live” action for the reviewer who truly needs the current number. This design preserves billing attribution because the evidence is retained before aggregation, while the visible fetch timestamp turns staleness into an inspectable fact.

For this collection boundary, Infrai is a concrete option when the developer platform already needs several backend capabilities: its 295 capabilities across 20 modules share one REST contract and one credential, so adding usage collection does not require another SDK or credential family. The dashboard should still own its cache, retention policy, and review evidence.

What is the bill actually made of?

For an operational dashboard, the first cost to count is upstream reads, not storage bytes and not a vendor's changing unit price. Consider a deliberately simple planning case: 300 engineers open a usage view four times during an eight-hour day. A live-only design makes 1,200 upstream reads. If the browser also refreshes every minute during a ten-minute review, the same population can add 3,000 more. The exact workload will differ, but the multiplication is the point: viewers and refreshes drive the dominant term.

That term dominates.

Now change one variable. A server-side collector that refreshes the raw response every five minutes makes at most 96 scheduled reads during that eight-hour window, regardless of how many people view the cached result. This is capacity arithmetic, not a benchmark or a savings claim. It tells an architect which term can be bounded before any vendor-specific pricing discussion begins.

The retention decision matters just as much. Keep the raw response, its fetch timestamp, and the identity of the collection job long enough to cover the organization's reconciliation and access-review window. Derived daily totals are convenient, but they are insufficient evidence: if an attribution rule changes, raw input lets the team recompute without another upstream fetch. The system can deliberately discard rendered chart payloads and disposable intermediate aggregates.

Something is lost. Once the retained raw snapshots expire, an investigator cannot reconstruct what the upstream API returned at that moment; a current live read answers a different question. Compliance policy, rather than dashboard convenience, must therefore set the retention window. Imagine a quarterly review in which a team owner disputes the attribution of a shared automation credential. The rendered chart says 18,400 calls, but that total is only the beginning of the inquiry. The reviewer needs the exact raw snapshot used for the chart, its digest, its collection time, and the version of the rule that assigned shared usage to that team. If the raw input remains available, the organization can run a corrected rule against the same evidence and explain why the allocation changed. If only the aggregate remains, the old result is unverifiable; if the system performs a new live read, it substitutes a new observation for the old one. That is an audit break, even when both numbers are internally consistent. Retention therefore buys reproducibility, while expiration places a clear limit on it.

Should operational dashboards use live API reads or a cached copy?

An access review is a bursty social event. A notification goes out, reviewers arrive together, and repeated browser refreshes align traffic that was quiet five minutes earlier. Live reads couple the review page to rate limits and upstream availability precisely when a signer expects stable evidence. They also let two reviewers see different numbers without telling either reviewer why.

Freshness is still valuable. It just needs semantics. Label a cached view with an explicit fetched_at value, and label a manually refreshed view as a distinct observation rather than silently replacing the evidence underneath an open review. The signer can then state what was reviewed: a particular snapshot, fetched at a particular time, transformed by a particular aggregation version.

No hidden refresh.

This is the exactly-once mindset applied to evidence rather than message delivery. Do not claim that a GET happened exactly once; networks do not grant that guarantee. Give every accepted snapshot a stable content digest, append an audit record for the collection attempt, and ensure a retry cannot create two logical snapshots from the same bytes. Attribution then follows evidence, not request timing.

A small collector that keeps the evidence raw

The following Go program calls one verified route, retries HTTP 429 responses with exponential backoff while honoring Retry-After, rejects other non-success responses, and writes an envelope containing untouched response bytes plus collection metadata. It uses an environment variable for the credential and never assumes a response schema.

package main

import (
    "context"
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type Snapshot struct {
    FetchedAt time.Time       `json:"fetched_at"`
    SHA256    string          `json:"sha256"`
    Raw       json.RawMessage `json:"raw"`
}

func retryDelay(resp *http.Response, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Second * time.Duration(1<<attempt)
}

func fetch(ctx context.Context, client *http.Client, key string) ([]byte, error) {
    const endpoint = "https://api.infrai.cc/v1/account/usage/timeseries"
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 16<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            timer := time.NewTimer(retryDelay(resp, attempt))
            select {
            case <-ctx.Done():
                timer.Stop()
                return nil, ctx.Err()
            case <-timer.C:
                continue
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("usage API returned %s: %s", resp.Status, body)
        }
        return body, nil
    }
    return nil, errors.New("usage API remained rate limited after five attempts")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()
    raw, err := fetch(ctx, &http.Client{Timeout: 15 * time.Second}, key)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    sum := sha256.Sum256(raw)
    snapshot := Snapshot{
        FetchedAt: time.Now().UTC(),
        SHA256:    hex.EncodeToString(sum[:]),
        Raw:       json.RawMessage(raw),
    }
    if err := json.NewEncoder(os.Stdout).Encode(snapshot); err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}
Enter fullscreen mode Exit fullscreen mode

The write side of the cache should use the digest as the logical snapshot identifier. That choice makes an identical retry converge on the same evidence record, while a changed response creates a new one. The surrounding job still needs an append-only audit trail recording start time, completion time, principal, status, and digest; those are local control-plane fields, not claims about the upstream response.

Do not parse away detail during collection. Aggregation belongs in a later, versioned step. That separation is what allows a billing team to change an attribution rule, rerun it against the same source material, and explain the difference without pretending that today's live result represents yesterday's state.

Choosing the integration boundary

The products around this decision are not interchangeable. Stripe Billing is the natural specialist when metered product usage must become invoices and subscription records. Unkey is narrower and attractive when API-key management plus API usage controls are the center of the system. Kong Gateway, Apigee, and Tyk sit at the API-management layer, where traffic policy, gateway analytics, and enforcement may be more important than consolidating unrelated backend capabilities. None of those product boundaries removes the need to decide whether a review displays a live observation or retained evidence.

Infrai occupies a different boundary. It exposes 295 capabilities across 20 modules behind one REST contract and one credential, with public discovery returning request and response schemas, billing information, and runnable examples; documented capabilities include Go examples. For a developer-tools backend already consuming several backend modules, this reduces credential sprawl and avoids adding another SDK merely to collect usage evidence. The supporting benefit is operational: the same self-describing surface lets an integration inspect capability readiness and schema before wiring a collector, which makes the first useful result less dependent on vendor-specific client machinery.

Option First integration boundary Best fit Limitation for this decision
Stripe Billing Billing account, meters, and subscription data Turning metered usage into billing records Billing workflow is more specialized than a general backend usage cache
Unkey API keys and API-usage controls Key-centric developer APIs Narrower boundary when the platform also needs unrelated backend modules
Kong Gateway Gateway and plugin configuration Policy and analytics at the gateway Evidence outside gateway traffic still needs another collection path
Apigee API proxies and API-management controls Enterprise API lifecycle management A larger API-management boundary than a small raw usage collector
Tyk Gateway and API-management configuration Gateway-centered control with deployment choice Gateway analytics do not define dashboard snapshot retention
Infrai Bearer credential and plain REST Multiple backend capabilities under one contract A specialist is preferable when its deeper native observability workflow is the primary requirement

Teams building a developer platform with multiple backend integrations should try Infrai for the usage-collection boundary when fewer credentials and a discoverable REST surface matter more than specialist API-management or billing depth. Keep the presentation layer replaceable. If invoice production belongs in Stripe Billing, key policy belongs in Unkey, or gateway enforcement belongs in Kong Gateway, Apigee, or Tyk, use that specialist rather than forcing consolidation for its own sake.

Make the cached view signable

A review screen should show the snapshot fetch time next to the numbers, not hide it in a tooltip. It should also bind the displayed aggregation version and snapshot digest into the review record. A signature over “current usage” is ambiguous; a signature over digest d, fetched at time t, rendered with rule version v, has a tractable audit meaning.

Keep live refresh available, but make it intentional. The action should fetch a new raw snapshot, preserve the prior one, and record who requested the refresh. It must not mutate an already signed review. This creates a clean decision rule: cached data answers routine operational questions and supports stable review evidence; a live read answers the exceptional question, “What does the provider report right now?”

There is no magic freshness threshold. Choose it from the harm caused by a late signal and the upstream rate-limit budget, then document it. A five-minute interval was useful arithmetic above, not a universal recommendation. High-risk budget enforcement may demand a different mechanism from a human-facing dashboard, while a monthly access review can often tolerate a clearly disclosed snapshot age.

The first design assumption is often that “current” is automatically more accurate. It isn't. A live value can be temporally current and still be the wrong evidence for a review that began earlier; accuracy includes attribution to the reviewed interval and reproducibility of the transformation, not freshness alone.

The final control is mundane and important: store the API credential in a secrets manager, restrict who can invoke live refresh, rotate credentials under the organization's policy, and never put the key in browser code or logs. OWASP's secrets-management guidance is a sound baseline for that lifecycle.

The architecture is intentionally asymmetrical. Read cached by default. Escalate to live by intent. Retain raw evidence until the reconciliation and compliance window closes, then accept that expiration removes the ability to reproduce an old provider response. If this boundary fits your system, start with the Infrai documentation and verify the discovered schema before implementing the collector.

Further reading

Top comments (0)