Short answer: keep raw usage events long enough to explain an access decision, then derive rolled-up totals for the dashboard and cache those totals on a schedule. For an edtech access review, the signed artifact should point from each total to an immutable event window, an actor, and a policy version. A fast chart without that trail is decoration, not evidence.
I care about this because usage systems fail in quiet ways. A retry can look like a student action. A delayed queue message can land in the next day. During an OTP review, I once saw a “zero sends” hour that was really a timezone conversion bug. The dashboard was green; the review was not defensible.
Start with the bill: what should the dashboard retain?
The dominant cost is usually retention and query work on raw events, not the few kilobytes of a daily total. Measure event count, payload bytes, index size, and query frequency before choosing a storage policy. In an edtech platform, a useful event might contain tenant_id, course_id, actor type, endpoint name, status class, request ID, and UTC timestamp. It should not contain a bearer token or a student's message body. I would also write down the accounting boundary before anyone tunes a database: does a retry count as an attempted call, a delivered response, or both? A 429 followed by a successful retry can otherwise make a school look like it exceeded a quota. For one access review, I split those states into attempted, accepted, and failed counters, then had the reviewer sign the definition along with the number. That small piece of prose prevented a week of arguing over whether the chart was “wrong.”
Keep the definition visible.
Keep the raw record immutable for the period your access policy and review cadence require. Store a compact rollup keyed by UTC hour, tenant, endpoint, and status class. The rollup is what the Node.js dashboard reads most often; the raw slice is what an auditor samples when a total is challenged. Retention is a decision, not a default.
The change that moves the bill is to stop asking the chart query to scan raw history. Write the event once, aggregate in a scheduled job, and serve the aggregate from a cache with an explicit freshness timestamp. You save repeated scans, but you deliberately stop keeping high-cardinality dimensions that nobody can explain in a signed review. When an incident needs that missing dimension, the cost is a narrower reconstruction and a slower decision.
How should an internal API usage dashboard choose raw timeseries, rolled-up totals, and a Node.js cache schedule?
Use a two-layer contract. Raw timeseries answers “which request happened?” Rolled-up totals answer “how much activity crossed this policy boundary?” The cache answers “when was this answer last computed?” Treat those as different claims and label them in the UI.
Here is a small Python model for the rollup boundary. It uses UTC buckets and keeps the source event window beside the number, so an export can be checked later.
from dataclasses import dataclass
from datetime import datetime, timezone
@dataclass(frozen=True)
class UsageEvent:
tenant_id: str
endpoint: str
status_class: int
occurred_at: datetime
def hour_bucket(value: datetime) -> datetime:
value = value.astimezone(timezone.utc)
return value.replace(minute=0, second=0, microsecond=0)
def roll_up(events: list[UsageEvent]) -> dict[tuple, int]:
totals: dict[tuple, int] = {}
for event in events:
key = (event.tenant_id, event.endpoint, event.status_class,
hour_bucket(event.occurred_at))
totals[key] = totals.get(key, 0) + 1
return totals
The schedule should be derived from lateness, not from a fashionable interval. If 99% of queue events arrive within six minutes, a fifteen-minute rollup may be reasonable; if imports arrive after an hour, mark earlier buckets provisional and re-open them. I’m not sure one schedule fits every school calendar. Your mileage may vary during enrollment peaks.
For a Node.js service, cache keys should include the policy version and the complete filter, while the value includes computed_at, window_start, window_end, and source_revision. A stale value can be displayed if it is labelled stale and the review workflow refuses to sign it. Cache invalidation is part of the evidence contract.
What makes a usage number auditable instead of merely plausible?
An access review needs reproducibility. Freeze the query definition, the timezone, the identity join, and the policy revision. Record who generated the export, when it was generated, and the raw-event interval used. Hash the exported rows or store an append-only manifest so a later reviewer can detect edits.
Do not join on an API key. Join on a workload identity and request ID, then map that identity to the school, service account, and approved role at the time of the event. OWASP’s Secrets Management Cheat Sheet recommends limiting secret exposure and auditing secret use; the same discipline applies to usage evidence. A secret value never belongs in a dashboard row.
Test the awkward cases: duplicate delivery, a retry after a 429, daylight-saving transitions, an event arriving after its bucket was signed, and a revoked service account making a queued request. I write these as deterministic fixtures, then compare the raw slice and rollup. Three words matter here: prove the join.
A retention decision that survives review
| Data layer | Keep for | Review purpose | Failure boundary |
|---|---|---|---|
| Immutable raw events | Policy-defined window | Explain one request and its actor | Storage and privacy load grow with volume |
| Hourly rollups | Longer reporting window | Compare totals and spot anomalies | Cannot answer dimensions omitted at aggregation |
| Cached dashboard result | Minutes to hours | Fast operator view | Can be stale unless freshness is explicit |
| Signed export manifest | Review lifecycle | Prove exactly what was approved | Does not restore deleted source events |
The catch is that longer retention is not automatically better. It can increase privacy exposure and review scope. This design is not suitable when regulations require immediate deletion of event-level data; use a minimal aggregate plus a short, access-controlled evidence window. Stick with raw retention when disputes, incident response, or chargeback rules require request-level proof.
I would sign an access review only when the dashboard can show its freshness, policy version, and source interval next to the number. If any of those are missing, the correct status is “pending evidence,” not a confident zero.
Top comments (0)