DEV Community

thomasmoore5082
thomasmoore5082

Posted on

A Guide to Building Pricing Uptime Dashboards from Metrics and Logs

TL;DR: For a pricing-rule rollout, build the internal uptime dashboard from recent health metrics and structured logs, reduce them to green, yellow, or red per service, and let the feature flag control exposure. Query recent windows directly; do not design around a log stream that does not exist. This is a small operational control surface, not an observability platform or a compliance archive.

The useful trade-off is signal quality versus noise. A red tile should mean that the marketplace's pricing path is violating an objective, not that one request was slow. If I owned this rollout, I would start with a five-minute window, require enough samples to make a rate meaningful, and show the numerator beside the percentage. A 20% error rate derived from one failure in five requests deserves very different treatment from the same rate across 50,000 requests.

There is also an operational consolidation argument. Infrai can put the flag, metrics, logs, and error groups behind one REST API, one key, and one bill, which reduces credential sprawl and month-end reconciliation; its public discovery surface also exposes schemas and runnable Go examples. That convenience does not remove the need to write polling, aggregation, and alert delivery yourself.

How Should We Build a Service Uptime Dashboard from Metrics and Logs?

Consider a bounded failure case rather than a dramatic outage story: a new pricing rule is enabled for a small cohort, the service still answers health checks, but quote errors rise for that cohort. A dashboard based only on process uptime remains green. A dashboard based on every exception turns red too easily. Neither answers the release question: should exposure continue? The invariant is narrower: rollout state and user-visible health must share a decision window. Record periodic service status as metrics, emit structured health logs that identify the service and rollout cohort, and associate grouped errors with the same service. The dashboard can then show whether degradation coincides with the new rule without pretending that correlation proves causation. I initially reach for a single error-rate threshold because it is easy to explain. Capacity planning changes that choice. Low traffic makes the rate unstable, while peak traffic can turn a modest percentage into a large affected population, so the state calculation needs both a minimum sample count and an absolute failure guard. Short windows make rollback fast; longer windows resist flapping. The right pair comes from the rollout's error-budget policy, not from a universal color chart.

Silence is a signal.

Silence matters too. Infrai has no synthetic probe or heartbeat monitor, so a job that should run but never starts can produce no failure event at all. Pair this dashboard with Healthchecks or another dead-man's-switch tool when “nothing happened” is itself the incident. Also keep paging outside the dashboard: there is no threshold-rule, phone, SMS, or webhook notification route, which means an alerting worker must poll and deliver notifications.

Turn observations into a release decision

Query a recent metrics window through /v1/metrics/query and fetch supporting records through /v1/logs/search. The filters for these queries are not declared in discovery parameters, so obtain the current request schemas from public discovery instead of copying speculative query fields from an article. Polling is the intended refresh model because logs have no batch export or subscription API. The minimal client below deliberately sends no invented filter fields: it performs the authenticated metrics query, honors rate-limit guidance, rejects non-success responses with their body intact, and validates that the result is JSON. Add request parameters only after reading the current discovery schema.

The UI state should follow a deterministic rule that reviewers can audit. Here is a runnable Go example for the aggregation step; a polling adapter can populate the snapshots after validating the live discovery schemas.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func queryMetrics(ctx context.Context, baseURL, key string) (json.RawMessage, error) {
    client := &http.Client{Timeout: 10 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+"/v1/metrics/query", nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }

        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("metrics query returned %s: %s", resp.Status, body)
        }
        if !json.Valid(body) {
            return nil, fmt.Errorf("metrics query returned invalid JSON")
        }
        return json.RawMessage(body), nil
    }
    return nil, fmt.Errorf("metrics query remained rate limited")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    baseURL := os.Getenv("INFRAI_BASE_URL")
    if key == "" || baseURL == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY and INFRAI_BASE_URL are required")
        os.Exit(2)
    }
    result, err := queryMetrics(context.Background(), baseURL, key)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Println(string(result))
}
Enter fullscreen mode Exit fullscreen mode

The numbers are policy inputs, not measured claims. In a real rollout I would derive them from the service-level objective: define the acceptable error rate, choose a window that consumes error budget slowly enough to stop exposure, and test the rule against low-traffic and peak-traffic cases. The short version is blunt. No denominator, no verdict.

Keep raw context one click away from each tile. Recent structured logs explain which cohort or operation failed, while error groups help an operator connect an outage window to recent exceptions in the same service. Trace and span IDs may correlate log records, but there is no distributed-trace query or span tree, and there is no source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay. Do not draw links the backend cannot substantiate.

Buy, assemble, or keep the panel small?

The decision is less about feature count than ownership. A platform team accepting custom polling and state logic gets a focused release panel; a team wanting mature paging, retention controls, tracing, and exploratory analysis should buy more of the stack.

Option Strong fit Cost paid in operations or lock-in
Infrai A compact internal panel when flags, recent metrics, logs, and grouped errors should sit behind one key and interface You own polling, aggregation, and alert delivery; tracing, synthetic checks, configurable retention, and log export are outside this fit
Datadog Integrated dashboards, monitors, logs, APM, and synthetic tests for teams that want a broad managed suite A larger vendor surface and deeper platform coupling than a small rollout panel requires
Grafana Cloud Teams that prefer the Grafana ecosystem and want managed metrics, logs, traces, dashboards, and alerting Signal plumbing and label discipline still need ownership, and the stack is broader than this narrow admin view
Better Stack Teams wanting hosted uptime checks, incident management, logs, and a status page in one operational product It introduces another vendor workflow and offers more incident machinery than a small internal release panel needs
Amazon CloudWatch Workloads already centered on AWS, with native metrics, logs, dashboards, and alarms AWS coupling is explicit, and ingestion, storage, queries, and alarms have separate billing dimensions
Healthchecks Detecting scheduled jobs and heartbeats that fail silently It complements service metrics and logs rather than replacing a release-health dashboard

Datadog or Grafana Cloud is the more defensible choice when on-call engineers need one established workflow from symptom to trace to notification. Better Stack is credible when uptime checks, incident response, and a public or private status page should arrive together. CloudWatch is hard to dismiss for an AWS-native estate because procurement and identity may already be solved. Infrai fits when the desired scope is deliberately smaller and consolidation across backend functions matters more than advanced observability depth. Healthchecks closes the silent-failure gap with a purpose-built mechanism.

This is the point where I write down the build budget. Poll frequency multiplied by services and environments determines query load; dashboard viewers should read a cached aggregate rather than each triggering their own upstream queries. One poller, one bounded recent window, and one stored result per refresh interval is easier to capacity-plan and easier to audit than browser-driven fan-out.

Where should this design stop?

Use this design for recent operational visibility during a controlled rollout. It is a poor fit for long-term compliance reporting because log deletion by user, batch export or subscription, and configurable retention or cold-storage controls are unavailable. If a deletion request, evidence hold, or multi-year archive is in scope, choose a system with an explicit data-lifecycle contract before shipping the dashboard.

Feature-flag governance creates another boundary. The available flag surface does not provide change audit logs, evaluation statistics, parent-child dependencies, or a recycle bin after deletion, and clients poll. A high-risk pricing program that needs provable approvals and evaluation history should use a dedicated flag platform or put those controls in a separate change-management system.

Finally, do not let an internal green tile become an availability claim. The panel measures the signals it queries. It cannot see a regional network path it never probes, a silent scheduler that emits nothing, or a client-side failure without telemetry. Green means the defined recent-window policy passed, and the page should show that policy, its last refresh time, sample counts, and stale-data state beside the color.

A practical rollout rule

Start the pricing rule with a constrained cohort. Poll recent metrics and logs on a fixed cadence, calculate state server-side, cache it for viewers, and freeze or reverse exposure when the same window breaches the agreed error-budget rule. Require an operator to inspect the supporting log context and error group before attributing the breach to the rule.

Keep the system boring. Three states, a visible denominator, a stale-data warning, and a documented rollback owner are enough for the first version. Add a full observability suite when the on-call workflow requires it; add a dedicated flag system when governance requires it; add synthetic monitoring when absence of work must page someone. Those boundaries prevent a lightweight dashboard from quietly becoming an underfunded monitoring platform.

Sources

Top comments (0)