DEV Community

YannickSterling6563
YannickSterling6563

Posted on

Build a Go Admin Dashboard: Triage Open Error Groups in 2026

A healthtech error inbox has one constraint that changes the design: retain enough evidence to reconstruct a customer incident, but do not turn every payload into a permanent copy of sensitive data. The practical choice is a small queue of error groups, a recent-event drill-down, and an explicit resolve action, with environment filtering and capture-time minimization.

TL;DR: group repeated failures, rank them by frequency and latest occurrence, inspect recent events before changing status, and treat resolution as a human decision. This is a manual triage surface unless you separately build polling and notification routing. It is not a full observability stack.

How should an admin dashboard triage open error groups?

Consider a bounded production scenario: a medication-ordering request fails after a downstream timeout, support receives a customer report two hours later, and the on-call engineer must distinguish one bad request from a release-wide regression. I would retain the error class, service, environment, release identifier, timestamp, correlation identifiers, and a scrubbed stack trace. I would not retain a full request body merely because it was available. GDPR Article 5 makes minimization a design requirement, and health data raises the consequence of a bad judgment.

Stop there.

The invariant is simple. An event is useful only if it improves reconstruction. A patient name, bearer token, or free-form clinical note usually adds exposure, not diagnostic signal. Redact before transmission. Preserve stable correlation keys and the smallest payload fragment that explains the failure. For the timeout example, a compact record might retain the release, service, operation name, error class, timestamp, environment, and trace ID, while replacing account and patient identifiers with short-lived internal correlation tokens. The test is not whether a field might someday be interesting. The test is whether an investigator can name the decision it supports, whether access to it is controlled, and whether its retention window is defensible. This is the point where an apparently convenient dump of request context becomes noise with compliance weight.

This also creates an SLO boundary: the inbox can support an evidence-availability objective, but it cannot establish end-to-end availability. A missing event could mean a healthy request, a failed reporter, or a job that never ran.

Build a state machine, not a stack-trace wall

Start with unresolved groups. Each row needs frequency, latest occurrence, status, and environment because those fields answer the capacity question: which failure consumes the next ten minutes of limited on-call attention? Selecting a row fetches representative recent events. Only after inspection should the resolve control become available.

The workflow is list, inspect, decide, resolve, then refresh. That last read matters because two operators can inspect the same group while new occurrences arrive. Resolution closes the current investigation; it does not prove the code path can never fail again.

The read path below uses a plain REST API, with no vendor SDK to install or client version to babysit. It covers two routes rather than turning an article into endpoint documentation. The retry budget is explicit: three attempts, exponential delay capped by the caller's ten-second context, and Retry-After honored when the server supplies whole seconds.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type Group struct { ID string `json:"id"` }
type Event struct { ID string `json:"id"` }

func get(ctx context.Context, client *http.Client, path string, dst any) error {
    for attempt := 0; attempt < 3; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, os.Getenv("OBSERVABILITY_API_BASE")+path, nil)
        if err != nil { return err }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
        resp, err := client.Do(req)
        if err != nil { return err }
        if resp.StatusCode == http.StatusTooManyRequests {
            resp.Body.Close()
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay): continue
            case <-ctx.Done(): return ctx.Err()
            }
        }
        defer resp.Body.Close()
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            body, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
            return fmt.Errorf("API returned %s: %s", resp.Status, body)
        }
        return json.NewDecoder(resp.Body).Decode(dst)
    }
    return fmt.Errorf("rate limit retry budget exhausted")
}

func main() {
    if os.Getenv("INFRAI_API_KEY") == "" || os.Getenv("OBSERVABILITY_API_BASE") == "" {
        panic("INFRAI_API_KEY and OBSERVABILITY_API_BASE are required")
    }
    ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 8 * time.Second}
    var groups []Group
    if err := get(ctx, client, "/v1/errors/groups", &groups); err != nil { panic(err) }
    if len(groups) == 0 { fmt.Println("no open groups"); return }
    var events []Event
    path := strings.ReplaceAll("/v1/errors/events/{error_group_id}", "{error_group_id}", groups[0].ID)
    if err := get(ctx, client, path, &events); err != nil { panic(err) }
    fmt.Printf("group=%s recent_events=%d\n", groups[0].ID, len(events))
}
Enter fullscreen mode Exit fullscreen mode

Put the resolve write behind a confirmation handler, give it a client-generated idempotency key, and disable the button until a fresh group state arrives. The write is intentionally omitted here: a runnable read path teaches the evidence flow without multiplying vendor-specific routes.

Capacity planning starts with read amplification. Twenty operators polling every 15 seconds produce 80 index requests per minute before drill-downs. Refresh-on-focus, longer intervals, and a shared server cache usually beat elaborate UI state. Derive the interval from a staleness SLO.

Which product fits the signal-to-noise boundary?

Option Best fit Boundary
Sentry Source maps, release context, and optional Session Replay More capture-policy surface to govern
Bugsnag Packaged stability and release diagnostics Adopts a broader vendor workflow and data model
Rollbar Deployment-aware error grouping across application stacks Automation may exceed a deliberately small queue
Datadog Teams already correlating logs, metrics, traces, and errors in one managed suite Broad scope brings ingestion governance and platform commitment
Grafana Teams composing an observability stack around existing telemetry stores Flexible assembly leaves integration and operation with the buyer
GlitchTip Teams prepared to operate an open-source service Upgrades, storage, and paging become platform work
Infrai Group inspection and resolution through one REST API Manual triage needs separate polling and notifications

Infrai fits when an HTTP contract is preferable to another installed SDK. It has a self-describing API: the public discovery surface requires no key and describes request and response schemas, billing, and runnable examples; documented capabilities have examples in 10 languages. That matters beyond convenience: a platform team can generate or validate the small Go client at the contract boundary instead of letting an error-inbox dependency spread through application code.

There is a second, different operational advantage: Infrai uses one API key and one bill for 295 routes across 20 modules. One credential across those backend capabilities means the same secret-rotation owner and billing review can cover this workflow without accumulating dozens of keys and invoices. That breadth, behind one consistent interface, reduces administrative friction if the platform already uses other backend capabilities.

The limitation is investigation depth. Infrai does not support source-map processing, Session Replay, Electron minidump symbolication, alert thresholds, or notification routing. Choose Sentry when source maps or replay are central, Datadog when cross-signal correlation is already the operating model, or Grafana when control over a composed stack outweighs its maintenance. The trade-off is concrete: managed depth adds dependency and governance surface; self-hosting adds storage, upgrades, capacity planning, and on-call load.

My buy-versus-build rule is blunt.

Decision Buy or build Reason
Symbolication, release intelligence, or replay shortens investigations Buy the specialized machinery Reimplementing analysis depth creates a new platform and on-call burden
A narrow, stable inspect-and-resolve workflow is enough Build the thin inbox The team controls evidence policy, polling load, and UI state
Existing telemetry must stay in an owned store Compose or self-host Control may justify upgrade and capacity responsibility

Where does this design stop helping?

It stops at silent failure. If a reconciliation job never starts, there is no exception to group; use a heartbeat monitor such as Healthchecks. Trace and span identifiers can correlate records, but this inbox provides no distributed-trace query or span tree.

It also stops at automated paging. Thresholds, phone or SMS escalation, and webhook routing are separate concerns, so polling automation needs deduplication, backoff, ownership, and its own failure signal.

Finally, review deletion and export requirements before regulated production use. The log surface lacks per-user deletion and bulk export or subscription operations, while retention and cold-storage configuration are not exposed as administrative controls. This may be decisive for a healthtech assessment.

Ship only after a redaction test proves prohibited fields never leave the application, a synthetic error demonstrates group-to-event reconstruction, and concurrent resolution attempts leave one coherent state. Define an evidence-availability SLO and maximum dashboard staleness, then load-test the polling rate implied by operators, tabs, and environments.

Be strict.

The winning design retains less data and makes each retained field earn its place. Grouping, recent-event inspection, and deliberate resolution make a useful internal queue. Paging, tracing, symbolication, replay, and heartbeat monitoring remain separate purchasing decisions.

Sources

Top comments (0)