DEV Community

CarterHughes6853
CarterHughes6853

Posted on

Property Agent Error Tracking: 6 Admin Page Signals for Event Groups

Build the error tracking admin page around a property-management agent loop captured as one trace, retain errors as immutable events, and make resolution a reversible annotation on a group rather than a mutation of its history. That decision rule gives an operator enough context to answer three questions together: which step delayed a tenant or property manager, what the loop cost, and what actually failed.

TL;DR: Capture six signals for every loop: a correlation ID, step name, monotonic duration, model usage, estimated cost under a versioned rate card, and outcome. Group errors by a normalized fingerprint, but keep every occurrence available for incident reconstruction. The admin page should default to unresolved groups, support structured search, open a representative event with its full timeline, and record who resolved or reopened the group. Averages alone are inadequate; alert and plan capacity against tail latency, error rate, and budget burn.

This is the minimum useful shape. Less context produces a fast-looking dashboard that cannot explain a slow maintenance-request workflow; indiscriminate capture creates a privacy and retention problem, especially when prompts contain tenant names, access instructions, or addresses.

How should an admin page list unresolved error groups?

Consider an agent that receives a report of a leaking sink, reads the lease and property policy, checks vendor availability, and drafts a response. A single user action may cross retrieval, model calls, tool calls, retries, and a final write. The visible failure may occur at the end even though the expensive delay began several steps earlier.

The event detail view therefore needs a timeline, not a bag of log lines. Each step should carry its start and end, status, attempt number, dependency identity, input and output sizes where they are safe to retain, and links to sanitized logs. Record wall-clock timestamps for ordering across services and monotonic durations for elapsed time inside a process. Clock skew can distort subtraction between wall-clock timestamps; a measured duration avoids that trap.

Six signals form the operational spine:

  1. correlation_id joins the request, trace, logs, and stored error occurrence.
  2. step identifies retrieval, inference, tool execution, or persistence without encoding tenant data.
  3. duration_ms exposes both the total and the slow component.
  4. usage records the provider-reported input and output units when available.
  5. estimated_cost_micros applies a named, versioned rate card so historical estimates remain explainable.
  6. outcome distinguishes success, timeout, cancellation, dependency rejection, and application error.

Do not stuff raw prompts into the group title. A useful fingerprint combines stable fields such as operation, error class, normalized message, and the failing frame or component. Strip request IDs, timestamps, addresses, and other high-cardinality values before hashing. Keep the original sanitized occurrence separately, because grouping is a lossy index and an operator eventually needs evidence.

Short loops still fail.

Build the event path before the dashboard

The safe implementation starts at emission. Applications should write logs as event streams; collection, routing, indexing, and retention belong outside the request handler. OpenTelemetry provides a vendor-neutral model for traces, metrics, logs, context propagation, and semantic conventions, while W3C Trace Context defines interoperable HTTP headers for passing trace identity between services.

The following Go example models a step record without tying the producer to a storage product. The cost value is an estimate, not an invoice: it is computed from observed usage and a rate-card version, then labeled so a later pricing change cannot silently rewrite history.

package telemetry

import (
    "context"
    "encoding/json"
    "io"
    "time"
)

type Usage struct {
    InputUnits  int64 `json:"input_units"`
    OutputUnits int64 `json:"output_units"`
}

type StepEvent struct {
    CorrelationID      string `json:"correlation_id"`
    TraceID            string `json:"trace_id"`
    Step               string `json:"step"`
    Attempt            int    `json:"attempt"`
    Outcome            string `json:"outcome"`
    DurationMS         int64  `json:"duration_ms"`
    Usage              Usage  `json:"usage"`
    EstimatedCostMicros int64  `json:"estimated_cost_micros"`
    RateCardVersion    string `json:"rate_card_version"`
    OccurredAt         string `json:"occurred_at"`
}

func ObserveStep(ctx context.Context, out io.Writer, base StepEvent, run func(context.Context) (Usage, int64, error)) error {
    started := time.Now()
    usage, costMicros, err := run(ctx)

    base.DurationMS = time.Since(started).Milliseconds()
    base.Usage = usage
    base.EstimatedCostMicros = costMicros
    base.Outcome = "ok"
    if err != nil {
        base.Outcome = "error"
    }
    base.OccurredAt = time.Now().UTC().Format(time.RFC3339Nano)
    _ = json.NewEncoder(out).Encode(base)
    return err
}
Enter fullscreen mode Exit fullscreen mode

Production code must handle encoder failures through a bounded, non-recursive fallback path; telemetry must not turn a recoverable model failure into request failure. It also needs explicit redaction before serialization. The example deliberately omits prompt text, response text, unit prices, and tenant identifiers.

On the ingestion side, append the occurrence first, then update the group index. A group row can cache first_seen, last_seen, occurrence count, affected correlation count, latest deployment, and resolution state. The occurrence remains immutable. Resolution belongs in an audit record containing group ID, actor, timestamp, reason, and the group fingerprint version; reopening adds another record instead of erasing the first.

Search should target documented fields such as status, environment, operation, deployment, error class, and time range. Free-text search is useful for sanitized messages, but it should not be the only way to find all unresolved production failures from a specific rollout. Parse filters into a query model, reject unknown fields, cap the time range, paginate with a stable cursor, and display the active filters. Otherwise an empty result is ambiguous: no errors, a typo, or a truncated query can look identical.

Turn the admin page into an incident instrument

The first screen should be a dense list of groups, ordered by recent impact rather than raw lifetime count. An operator needs unresolved state, last seen, recent occurrences, affected loops, representative outcome, deployment, and tail latency contribution in the same row. Clicking a row opens one event and its surrounding loop timeline; moving between occurrences must preserve the search context.

Severity is not a synonym for frequency. One repeated validation error may be noisy but harmless, while a rare failure after a work order is dispatched can leave the property manager uncertain about whether a contractor was contacted. Define impact using the user-visible outcome and workflow boundary, then add frequency and recency. This is where SLO language matters: page when the error consumes error budget or threatens a critical workflow objective, not merely because a counter crossed an arbitrary integer.

For latency, show distributions by operation and step. Percentiles are useful only with enough samples and an explicit window; for a sparse building portfolio, list the slowest recent traces alongside the distribution instead of presenting a shaky percentile as certainty. For cost, aggregate estimated micro-units by successful loop, failed loop, retry, model operation, and property workflow. Averages conceal retry storms. Track the upper tail and total burn together.

The ownership decision is less glamorous, but it determines whether the console survives contact with on-call reality:

Capability Build and operate Managed service Decision test
Event capture Standards-based schema and full control Faster initial integration Can data leave the chosen boundary?
Grouping Tunable fingerprints, ongoing maintenance Mature defaults, less control Can fingerprints be versioned and explained?
Search and retention Capacity and index work stay in-house Operational load moves outward Can evidence be exported before expiry?
Resolution audit Fits internal authorization exactly Depends on available roles and audit model Can every state change be attributed and reversed?
Cost attribution Custom rate cards and workflow dimensions May provide prebuilt usage views Can estimates be reconciled to invoices?

I would keep the producer schema and correlation identifiers portable even when buying storage and search. That isn't an ideological preference for self-hosting; it limits migration risk and preserves the option to route security-sensitive events differently. The capacity plan still needs estimated events per loop, peak loops per second, average encoded bytes, index expansion, retention days, and replay headroom. If nobody owns those numbers, the architecture decision is unfinished.

There is a real limitation here: building this event path is unsuitable for a small team that cannot own ingestion, index capacity, privacy deletion, and an on-call rotation for the console itself. A managed backend can remove much of that operational load, but it trades away some control over grouping behavior, retention boundaries, and export. Conversely, a self-hosted store may fit strict data-boundary requirements while consuming engineering time that would otherwise improve the property workflow. Choose against those constraints, not against a screenshot.

Verify failure paths, then practice rollback

Test incident reconstruction as an acceptance criterion. Generate a synthetic maintenance request, force a timeout in one tool step, allow one bounded retry, and verify that the list shows one group with two distinct occurrences or attempts as designed. Open the event and confirm that the trace timeline explains the latency, the usage totals do not double-count a completed step, and redacted fields cannot be recovered through search, export, or cached previews.

Then test the controls that are easy to trust without evidence:

  • Resolve a group, send a matching new occurrence, and verify the documented reopen policy.
  • Change the fingerprint algorithm and confirm that old groups retain their version and audit history.
  • Drop or delay telemetry delivery; the property workflow must continue, and the loss must surface through collector health metrics.
  • Apply retention deletion and verify removal from primary storage, search indexes, exports, and derived previews according to policy.
  • Roll back an application deployment and confirm that events retain both deployment identity and trace continuity.

A safe rollout begins with shadow emission and schema validation, then enables indexing for a small traffic slice before turning on alerts. Watch rejected-event rate, queue depth, ingest latency, index growth, and application overhead. Set rollback thresholds before deployment. If overhead or rejection breaches them, disable the new emitter path through a controlled configuration change while leaving the previous event stream intact.

Do not use resolution as deletion. During an incident, the reversible action is to acknowledge or resolve the group with a reason; retention and privacy deletion are separate governed operations. Mixing them makes the button dangerous and breaks the evidence chain needed to determine whether the same failure returned.

Operate against budgets, not dashboard aesthetics

The finished system should let an on-call engineer move from an unresolved group to a representative event, then to the complete agent-loop timeline, without inventing joins by hand. It should also let the platform owner answer a harder question: are latency and cost within budgets for successful property workflows, or are retries and partial failures consuming capacity without producing an outcome?

Set an availability or completion SLO at the workflow boundary, define latency objectives for the whole loop and its important steps, and maintain a cost budget per completed workflow using a versioned estimator. Review false merges and false splits in grouping after deployments. Sample successful traces intelligently, but retain error traces under a policy that reflects privacy, investigation needs, and storage capacity.

The console is good enough when it supports a defensible reconstruction, not when it has the most charts. Preserve the event, explain the group, audit the state change, and make retreat possible.

References

Top comments (0)