DEV Community

RhettMurray8263
RhettMurray8263

Posted on

Marketplace Agent Logs — Production API JSON Search for Incident Reconstruction

TL;DR: Optimize log management for reconstruction, not collection volume. A marketplace agent run should leave one correlated sequence of request, model, tool, retry, and outcome events, with elapsed time and usage recorded at their actual boundaries. Choose a system only after a replay exercise proves that a beginner can search structured fields, explain a slow or costly run, and apply retention or deletion policy without reading raw prompt text.

That decision makes the dashboard secondary. The durable asset is the event contract: small JSON records linked by stable identifiers. A dashboard can change later; ambiguous evidence cannot be repaired after an incident.

How should beginners manage Express API production logs and errors?

Imagine a buyer asks a marketplace assistant to find three in-stock cameras, compare delivery dates, and stay under a budget. One API request may trigger retrieval, several model calls, inventory tools, and a retry. A useful incident view must answer: which step waited, which step failed, how much model usage belonged to the run, and whether the final response reflected the latest tool result?

Start with four identifiers: request_id for the incoming API call, run_id for the agent loop, step_id for one operation, and parent_step_id for causality. Add an event name, UTC timestamp, duration, status, model-usage fields when present, and a small set of marketplace dimensions such as tool name or result count.

Keep the schema boring. That is a feature.

Do not make the prompt the primary evidence. Prompts and tool payloads can contain buyer details, seller data, or free-form text with unpredictable sensitivity. Record bounded metadata by default and place any approved content capture behind a separate policy. GDPR Article 17 is also a practical warning against treating logs as an exemption from deletion obligations. The storage design needs a way to locate data connected to an erasure request and apply the relevant policy.

Build the reconstruction path before the dashboard

The data flow is straightforward. The API creates correlation identifiers, every agent boundary emits one JSON event, an asynchronous shipper moves those events away from the request path, and storage indexes the few fields used during incidents. Metrics are derived from the same boundaries for trend detection, while event records retain the detail needed to reconstruct an individual run. OpenTelemetry describes a metric as a measurement about a service captured at runtime; metrics can reveal a latency or usage shift, but ordered events carry the run-level explanation.

Here is a runnable local model of that contract. It uses a list as the sink so the example can be exercised in a notebook before the sink becomes a production queue or collector. The sample values are fixtures, not benchmark results.

from __future__ import annotations

import json
from collections import defaultdict
from datetime import datetime, timezone
from decimal import Decimal
from typing import Any
from uuid import uuid4


def emit(events: list[dict[str, Any]], **fields: Any) -> None:
    event = {
        "timestamp": datetime.now(timezone.utc).isoformat(),
        "schema_version": 1,
        **fields,
    }
    events.append(event)
    print(json.dumps(event, separators=(",", ":"), sort_keys=True))


def summarize(events: list[dict[str, Any]], run_id: str) -> dict[str, Any]:
    selected = [event for event in events if event["run_id"] == run_id]
    selected.sort(key=lambda event: event["timestamp"])

    duration_by_kind: dict[str, int] = defaultdict(int)
    input_tokens = 0
    output_tokens = 0
    estimated_cost = Decimal("0")

    for event in selected:
        duration_by_kind[event["kind"]] += int(event.get("duration_ms", 0))
        usage = event.get("usage", {})
        input_tokens += int(usage.get("input_tokens", 0))
        output_tokens += int(usage.get("output_tokens", 0))
        estimated_cost += Decimal(str(usage.get("estimated_cost", "0")))

    return {
        "run_id": run_id,
        "event_count": len(selected),
        "duration_ms_by_kind": dict(duration_by_kind),
        "input_tokens": input_tokens,
        "output_tokens": output_tokens,
        "estimated_cost": str(estimated_cost),
        "statuses": [event["status"] for event in selected],
    }


events: list[dict[str, Any]] = []
request_id = str(uuid4())
run_id = str(uuid4())
root_step_id = str(uuid4())

emit(
    events,
    event_name="agent.step.completed",
    kind="model",
    request_id=request_id,
    run_id=run_id,
    step_id=root_step_id,
    parent_step_id=None,
    status="ok",
    duration_ms=420,
    usage={
        "input_tokens": 610,
        "output_tokens": 84,
        "estimated_cost": "0.0042",
        "currency": "USD",
        "pricing_version": "fixture-v1",
    },
)
emit(
    events,
    event_name="agent.step.completed",
    kind="tool",
    request_id=request_id,
    run_id=run_id,
    step_id=str(uuid4()),
    parent_step_id=root_step_id,
    status="ok",
    duration_ms=95,
    tool_name="inventory_lookup",
    result_count=3,
)

print(json.dumps(summarize(events, run_id), indent=2, sort_keys=True))
Enter fullscreen mode Exit fullscreen mode

The cost field is deliberately labeled estimated_cost, paired with a currency and a pricing-version label. Token counts are observations; converting them into money is a model that can change. Keeping those concepts separate lets an eval harness compare prompt variants on stable usage measurements while finance logic evolves independently.

One trap deserves emphasis: summing every duration does not necessarily produce wall-clock latency. Agent steps may overlap. Store start and end boundaries, then calculate critical-path latency separately from cumulative work. A retry should be a new step with an attempt number, not an overwrite of the failed event. The failed attempt is often the clue.

Evaluate search with incident questions

A beginner-friendly search box is helpful, but syntax is not the acceptance test. Seed a staging dataset with a slow inventory call, a model retry, a cancelled request, and two concurrent tool calls. Then hand an operator a request_id and time the reconstruction. Can they find the whole run, order its events, isolate the longest boundary, total model usage once, and distinguish cumulative work from wall time?

This turns selection into an eval instead of a screenshot contest. Score candidates against the same questions and the same exported JSON. Check exact-match filters for identifiers and status, numeric range filters for duration and usage, time-zone handling, saved views, access controls, retention controls, deletion workflow, and export fidelity. Also simulate malformed events and sink unavailability: application work needs an explicit policy for buffering or dropping telemetry, and telemetry failure must not silently change the business response.

Prefer the option that preserves evidence under stress and makes uncertainty visible. A polished chart that double-counts retries is worse than a plain table with correct lineage.

Test cardinality before rollout. IDs are valuable for point lookup but poor grouping dimensions in aggregate charts. Group trends by bounded fields such as event name, status, tool name, or model role; reserve high-cardinality identifiers for event search. OpenTelemetry's metrics guidance discusses attributes as dimensions and warns that large numbers of combinations can create a cardinality explosion.

Separate trend signals from case evidence

Use metrics to notice drift: agent-run latency, step latency by kind, error counts, retry counts, and token totals over a chosen interval. Use structured events to explain the selected case. Trying to force both jobs into one representation either removes incident detail or creates noisy aggregations.

The boundary is clean. A metric says that tool latency changed. A correlated event sequence says which inventory call, attempt, and downstream model step shaped one buyer's response. Keep both tied to the same low-cardinality vocabulary, then validate them against the eval fixture during deployment.

Tiny records win.

This event-first design has a real limitation. It is a poor fit when the only requirement is a short-lived local debug stream, because schema ownership and correlation add work before they add value. It also cannot replace traces for critical-path visualization or metrics for low-cost aggregate alerting. The trade-off is deliberate: keep searchable case evidence in events, use metrics for trends, and add tracing when concurrency makes parent-child timing hard to reason about from timestamps alone.

Operate the contract, not a pile of fields

Before production, publish the event schema and name its owner. Reject or quarantine events with missing correlation fields, cap arbitrary attributes, redact at the producer, and document what must never be captured. Roll out schema changes with a version field and keep readers compatible across the migration window. The operational checklist should live beside the eval: confirm timestamps are UTC, retries remain distinct, asynchronous work carries context, sampling retains errors according to policy, cost estimates name their pricing version, and deletion can reach every relevant store and derived copy.

During an incident, begin with the request or run identifier, construct the ordered tree, compare wall time with cumulative step time, and inspect errors before payload content. Afterward, convert the failure into a fixture and a query assertion. That loop is how notebook exploration becomes production discipline: the same compact records support local experiments, deployment gates, alert trends, and a defensible reconstruction without coupling the application to a particular dashboard.

Sources

References used for the standards and policy boundaries:

Top comments (0)