DEV Community

MidnightEcho794261
MidnightEcho794261

Posted on

Node.js Express Error Capture: Request Context for Logistics Cost Tracking

A useful error record for a logistics agent must connect the failure to one shipment, one request, and the work already consumed by the agent loop. Capture request context and cumulative usage at the same boundary that records the stack trace. Otherwise, a rejected model call may be visible while the latency and token spend that preceded it disappear into aggregate charts. The implementation below uses Express, AsyncLocalStorage, and a small structured event sink; it keeps the transport replaceable and the cost ledger explicit.

TL;DR: give every inbound request a correlation ID, maintain per-request agent totals in async-local state, turn route failures into structured events in the final Express error handler, and reserve process-level handlers for last-chance reporting plus shutdown. Do not put raw prompts, authorization headers, or shipping addresses into that context.

How can Node.js Express capture an API error with request context?

For a shipment-planning endpoint, the record should answer four operational questions: which request failed, where it failed, how long the agent loop ran, and how much billable usage accumulated before failure. That is enough to group failures by route and agent step while attributing work to a shipment's internal reference. It is deliberately not a copy of the entire request.

Context is the ledger.

Use stable, bounded fields for aggregation: route template, HTTP method, error type, agent step, and model family. Keep high-cardinality values such as requestId and shipmentRef in event fields, not metric names or labels. Prometheus naming guidance recommends a common prefix, base units, and labels that preserve meaningful aggregation. A metric such as logistics_agent_step_duration_seconds can aggregate; a metric name containing a shipment ID cannot.

Severity also needs a boring rule. A handled validation error is not the same operational event as a process-ending exception. RFC 5424 defines severity levels, including Error, Critical, Alert, and Emergency. Map the application's small internal severity set to those semantics at the export boundary rather than inventing a new meaning in each handler.

Keep that mapping dull.

Build the capture path first

This example assumes Node.js 20 or later, Express, and TypeScript. The event destination is an interface, so the same capture path can write to a local structured log, a queue, or an observability backend without changing route code. The request carries a client-supplied ID only when it matches a conservative character set; otherwise the server creates one.

import express, { ErrorRequestHandler, RequestHandler } from "express";
import { AsyncLocalStorage } from "node:async_hooks";
import { randomUUID } from "node:crypto";

type Usage = { inputTokens: number; outputTokens: number };
type RequestContext = {
  requestId: string;
  shipmentRef?: string;
  route?: string;
  startedAtMs: number;
  agentStep?: string;
  usage: Usage;
};

type ErrorEvent = {
  timestamp: string;
  severity: "error" | "critical";
  message: string;
  errorType: string;
  stack?: string;
  request?: {
    id: string;
    method?: string;
    route?: string;
    shipmentRef?: string;
  };
  agent?: { step?: string; inputTokens: number; outputTokens: number };
  elapsedMs?: number;
};

interface EventSink {
  capture(event: ErrorEvent): Promise<void>;
  flush(deadlineMs: number): Promise<void>;
}

const sink: EventSink = {
  async capture(event) {
    process.stderr.write(`${JSON.stringify(event)}\n`);
  },
  async flush() {
    // stderr is the transport in this example, so there is no remote buffer.
  },
};

const context = new AsyncLocalStorage<RequestContext>();
Enter fullscreen mode Exit fullscreen mode

The middleware establishes context before any route work starts. Notice what is absent: request bodies, prompt text, carrier credentials, and customer addresses. The shipment reference should be an internal opaque identifier, not a tracking number exposed to a customer.

const requestContext: RequestHandler = (req, res, next) => {
  const supplied = req.header("x-request-id");
  const requestId =
    supplied && /^[A-Za-z0-9._-]{1,80}$/.test(supplied)
      ? supplied
      : randomUUID();

  res.setHeader("x-request-id", requestId);
  context.run(
    {
      requestId,
      startedAtMs: Date.now(),
      usage: { inputTokens: 0, outputTokens: 0 },
    },
    next,
  );
};

function addUsage(step: string, usage: Usage): void {
  const current = context.getStore();
  if (!current) return;
  current.agentStep = step;
  current.usage.inputTokens += usage.inputTokens;
  current.usage.outputTokens += usage.outputTokens;
}

function asError(reason: unknown): Error {
  if (reason instanceof Error) return reason;
  if (typeof reason === "string") return new Error(reason);
  return new Error("Non-Error rejection reason");
}

function toEvent(error: Error, severity: ErrorEvent["severity"]): ErrorEvent {
  const current = context.getStore();
  return {
    timestamp: new Date().toISOString(),
    severity,
    message: error.message,
    errorType: error.name,
    stack: error.stack,
    request: current
      ? {
          id: current.requestId,
          route: current.route,
          shipmentRef: current.shipmentRef,
        }
      : undefined,
    agent: current
      ? { step: current.agentStep, ...current.usage }
      : undefined,
    elapsedMs: current ? Date.now() - current.startedAtMs : undefined,
  };
}
Enter fullscreen mode Exit fullscreen mode

The agent adapter should return usage alongside its result. Add that usage immediately after each completed call, not at the end of the whole loop. If step three rejects, steps one and two still count.

type AgentReply = { text: string; usage: Usage };

async function callAgentStep(input: string): Promise<AgentReply> {
  // Replace with an adapter that returns the provider's reported usage.
  return { text: input, usage: { inputTokens: 12, outputTokens: 8 } };
}

const planShipment: RequestHandler = async (req, res, next) => {
  try {
    const current = context.getStore();
    if (current) {
      current.route = "/shipments/:shipmentRef/plan";
      current.shipmentRef = String(req.params.shipmentRef);
    }

    const reply = await callAgentStep("Choose the next planning action");
    addUsage("choose-action", reply.usage);
    res.json({ requestId: current?.requestId, plan: reply.text });
  } catch (reason) {
    next(reason);
  }
};

const captureExpressError: ErrorRequestHandler = async (
  reason,
  req,
  res,
  _next,
) => {
  const error = asError(reason);
  const event = toEvent(error, "error");
  if (event.request) event.request.method = req.method;

  try {
    await sink.capture(event);
  } finally {
    if (!res.headersSent) {
      res.status(500).json({
        error: "Internal server error",
        requestId: context.getStore()?.requestId,
      });
    }
  }
};

const app = express();
app.use(express.json({ limit: "64kb" }));
app.use(requestContext);
app.post("/shipments/:shipmentRef/plan", planShipment);
app.use(captureExpressError);
app.listen(3000);
Enter fullscreen mode Exit fullscreen mode

The synthetic usage values make the example runnable; they are not pricing or a benchmark. In production, record the usage returned by the model adapter. Monetary attribution belongs in a separate calculation keyed by model and effective rate version, because a hard-coded dollar figure inside an error event becomes stale and cannot be audited later.

Failure still consumed work.

Where do unhandled failures belong?

Express error middleware is the normal path for request failures. Process events are a last line of evidence, and their context may be absent because not every asynchronous failure originates inside an HTTP request. Node's documentation says uncaughtException is a crude mechanism intended for synchronous cleanup before shutdown; resuming normal operation after one is unsafe. Treat unhandledRejection with the same operational caution instead of trying to manufacture an HTTP response from a process handler.

let shuttingDown = false;

async function reportFatal(reason: unknown, origin: string): Promise<void> {
  if (shuttingDown) return;
  shuttingDown = true;

  const error = asError(reason);
  await sink.capture({
    ...toEvent(error, "critical"),
    message: `${origin}: ${error.message}`,
  });

  await Promise.race([
    sink.flush(1_500),
    new Promise<void>((resolve) => setTimeout(resolve, 1_500)),
  ]);
  process.exitCode = 1;
}

process.on("uncaughtException", (error, origin) => {
  void reportFatal(error, origin).finally(() => process.exit(1));
});

process.on("unhandledRejection", (reason) => {
  void reportFatal(reason, "unhandledRejection").finally(() => process.exit(1));
});
Enter fullscreen mode Exit fullscreen mode

One trap is easy to miss: the reporter can fail too. A production sink therefore needs a bounded flush, a non-network fallback such as stderr, and recursion protection. Keep this path tiny. Do not perform normal business cleanup or start another agent call after the process has entered fatal shutdown.

Then stop.

There is also a timing trade-off. A 1,500 ms ceiling gives a buffered exporter a chance to drain, but it delays replacement of the unhealthy process by the same maximum. Pick that ceiling from the deployment platform's termination budget and test it by forcing a real child process to fail. The number above is an example policy, not a universal default.

Cost attribution changes the schema

Latency and token counts answer different questions. Duration shows the user-visible and infrastructure impact; input and output tokens describe model work. Store all three as numeric fields. Do not collapse them into one score.

Field Event role Aggregation rule
elapsedMs Request time consumed before capture Convert to seconds for a duration metric
inputTokens Accumulated input usage Sum by model family, route, and agent step
outputTokens Accumulated output usage Sum separately from input usage
shipmentRef Join key for internal cost allocation Keep out of metric labels
requestId Trace one execution path Search events; never aggregate it

Usage must be accumulated after every successful step. A final-only counter undercharges failed loops, while an increment before the response arrives can charge work the provider never reported. For streamed output, the adapter should finalize usage according to the API's documented response semantics and mark missing usage as unknown. Zero is a real value; unknown is not zero.

This separation also keeps vendor switching practical. The route sees { inputTokens, outputTokens }, while the adapter owns provider-specific response parsing. A finance job can join usage to a dated rate table later. Error capture stays focused on evidence available at failure time.

Ship the failure tests with the endpoint

Before deployment, exercise four paths: a thrown route error, a rejected awaited promise, a rejection outside the Express chain, and an uncaught exception in a child process. Assert that handled request failures return a request ID, include a stack in the event, and contain no prompt or address data. For fatal cases, assert one critical event and a nonzero exit.

Then inspect cardinality. Route templates and agent-step names should remain bounded as shipment volume grows. Verify the shutdown deadline under a disconnected sink, because a reporter that waits forever can defeat process replacement. Finally, sample successful requests separately from errors; error-only latency data exaggerates slow paths and cannot describe the service baseline.

The operational rule is compact: capture close to the request, count usage close to the model call, aggregate only bounded dimensions, and terminate after a process-level failure. This gives a solo operator enough evidence to answer which logistics workflow failed and what work it consumed, without tying the application to an exporter or leaking the shipment itself.

Sources

Top comments (0)