DEV Community

Paul Crinigan
Paul Crinigan

Posted on

How to Log an AI Agent So You Can Actually Debug It

If you have shipped an agent and then tried to explain why it failed on one specific run, this one is for you. It is a practical logging setup to put in place before the first real user touches the agent.

An agent that fails on a fraction of its runs leaves almost nothing behind if all you log is the input and the final answer. The model planned, called tools, read the results and changed its mind along the way, and none of that is on record. Agents are non deterministic, so replaying the input will not reproduce the failure. The only way to debug it later is to have written down every step while it happened.

Start With One Event Schema

Every event the agent emits should share the same core fields: a timestamp with milliseconds, a session id for the task, a step index that orders events within that task, an event type from a fixed list (llm_call, tool_call, tool_response, planning, evaluation, error, completion), and a log level.

Each event type then adds its own payload. An llm_call carries the model, input and output tokens, and latency. A tool_call carries the tool name and its arguments. A completion carries the outcome, total steps, total tokens and total cost. Errors always carry the full payload of the failing step, whatever the log level is set to.

Validate the schema where logs are ingested, even if a missing field only raises a warning. The step by step logging setup lists the full fields for each event type.

Wrap Every LLM Call and Every Tool

Put one wrapper around your LLM client and make it the only way the agent can reach the model. It records the start time, pulls token counts from the response, computes latency and emits the event, so no call can ever go unlogged. At INFO level, log the first and last hundred characters of the prompt plus its length, which is usually enough to identify the prompt version. Save the full prompt and response for DEBUG.

Tools get the same treatment with one important difference: log the call and the response as two separate events. If a tool call starts and its response never arrives, that gap only shows up when the two are logged apart.

Turn the Logs Into a Trace

Once every event carries a session id and a step index, you can assemble a trace for each task: a root span for the whole task, a child span for each LLM call, and the tool calls nested under the LLM call that triggered them. A retry becomes a sibling span under the same parent, so you can see at a glance that the first attempt failed.

What makes agent traces different from microservice traces is the reasoning context. A span that only says the agent made three LLM calls and two tool calls tells you what happened. A span that also records that the first search returned nothing, so the model rewrote the query, tells you why. This walkthrough on tracing agent decisions goes deeper on span design.

Watch Three Numbers Before Anything Else

Teams that try to build complete monitoring before launch often end up launching with none. Three metrics catch the most common ways an agent deployment goes wrong: task success rate, cost per task and error rate. Cost is the one that surprises people, since real traffic often costs two to five times what testing predicted, so a simple daily cost alert is worth setting up on day one.

Add LLM calls per task and tool success rate when one of the three moves and you need to know why. This guide on what to monitor first lays out the order for adding the rest.

The Takeaway

Logging is the layer everything else sits on. Metrics are computed from it, traces are assembled from it, and every investigation ends up querying it. Getting the schema and the wrappers right early costs a day, and skipping it costs every bit of data you would have collected in the meantime. The full guide to AI agent observability covers dashboards, cost tracking and the common failure patterns that monitoring catches.

Top comments (1)

Collapse
 
reidmarlow profile image
Reid Marlow •

Splitting the tool call from the tool response is what saves you during hung subprocesses. If a sandbox command freezes or an external API times out without an error payload, a single combined log event never gets written. When the call and return are recorded as distinct events with their own timestamps, the orphaned start event immediately flags the exact boundary where execution stalled.