DEV Community

Cover image for Part 9: Supporting an agentic flow in production: observability, runbooks and failure modes
Shakar Bisetty
Shakar Bisetty

Posted on

Part 9: Supporting an agentic flow in production: observability, runbooks and failure modes

Part 9 of 10 · Building an Agentic Change-Approval MVP on MuleSoft

Part 8 shipped the system. This part is about the day after go-live: what the support team watches, what actually breaks, and what they do about it. An agentic flow fails differently from a normal integration. Most failures don't throw an error. The agent just does something slightly wrong, or nothing at all.

One trace per change request

The first decision paid for itself many times over: the change request number is the correlation ID, end to end. It's set when the ITSM ticket arrives and passed through the broker, every agent, every MCP tool call, the Process APIs and the SAP System API.

That means a support engineer can answer "what happened to CR-1042?" with a single search across the logs, instead of piecing together five applications by timestamp. Every log line is structured JSON with the CR number, the agent name, the tool name and the step in the broker's graph.

Six failure modes, and how we catch each one

Watch four groups of signals

We kept the production dashboard deliberately small. Four groups, one row each:

  • Flow health. Requests in flight, time spent in each stage of the broker's graph, and the escalation rate to a person. A rising escalation rate is usually the first sign that something upstream changed.
  • Tool layer. Calls per tool and named errors per tool. A spike in APPROVAL_MISSING is worth a look: it often means approvers are being chased before the approval request was sent.
  • Gateway. Denials per agent, PII detections and rate-limit responses from the Omni Gateway policies in Part 5.
  • LLM. Tokens per request per agent, and the trend of the quality scores from the judge LLM introduced in Part 7.

Each alert on this dashboard points to one of the failure modes below, and each one has a runbook.

The six failure modes

1. The agent loops. The same tool is called with the same input three or more times for one request. The broker's graph caps retries at two and escalates to a person, so the request doesn't get stuck. The alert exists so someone finds out why: usually a tool description that doesn't say what to do after an error.

2. A tool isn't found. After a deployment, an agent reports it can't complete a step, and the trace shows zero tool calls for that stage. This is almost always deployment order (agents deployed before their tools, as described in Part 8). We now run a smoke check after every deployment that lists each agent's tools and compares them with the access matrix.

3. The token budget is hit. The LLM rate-limit policy returns 429s for one agent. The request waits and resumes, so nothing is lost. The real question is why: in our case it was a prompt change that pulled the full transport history into every call.

4. An SAP transport is locked. TRANSPORT_LOCKED errors come back from SAP. The Process API retries with backoff. The agent never retries on its own, because an agent retrying a locked import is how you get duplicate imports.

5. Quality drifts. The judge scores fall slowly over a few weeks with no change on our side. This can happen when the model provider updates a model behind the same name. We rerun the golden set, pin a specific model version, and treat any model change as a release, as in Part 8.

6. Policy denials spike. One agent suddenly gets 403s from the gateway. We compare its live tool list with the access matrix and check the most recent policy change. So far it has always been configuration, never an agent misbehaving.

Runbooks with four sections

Every runbook has the same shape, so the person on call doesn't need to learn a new format at 2 a.m.:

  1. Symptom: what the alert or the user report looks like.
  2. Check: the exact log search or dashboard panel, filtered by CR number.
  3. Act: the safe actions, in order. Resubmit the request, reassign it to a person, or pause intake.
  4. Escalate: who owns it. The integration team for tools, the agent team for prompts and evaluations, the platform team for gateways.

The runbook for "the agent loops" is the most used. The fix is almost never in the agent's prompt. It's in the tool description, or in a named error that the agent can't act on.

Pause intake: the kill switch

Every agentic system needs a way to stop safely. Ours is a single switch that stops the broker from accepting new requests. Requests already in flight finish their current step and park. New tickets go to the manual queue.

This only works because the manual process never went away. It's documented, and the change analysts still know it. We tested the switch in production in a planned window before go-live, and again after the first month.

Who supports what

  • Service desk (L1): request status by CR number, and the pause switch.
  • Integration team (L2): MCP tools, Process and System APIs, SAP errors.
  • Agent team (L2): prompts, the broker's graph, golden-set reruns.
  • Platform team: Omni Gateway, Agent Fabric, credentials and deployments.

Once a week, the agent team reviews ten completed requests by hand, the judge score trend and the token spend per agent. It takes an hour and it has caught every slow drift so far.

Next

Part 10, the last part, covers the lessons learned and a reusable blueprint you can take to your own change process.

What's the first runbook you'd write for an agent in production?


This series describes a reference model built on a fictional company. Product capabilities are based on MuleSoft documentation as of October 2026; check current docs before you build.

Top comments (0)