DEV Community

Cover image for Not every step should be an agent: a 2x2 for deciding what to automate
Shakar Bisetty
Shakar Bisetty

Posted on

Not every step should be an agent: a 2x2 for deciding what to automate

Companion to Part 1 · Building an Agentic Change-Approval MVP on MuleSoft

The fastest way to over-engineer an agentic system is to make every step an agent.

In Part 1 I introduced the use case for this series: automating the approval and promotion of SAP changes from lower to upper environments in a regulated company. The process has about 21 steps across 11 stages. Before designing anything, we had to decide which of those steps an agent should touch at all.

This is the tool we used. It's simple enough to run on a whiteboard in 30 minutes with the process owners.

Two questions per step

For every step in the process, ask:

  1. Does it follow rules, or does it need judgement? If you can write the logic as an if-statement or a mapping, it follows rules. If a person reads something messy and decides, it needs judgement.
  2. If it goes wrong, how bad is it, and can we undo it? Drafting a document badly is cheap to fix. Importing the wrong transport into production is not.

Plot each step on those two axes and you get four quadrants.

Not every step should be an agent

The four answers

Rules, low risk: integration tool, fully automated.
This is an MCP tool over a Mule API, with no LLM in the path. Creating the change record, reading transport objects and fetching import logs all live here. They're deterministic, fast, cheap and easy to audit. An agent can call these tools, but the tool itself doesn't reason.

Judgement, low risk: the agent acts and a human spot-checks.
This is where agents earn their keep. Is the change request complete? Does the description match the transport contents? What do these import logs actually mean for the approver? Mistakes here are visible and easy to correct, so the agent can act and a person reviews samples.

Rules, high risk: still a tool, but only after approval.
Importing to production is a deterministic call, so it shouldn't involve an LLM. But it's hard to undo, so the tool only runs once an approval is on record. In a MuleSoft setup that's enforced in the integration layer and at the gateway, not by asking the agent nicely in a prompt.

Judgement, high risk: a human decides, the agent prepares.
Business approval, validation review and the CAB decision stay with people. The agent's job is to hand them a complete package with a summary, the evidence and the risks, so the decision is fast and well-informed.

What we learned

When we sorted our process, more steps landed in the "tool" quadrants than in the "agent" quadrants. That surprised some of the team, but it's a good sign:

  • Deterministic steps are cheaper to run and easier to test.
  • Auditors understand them.
  • Every step that doesn't need an LLM is one less thing that can hallucinate.

It also split the work cleanly between teams. The integration team owns the tool quadrants: MCP tools, Mule APIs, and SAP and ITSM connectivity. The agent team owns the judgement quadrants. The platform team owns the guardrails that keep the high-risk row behind approvals.

The borderline cases

Some steps don't sort cleanly. "Analyse import logs" is the one we argued about most. You could write rules for known return codes, but real logs are messy and the value is in explaining them to a non-technical approver. We put it with the agent, with the raw return codes still passed through deterministically.

Where would you put it? I'd like to hear the case for a rules engine.

Next

In Part 2 I'll put the whole reference architecture on one page: Agent Fabric, Omni Gateway, MCP tools and the SAP and ITSM back ends, colour-coded by which team owns what.


This series describes a reference model built on a fictional company. Product capabilities are based on MuleSoft documentation as of September 2026.

Top comments (0)