DEV Community

Cover image for Part 7: Testing an agentic flow: MUnit, evaluations and the evidence auditors ask for
Shakar Bisetty
Shakar Bisetty

Posted on

Part 7: Testing an agentic flow: MUnit, evaluations and the evidence auditors ask for

Part 7 of 10 · Building an Agentic Change-Approval MVP on MuleSoft

Part 6 covered how the team builds the tools. This part covers how we test them, and how we test the agents that call them. In a regulated change process, "it worked in the demo" isn't evidence. The quality team will ask what was tested, how, against which version, and who signed off.

Test deterministic things deterministically

The single most useful rule we followed: anything that should always give the same answer gets an exact test, and only the LLM steps get scored. That splits the system cleanly, and it means most of the test suite is ordinary integration testing.

A test pyramid for an agentic flow

Layer 1: MUnit for the tool layer

Every MCP tool flow from Part 4 gets MUnit tests, generated by the /new-mcp-tool workflow and finished by hand:

  • The happy path, with the Process API mocked.
  • One test per named error, such as APPROVAL_MISSING and TRANSPORT_LOCKED, checking the error code the agent will see.
  • Schema tests: malformed input is rejected before any call goes out.

The most important test in the whole suite is a negative one. For import_transport_to_qa, we call the Process API with no valid approval and assert two things: the result is APPROVAL_MISSING, and the SAP System API mock was called zero times (MUnit's verify-call with times set to 0). That one test proves the approval gate from the human-in-the-loop post holds in code.

Layer 2: policy tests at the gateway

The Omni Gateway rules from Part 5 are configuration, and configuration drifts. So we test them in the test environment with a small script:

  • For each agent identity, list the tools it can see and compare with the access matrix.
  • For each agent, try one tool it should not have and expect a denial.
  • Send a prompt with planted test PII and confirm the PII policies catch it.

These run after every gateway change, not just at release.

Layer 3: scenario tests on the broker graph

The broker's graph has branches, loops and gates, and each one gets a scenario with fixed inputs:

  • An incomplete request loops back to the requester twice, then escalates to a person.
  • A rejected approval closes the request with a reason and never reaches create_sap_change_doc.
  • A complete, approved request reaches the stage 4 hand-off with a full record.

We're testing the routing here, not the quality of the agents' reasoning, so the inputs are chosen to make the agents' decisions obvious.

Layer 4: evaluations for the agents

This is the only layer where answers vary. We built a golden set of 40 historical change requests, anonymised, with the answers a senior analyst gave: completeness, risk class and the transport sequence.

Each agent version runs against the full set:

  • Structured fields are checked exactly. Risk class must match; the transport sequence must respect every dependency.
  • Free-text output is scored. The A2A Quality Evaluation policy in Omni Gateway sends each response to a judge LLM and returns a yes/no on whether the request was addressed, a 1–3 quality score and an explanation, as response headers and metrics. It doesn't block anything, so we collect the scores and set our own release threshold.

A new prompt or model version only ships if it matches or beats the previous version on the golden set. The scores per version go into the evidence pack.

Layer 5: human sign-off

Before go-live, business owners run a sample of real-shaped requests in the test environment and sign off on the packages the Approval agent builds. This is the same human gate the process has always had, applied to the system itself.

Test data without touching production

All of this runs against a non-production SAP client with synthetic transports, a test ITSM instance and test approvers. Production systems never appear in a test, and the gateway's prompt guard blocks production system names even if someone tries.

The evidence pack

What we hand the quality team for each release:

  1. MUnit results for every tool and Process API, including the zero-call SAP test.
  2. Gateway policy test results against the access matrix.
  3. Graph scenario results.
  4. Golden-set scores for each agent, compared with the previous version.
  5. Signed UAT records from the business owners.

Who owns what

  • Integration team: MUnit and the Process API tests.
  • Platform team: gateway policy tests.
  • Agent team: graph scenarios, the golden set and evaluation thresholds.
  • Quality and business owners: the sign-off.

Next

Part 8 covers deployment: CI/CD for the Mule apps and the agent network, promoting through environments, and how to roll back an agent.

What would you put in your golden set first?


This series describes a reference model built on a fictional company. Product capabilities are based on MuleSoft documentation as of October 2026; check current docs before you build.

Top comments (0)