DEV Community

Corpable
Corpable

Posted on

Designing a Human-in-the-Loop Approval Matrix for an AI Agent That Drafts B2B Quote Replies

If you let an LLM agent draft replies to inbound B2B quote requests (RFQs), the hard part is not the prompt. It is deciding, in code, which drafts may go out on their own, which need a human click, and which the agent must never write at all.

This post walks through the approval matrix pattern we use with small export sales desks: a declarative policy file, a field-level classifier, and a gate function that sits between the model and the outbox. The examples are in Python and YAML, but the pattern is tool-agnostic.

The failure mode we are designing against

An RFQ reply looks like prose, but it is really a bundle of commitments: a unit price, a quantity tier, a lead time, an Incoterm (FOB, CIF, DDP...), payment terms, maybe a certification claim. A model asked to "be helpful" will happily fill gaps. It will match the buyer's requested Incoterm, round a lead time down, or say "yes, we have CE" because that sounds like a good answer.

None of those are hallucinations in the classic sense. They are unauthorized commitments. So the control is not "make the model more accurate". It is "make sure the model can't commit to anything nobody approved".

Step 1: Classify the fields, not the message

Approving or rejecting a whole message is too coarse. A reply that confirms a catalogue link and asks two qualifying questions is low-risk. The same reply with one extra sentence ("we can do net-60") is not.

So we tag each commercial field with one of three tiers:

  • auto: the agent may draft and send, as long as the value comes from an approved source (catalogue, price book, FAQ).
  • review: the agent may draft, but a named human approves before sending.
  • human_only: the agent must not produce a value at all. It writes a holding reply and escalates.

Step 2: Put the matrix in a policy file

Keep the policy out of the prompt. Prompts drift; config files get code review.

# rfq_policy.yaml
version: 3
owners:
  sales_lead: "sales@desk.example"
  ops: "ops@desk.example"
  finance: "finance@desk.example"

fields:
  acknowledgement:   { tier: auto }
  catalogue_link:    { tier: auto }
  qualifying_question: { tier: auto }
  moq:               { tier: auto,   source: catalogue }
  lead_time_range:   { tier: auto,   source: catalogue }
  unit_price:        { tier: review, source: price_book, approver: sales_lead }
  quantity_tiers:    { tier: review, source: price_book, approver: sales_lead }
  incoterm:          { tier: review, source: sku_default, approver: ops,
                       allowed: [EXW, FOB, FCA] }
  sample_terms:      { tier: review, approver: sales_lead }
  certification:     { tier: review, source: cert_registry, approver: ops }
  payment_terms:     { tier: human_only, approver: finance }
  discount:          { tier: human_only, approver: sales_lead }
  exclusivity:       { tier: human_only, approver: sales_lead }
  complaint_or_claim: { tier: human_only, approver: sales_lead }

hard_stops:
  - off_platform_payment_request
  - bank_detail_change
  - ship_before_payment
Enter fullscreen mode Exit fullscreen mode

Two details matter here. Every non-auto field has a named approver, so escalations don't land in a shared void. And source declares where an allowed value may come from. A value with no matching source record is treated as invented, even if it happens to be correct.

Step 3: Have the model emit structure, then render prose

Instead of asking the model for a finished email, ask for a JSON object: the fields it wants to assert plus the prose around them. Most current LLM APIs support schema-constrained output, which makes this reliable.

{
  "fields": {
    "moq": {"value": "500 pcs", "source_ref": "catalogue:SKU-1182"},
    "incoterm": {"value": "CIF Hamburg", "source_ref": null}
  },
  "body_template": "Thanks for your RFQ for SKU-1182. MOQ is {moq}. ..."
}
Enter fullscreen mode Exit fullscreen mode

Now the gate can reason about exactly what is being promised.

Step 4: The gate

from dataclasses import dataclass, field

TIER_ORDER = {"auto": 0, "review": 1, "human_only": 2}

@dataclass
class Decision:
    action: str                 # "send" | "queue_review" | "escalate" | "block"
    reasons: list = field(default_factory=list)
    approvers: set = field(default_factory=set)

def gate(draft: dict, policy: dict, signals: set) -> Decision:
    stops = signals & set(policy["hard_stops"])
    if stops:
        return Decision("block", [f"hard stop: {s}" for s in sorted(stops)])

    worst, d = "auto", Decision("send")
    for name, f in draft["fields"].items():
        rule = policy["fields"].get(name)
        if rule is None:
            # unknown field = the model invented a commitment type
            rule = {"tier": "human_only", "approver": "sales_lead"}
            d.reasons.append(f"unknown field {name}")
        tier = rule["tier"]
        if rule.get("source") and not f.get("source_ref"):
            tier = max(tier, "review", key=TIER_ORDER.get)
            d.reasons.append(f"{name}: no source record")
        allowed = rule.get("allowed")
        if allowed and f["value"].split()[0] not in allowed:
            tier = "human_only"
            d.reasons.append(f"{name}: {f['value']} not in approved menu")
        if tier != "auto":
            d.approvers.add(rule["approver"])
        worst = max(worst, tier, key=TIER_ORDER.get)

    d.action = {"auto": "send", "review": "queue_review",
                "human_only": "escalate"}[worst]
    return d
Enter fullscreen mode Exit fullscreen mode

With the sample draft above, incoterm has no source and CIF isn't in the approved menu, so the decision is escalate with ops as approver. The agent then sends only a holding reply: acknowledge, confirm MOQ, say the shipping terms will follow from a colleague.

A few design choices worth copying:

  • Unknown fields fail closed. If the model starts asserting something your schema doesn't know about, that's a policy gap, not a pass.
  • Missing provenance escalates one tier. That one rule catches most "confident but made-up" values.
  • Hard stops run first and are driven by a separate classifier (or plain rules) on the inbound message, not by the drafting model.

Step 5: Log decisions, measure the right things

Store every gate decision with the policy version, the draft, the approver, and the edit diff if a human changed something. After a few weeks you can answer useful questions:

  • Which review fields are approved unchanged 95%+ of the time? Those are candidates to promote to auto.
  • Which fields get edited most? That usually means the source data (price book, lead times) is stale, not the model.
  • What is the median first-reply time for qualified RFQs? Reply volume on its own tells you almost nothing.

Promotion between tiers should be a reviewed change to rfq_policy.yaml, never a prompt tweak.

Where this fits with off-the-shelf agents

Packaged agents for marketplace sellers increasingly ship with their own approval settings. For example, Alibaba.com's Accio Work agent can handle first-draft reception for a store, and we've written up how teams configure Accio Work with human approval steps. Even then, I'd keep an explicit matrix like the one above as the source of truth and map the vendor settings onto it, so the policy survives a tool change.

If you're building this from scratch, start with the data before the model. A short pre-launch checklist for RFQ desks (SKU defaults, Incoterm menu, cert registry, named approvers) will save you more incidents than any prompt engineering.

TL;DR

  • Classify fields, not messages: auto, review, or human_only.
  • Keep the matrix in a versioned config file with named approvers and required sources.
  • Make the model emit structured fields, and gate them before rendering prose.
  • Fail closed on unknown fields and missing provenance.
  • Promote fields to auto based on logged approval data, through code review.

How are you gating agent output in your workflows? I'd be curious whether others use field-level policies or something coarser.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

Official Platform Update

Security protocols have been updated for all developer accounts.

  • tr.ee/dev-to