One AI agent answering every support ticket is a good demo. We know because we built that first, showed it around internally, and got exactly the reaction you'd expect — "cool, when's it live." Then we actually thought about what "live" meant. Billing questions need real invoice data. Technical questions need product docs and sometimes an actual diagnostic call. And anything involving a refund above a certain amount needs a human to look at it before it goes out, because we were not comfortable with a language model unilaterally deciding to give someone their money back.
Cramming all of that into one system prompt and one model call gets you something mediocre at everything instead of good at anything. So we didn't. Here's the actual build — three specialized agents, a shared session, and a human approval gate that we do not let anything skip.
What we're building
customer message
│
▼
[router agent] ──reads session memory, classifies──┐
│ │
┌───┴────┬─────────────┐ │
▼ ▼ ▼ │
[billing] [technical] [escalation] ──always flags for human review
│ │ │
└─────────┴─────────────┘
│
▼
draft response + needsHumanApproval flag
│
┌────────┴────────┐
▼ ▼
auto-send held for review
(routine) (sensitive)
One shared session, three specialists, one gate that anything sensitive has to pass through before a customer ever sees it.
One ticket, one shared session, three specialized agents, and a human who signs off before anything sensitive ships.
Defining the agents
DNotifier's agent primitive is DNotifier.defineAgent({ name, model, run(ctx) }). The ctx your run function gets handed gives you the shared workflow state, this step's input, and a sendAI() scoped to whatever model you put on that specific agent.
import { DNotifier } from "@dnotifier-realtime/dnotifier";
const routerAgent = DNotifier.defineAgent({
name: "router-agent",
model: "gpt-4o-mini", // cheap and fast — this runs on every single message
async run(ctx) {
const classification = await ctx.sendAI({
message: {
text: `Classify this support message into one category:
billing, technical, or escalation. Message: "${ctx.input.message}"
Respond with just the category word.`,
},
});
ctx.state.category = classification.text.trim().toLowerCase();
return ctx.state.category;
},
});
const billingAgent = DNotifier.defineAgent({
name: "billing-agent",
model: "gpt-4o",
async run(ctx) {
const answer = await ctx.sendAI({ message: { text: ctx.input.message } });
ctx.state.draftResponse = answer.text;
return answer.text;
},
});
const technicalAgent = DNotifier.defineAgent({
name: "technical-agent",
model: "gpt-4o",
async run(ctx) {
const answer = await ctx.sendAI({ message: { text: ctx.input.message } });
ctx.state.draftResponse = answer.text;
return answer.text;
},
});
const escalationAgent = DNotifier.defineAgent({
name: "escalation-agent",
model: "gpt-4o",
async run(ctx) {
const draft = await ctx.sendAI({
message: {
text: `Draft a careful, empathetic response to this escalated
issue, and flag it for human review before sending:
"${ctx.input.message}"`,
},
});
ctx.state.draftResponse = draft.text;
ctx.state.needsHumanApproval = true;
return draft.text;
},
});
The model choice per agent isn't an afterthought — it's the whole point. The router runs on gpt-4o-mini because deciding "is this a billing question or a technical one" doesn't need a flagship model, and it's going to run on literally every incoming message, so it had better be cheap. The specialists run on gpt-4o because they're doing the reasoning a customer will actually judge. If we ran everything through the expensive model just because the escalation path needed it, we'd be paying flagship prices for "what are your support hours" questions all day.
Tying it into a workflow
const supportWorkflow = new DNotifier.Workflow({
name: "customer-support-pipeline",
description: "Routes and resolves inbound support tickets across three specialist agents",
observability: true,
async entry(ctx) {
const category = await ctx.agents.run(routerAgent, { message: ctx.input.message });
let result;
if (category === "billing") {
result = await ctx.agents.run(billingAgent, { message: ctx.input.message });
} else if (category === "technical") {
result = await ctx.agents.run(technicalAgent, { message: ctx.input.message });
} else {
result = await ctx.agents.run(escalationAgent, { message: ctx.input.message });
}
return {
category,
response: result,
needsHumanApproval: ctx.state.needsHumanApproval ?? false,
};
},
});
supportWorkflow.registerAgents([routerAgent, billingAgent, technicalAgent, escalationAgent]);
observability: true is one line and it's saved us more debugging time than almost anything else in this system. The first time someone from support Slacked us "the bot gave a weird answer on ticket #4471," we didn't reconstruct anything from a transcript. We opened the run, saw exactly which agent handled it, on which model, with what input, and had an answer back in under a minute.
Running it against a real ticket
const notifier = new DNotifier({
appId: process.env.DNOTIFIER_APP_ID,
secret: process.env.DNOTIFIER_SECRET,
userId: "support-system",
transport: "ws",
WebSocketImpl: WebSocket,
});
await notifier.connect();
const outcome = await notifier.runWorkflow(supportWorkflow, {
message: "I was charged twice for my subscription this month.",
sessionId: "ticket-88213",
});
console.log(outcome);
// { category: 'billing', response: '...', needsHumanApproval: false }
Routine billing question, routes to the billing agent, resolves without anyone in the loop. Now the same customer, same ticket, a follow-up:
const followUp = await notifier.runWorkflow(supportWorkflow, {
message: "This is the third time this has happened and I want to cancel and get a full refund for the year.",
sessionId: "ticket-88213",
});
// { category: 'escalation', response: '...', needsHumanApproval: true }
Because it's the same sessionId, the escalation agent isn't starting from zero — it already has the prior turn's context about the duplicate charge, without our application code having to reassemble and re-pass that history manually. That part genuinely surprised us the first time we tested it; we expected to have to wire up our own history-stitching logic and didn't need to.
The part we actually care about most: the approval gate
needsHumanApproval coming back true is a hard stop, not a suggestion. This is where the drafted response gets pushed into whatever review surface support already uses — for us that's a Slack channel with an approve/reject button — and the "send to customer" function only fires once a human actually clicks approve.
if (outcome.needsHumanApproval) {
await notifyReviewQueue({
ticketId: "ticket-88213",
draft: outcome.response,
});
// held here. nothing goes out until a human approves it.
} else {
await sendToCustomer(outcome.response);
}
We'll be blunt about why this matters more than it might look on the page: a system that can draft a refund confirmation and a system that can send one without anyone checking it first are very different systems from a risk standpoint. "Which of our agents can take an irreversible action without a human in the loop" is a question we wanted a deliberate answer to, not one we backed into by accident because the demo worked and nobody circled back to add the guardrail.
Why three agents instead of one really good prompt
We got this pushback internally more than once, and it's a fair question — doesn't one sufficiently detailed prompt handle all of this? For a while it kind of does. It falls apart for a few concrete reasons once you're past the prototype stage.
The knowledge base gets muddier fast. A billing agent grounded only in invoice and pricing docs gives sharper answers than one generalist searching across billing, technical, and policy docs at once — less noise for the retrieval step to sort through.
Cost stops making sense. Routing "what's your refund policy" through the same expensive model handling genuine escalations is money spent for zero quality gain on the easy majority of traffic.
Debugging gets genuinely harder. One sprawling prompt that goes wrong means untangling a single enormous set of instructions. Three focused agents means the observability dashboard tells you exactly which one misfired, every time.
And the approval logic gets fuzzy. "Escalate to a human if the refund's over $500" is a clean, auditable rule sitting in a dedicated escalation agent. It's a rule that's a lot easier to lose track of buried inside one long prompt that's also handling billing lookups and troubleshooting steps.
None of this is specific to support tickets, either. The same router-then-specialists-then-shared-state shape, with an optional human gate bolted on, shows up in content review pipelines and internal ops tooling we've built since. Support was just the clearest place to build it first and see if the pattern actually held up under real traffic. It did.
If you're staring down a single monolithic support prompt that's slowly turning into a maintenance headache, this is the refactor. It took us about a day to build the version above, and the approval gate alone was worth it the first week — it caught a draft response none of us would have wanted to send automatically.
Top comments (0)