In our (currently) in-house business automation agent, we route prompts to the right agent. This is done using an agent router.
Take-away: before a prompt reaches the active agent, two small local models ask one question: is this the right agent? Either says yes, the prompt stays. Both say no, Kev-4B picks the three best agents to switch to. No data leaves the premises, and a call costs $0.
Why two models
- bge-small (6 ms) is very accurate at catching prompts that don't belong. But it is not good at picking the right agent (in our experiments).
- Kev-4B (local, ~0.12 s) is slightly lower at calling out prompts that don't belong, but good at picking the right target agent: it is in its top 3 for 97% of prompts, against 94% for bge-small.
- Jev is the most accurate on both counts, but it runs over the network and the prompt leaves the premises. It stays as a configurable alternative.
Results *4
150 agent-goals *3, 100 prompts each of three kinds: on-goal (matches the agent's goal), other-goal (belongs to another agent) and out-of-scope (belongs to no agent). A prompt that doesn't match the agent's goal is off-goal, so other-goal and out-of-scope are both off-goal.
- Detection: an off-goal prompt is flagged.
- False alarm: an on-goal prompt is wrongly flagged.
- Right agent for the prompt: an other-goal prompt is sent to the correct agent.
| Option | Detection | False alarm | Right agent for the prompt | Latency *1 | Cost / call |
|---|---|---|---|---|---|
| Jev (top-10 shortlist) | ~100% | 1% | 91% | ~0.46 s (network) | ~$0.00003 |
| Kev-4B *2, flag if P(current) < 0.3 | 99.5% | 8% | 85% (97% in top 3) | ~0.12 s | $0 |
| bge-small, similarity < 0.65 | 100% | 7% | 83% (94% in top 3) | ~0.006 s | $0 |
| bge-small, current goal not in top 5 | 97.5% | 1% | 82% (93% in top 3) | ~0.006 s | $0 |
| Laya *2 (shortlist of 10) | 100% | 19% | 86% | ~0.01 s | $0 |
| Cross-encoder (MiniLM-L6) | 99% | 29% | 75% | ~0.01 s | $0 |
| Cross-encoder (bge-reranker-base) | 99% | 36% | 65% | ~0.02 s | $0 |
Out-of-scope prompts: Kev flags 99 of 100 and correctly shows no suggestion for 80.
*1 Measured on one machine (RTX 5070). Jev's latency includes the network.
*2 Used as they come, not fine-tuned.
*3 Synthetic data (hand-written business prompts) and CLINC150 (a public intent dataset). Not tested in deployment with real customer prompts or usage.
*4 Limited dataset, so the numbers may be optimistic.

Top comments (1)
tr.ee/dev-to