Everyone has been stuck with a support bot that keeps saying "Sorry, I didn't understand that" while you get angrier and angrier.
I wanted to build the opposite: a bot that answers what it can, and quietly steps aside when a real person is needed. It works on web, Discord, and WhatsApp, and the code is open source: github.com/Shahzaib30/ai-support-agent
The Big Picture
The system has four layers. Customers talk to it from any channel, one FastAPI backend handles everything, a few services store state and send alerts, and monitoring watches it all.
Inside the FastAPI core there are three gates in front of the answer engine: a HITL gate (human-in-the-loop, which checks whether a human is already handling the chat), an explicit escalation check, and a sentiment escalation check. PostgreSQL remembers conversations, Redis caches answers, and everything runs with one Docker Compose command.
How the Bot Finds Answers
Most RAG bots search once and hope. Mine runs every question through six stages:
- The question arrives.
- A query condenser rewrites it. Customers say things like "what about sale items?", which means nothing alone. A DeepSeek step turns follow-ups into standalone questions first.
- Two searches run at once. FAISS with BGE-small embeddings finds passages by meaning. BM25 finds passages by exact words. Each returns its top 20.
- Reciprocal Rank Fusion merges them into one list of 20 candidates.
- A cross-encoder reranks them. It reads the question and each passage together, and only the top 5 move on.
- DeepSeek writes the answer using those 5 chunks, the chat history, and long-term memory.
Why two searches? Vector search understands meaning but is weak on exact things like prices, emails, and phone numbers. Ask "What does the Pro plan cost?" and it may return a nice passage about pricing in general instead of the one with the number. BM25 is the opposite. Together they cover each other's weak spots, and for questions about fixed facts, hybrid clearly beat vector search alone.
When Does the Bot Give Up?
There are two triggers, and I built them differently on purpose.
1. The customer is clearly upset (checked by an LLM)
Every message goes to the DeepSeek API with a strict prompt that replies in JSON, a label and a score from -1.0 to 1.0:
Rules:
- negative: customer is CLEARLY angry, frustrated, or complaining
- neutral: casual phrases, short responses, confusion, greetings
- positive: happy, satisfied, thankful
"oh no" → neutral
"wtf" → neutral (not clearly angry enough)
"this is terrible I want a refund NOW" → negative
Reply ONLY with JSON: {"label": ..., "score": ...}
The examples are what make it work. People say "oh no" and "wtf" casually, so those count as neutral. Only clear anger counts as negative.
One angry message doesn't trigger anything. The bot escalates after three negative messages in a row, so a customer who calms down isn't punished for an early bad moment.
2. The customer asks for a human (plain code, no LLM)
def wants_human(text: str) -> bool:
text = text.lower()
return any(p in text for p in HUMAN_REQUEST_PHRASES)
# "talk to a human", "real person", "live agent", ...
If the message contains a phrase like "talk to a human" or "real person," the bot escalates immediately. An API call to understand "I want a human" would cost money and time for no gain, so a simple phrase list does the job for free.
That became my main rule: use an LLM when you need judgment, and plain code when the rule is clear.
One Conversation, Three States
Before answering anything, the HITL gate checks what state the conversation is in:
- AI Active: normal RAG flow, the bot answers.
- Human Pending: an escalation was sent but no human has replied yet. If nobody answers within 5 minutes, the bot takes the chat back so the customer isn't left hanging.
- Human Active: a human is handling it, so the bot stays quiet and forwards the customer's messages to the agent. The customer never gets two answers.
When the case is resolved, the customer goes straight back to chatting with the AI.
The Slack Side
Sending the alert was the easy part. When either trigger fires, the bot posts to an #escalations channel in Slack with the customer, the channel, the reason, the sentiment score, and the conversation ID:
A human replies in the thread, and no special format is needed. The customer's new messages also show up in that thread, so the whole conversation lives in one place. To hand the chat back to the bot, the human types resolved.
The Hard Half: Getting the Reply Back to the Customer
Reading a Slack reply and delivering it to someone on Discord was the tricky part. I solved it with an n8n workflow:
- Slack Trigger fires on a new message in the thread.
- Is real user reply checks it's a person and not the bot's own message.
- Extract conversation_id and message pulls out the two things the API needs.
-
Is resolved branches:
- Yes: an API call resolves the conversation in PostgreSQL.
-
No: an API call (
/human_reply) stores the reply.
On the Discord side, a small bot checks the API every 5 seconds and DMs any new human reply to the customer. Here's what the customer sees:
, answers took about 1.4 to 2.4 seconds end to end. Those numbers come from my own test conversations, not real customers.
What I Learned
- Not everything needs an LLM. Use it for judgment, like reading sentiment. Use plain code for clear rules, like phrase matching.
- Tracing saves hours. LangSmith showed me what each request actually did, so I stopped guessing where things broke.
- Connecting tools is most of the work. Slack, Discord, n8n, and the database each have their own quirks, and the real effort goes into making them talk to each other.
- Use n8n when you need a fast integration. It makes tricky steps easy and keeps the workflow readable for other people. Skip it when you need deep customization, low latency, or custom logic. For those, write code.
Try It or Hire Me
- Code: github.com/Shahzaib30/ai-support-agent
- Upwork: My Upwork profile
- Linkedin: My Linkedin profile
I'm a Full Stack AI Engineer who builds AI agents, RAG systems, and automation. If you're building something similar, send me a message.
What's the worst support bot you've ever dealt with? Tell me in the comments.
Top comments (0)