Every agency says "we build AI chatbots."
Very few show numbers.
This is the real story of how we built an AI chatbot that today handles 60% of our clients' customer queries โ fully automated, 24/7, zero human intervention for most conversations.
I'm Umaar Ahmed, CEO at GTSol360. Here's the architecture, the code, and the lessons that cost us months.
|
|
|
๐ฏ The Client Problem
One e-commerce client was drowning:
- 12,000+ support tickets/month
- 3 full-time agents working 9-6
- 4-hour average response time
- 47% cart abandonment on pre-sale questions
Hiring 5 more agents = ~$15K/month. Still no night or weekend coverage.
They asked: "Can AI handle this?"
I said yes. We had 6 weeks.
โ Why Most AI Chatbots Fail
They're built as glorified FAQs.
You ask a real question, they reply:
"I'm sorry, I don't understand. Would you like to speak to a human?"
That's not a chatbot โ that's a fancy contact form.
The core problem: keyword matching. If the user's exact phrase isn't in training data, the bot dies.
Our v1 had 400 predefined Q&As. Customer satisfaction: 34%. Worse than no bot.
We threw it away and rebuilt.
๐๏ธ The Architecture That Works
USER MESSAGE
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ 1. Intent Classifier โ โ support / sales / general / human
โ (GPT-4o-mini) โ
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโ
โ
โโโโโโโโโดโโโโโโโโโ
โผ โผ
โโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ
โ 2a. RAG โ โ 2b. Human โ
โ Pipeline โ โ Handoff โ
โ (pgvector) โ โ (Slack) โ
โโโโโโโโฌโโโโโโโ โโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ 3. Response Generator โ
โ (GPT-4o) โ
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ 4. Guardrails โ โ PII, sentiment, legal triggers
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโ
โผ
RESPONSE
Four layers. Let me break them down.
๐ Layer 1: Intent Classification
First job: know what NOT to answer. Some messages need a human. Some need sales.
const INTENT_PROMPT = `
Classify the user message into ONE category:
SUPPORT | SALES | GENERAL | HUMAN
Respond with only the category name.
Message: "{message}"
`;
async function classifyIntent(message: string) {
const res = await openai.chat.completions.create({
model: "gpt-4o-mini",
messages: [
{ role: "user", content: INTENT_PROMPT.replace("{message}", message) }
],
temperature: 0,
max_tokens: 10,
});
return res.choices[0].message.content;
}
Cost: ~$0.0001 per message. Saves us from hallucinating on 40% of queries.
๐ Layer 2: RAG Pipeline
Instead of training on your data (slow, expensive), we retrieve relevant docs at query time and pass them to the AI.
We use Supabase pgvector โ same DB, no new infra.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE documents (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
content TEXT NOT NULL,
embedding VECTOR(1536),
source TEXT,
created_at TIMESTAMPTZ DEFAULT NOW()
);
CREATE OR REPLACE FUNCTION match_documents(
query_embedding VECTOR(1536),
match_threshold FLOAT,
match_count INT
)
RETURNS TABLE (id UUID, content TEXT, similarity FLOAT)
LANGUAGE plpgsql AS $$
BEGIN
RETURN QUERY
SELECT
documents.id,
documents.content,
1 - (documents.embedding <=> query_embedding) AS similarity
FROM documents
WHERE 1 - (documents.embedding <=> query_embedding) > match_threshold
ORDER BY documents.embedding <=> query_embedding
LIMIT match_count;
END;
$$;
Chunk size matters: 500-token chunks with 100-token overlap. Too big = diluted context. Too small = missing nuance.
โก Layer 3: Response Generation
Now we have:
- User's question
- 5 most relevant docs
- Conversation history
const SYSTEM_PROMPT = `
You are a support assistant for {company}.
RULES:
1. Answer ONLY from the provided context
2. If context lacks the answer, say so honestly
3. Be concise โ 2-3 sentences
4. Never make up policies, prices, or promises
CONTEXT:
{context}
`;
async function generateResponse(
message: string,
context: string[],
history: Message[]
) {
const res = await openai.chat.completions.create({
model: "gpt-4o",
messages: [
{ role: "system", content: SYSTEM_PROMPT },
...history,
{ role: "user", content: message },
],
temperature: 0.3,
});
return res.choices[0].message.content;
}
Why GPT-4o, not 3.5: Quality dropped ~30% on 3.5. Client churn risk was higher than the cost savings.
๐ก๏ธ Layer 4: Guardrails
Most agencies skip this. Then get burned.
function applyGuardrails(response: string, userMessage: string) {
// PII redaction
response = response.replace(/\b\d{16}\b/g, "[REDACTED]");
// Legal triggers
const legal = ["sue", "lawyer", "refund", "legal"];
if (legal.some(t => userMessage.toLowerCase().includes(t))) {
return { response, escalate: true, reason: "LEGAL" };
}
// Angry customers
if (analyzeSentiment(userMessage) < -0.5) {
return { response, escalate: true, reason: "ANGRY" };
}
return { response, escalate: false };
}
We shipped without guardrails once. Day 3, the bot promised a discount that didn't exist. Cost: $2,400.
Never again.
๐ The Numbers After 8 Months
Rolled out to 4 clients. Real results:
| Metric | Before | After |
|---|---|---|
| Response time | 4 hours | 2.3 seconds |
| Auto-resolved tickets | 0% | 60% |
| Support staff | 3 FTE | 1 FTE |
| Customer satisfaction | 3.2/5 | 4.6/5 |
| After-hours coverage | None | 24/7 |
| Monthly AI cost | $0 | $340 |
Client retention: 67% โ 94%.
They referred us 3 more clients.
๐ก 5 Lessons That Cost Us Months
1. Chunking > Model Choice
We spent 6 weeks testing models. Real upgrade came from fixing chunk size (2000 โ 500 tokens).
2. Retrieval Threshold > Count
We fetched 20 docs initially. Diluted context. Now: top 5, similarity > 0.75.
3. Guardrails Are Not Optional
Ship without them once, you'll never ship without them again.
4. Store History Properly
Last 10 messages in Postgres per conversation_id. Context without token bloat.
5. Evaluations Beat Vibes
200-query test suite. Every change gets evaluated. Regression caught instantly.
๐ ๏ธ The Stack We Shipped
Backend: Next.js 15 ยท TypeScript ยท Supabase (Postgres + pgvector + RLS)
AI: OpenAI GPT-4o ยท GPT-4o-mini ยท text-embedding-3-small
Infra: Vercel Edge ยท Cloudflare ยท Upstash Redis
Observability: Sentry ยท Vercel Analytics
No exotic tools. Everything battle-tested.
๐ What's Next
v2 in progress:
- Voice AI for phone support (Whisper + ElevenLabs)
- Multi-language (Urdu, Arabic, Spanish)
- Agentic workflows โ not just answering, but acting
๐ค If You're Building This
Building a production AI chatbot is not a weekend project. It's also not 6 months. It's 6-8 weeks of focused work if you know what you're doing.
We build these at GTSol360 for clients worldwide.
Book a free consultation โ โ no pitch, just a real conversation.
I'm Umaar Ahmed, CEO at GTSol360. We build AI chatbots, web platforms, and mobile apps. Follow me on LinkedIn and dev.to.
Top comments (0)