A support chatbot is only as good as its knowledge base, and a knowledge base is only as good as its seed corpus. The good news: you don't need hundreds of entries before go-live. Twenty to thirty well-chosen entries can absorb most repetitive questions on day one, and the gaps get filled from miss logs afterwards. Cold start is a selection problem, not a volume problem.
Three source types
Chat logs. Export the last month of real conversations, group them by question type, and the top-frequency clusters are your first entries: pricing, shipping time, courier, discounts, returns, invoicing. Real phrasing beats invented phrasing every time.
After-sales policy. This lives in your shop rules, not in chat logs: when the 7-day return window starts, who pays return shipping, what happens with damaged goods, refund timelines. These entries must be written as fixed policy — never let the model improvise, because money and liability are involved.
Product parameters. Size, material, color, compatibility, care instructions. Pattern: many ways to ask, one short answer. Store them as structured entries — one parameter with a group of synonym questions attached.
How many entries for day one
Twenty to thirty, chosen by frequency, not by coverage. Sort the last month's conversations by question type, take the top clusters, and stop when you hit the long tail — the tail is low-frequency and cheap to handle manually.
Make sure the first batch includes returns, refunds, and shipping-time questions. They're both the most frequent and the most dispute-prone, and a system answering them with one consistent policy beats three agents answering three different ways.
Anatomy of one entry
Each entry has four parts:
- The question, in the buyer's words. "When will my order ship?" — not "Logistics SLA overview."
- The answer, conclusion first. "Orders before 16:00 ship today; after that, next business day" — conditions after the verdict.
- A keyword group. All common phrasings of the same question: how long, when does it ship, how many days, where's my tracking number.
- A boundary. When this entry does NOT apply and a human should take over. "Custom orders never qualify for same-day shipping — always hand off for scheduling."
Without the boundary, the bot will answer in scenes where it shouldn't. That's the failure mode users remember.
Two underrated details
Synonyms outnumber canonical phrasings. For "shipping" alone, buyers write: how long does it take, has it shipped, where's my tracking number, why so slow. Organize the corpus by topic, not by sentence — one topic, one group of phrasings.
Negative phrasings need entries too. "Never mind", "don't ship it", "this isn't a refund request" differ from positive questions by a character or two and are classic false triggers. Put them into exclusion terms during seeding — far cheaper than firefighting after launch.
The cadence after launch
- Weekly: pick entries from the miss log — questions that didn't match, and matches that answered wrong.
- Biweekly: prune. Expired campaigns, changed specs, updated policy. Stale content is worse than missing content.
- Monthly: regression-test the seed entries against current rules.
Do that and the knowledge base becomes a corpus that keeps getting more accurate, instead of a document that goes stale the week after launch. We run our own support-automation stack this way — the trigger-rules setup and the missed-message fallback are both written up in detail, including how we log whether each reply came from the knowledge base or a fallback template.
Top comments (0)