I ran 65 questions against our own knowledge base and checked which ones the retrieved text actually answered. The similarity scores overlapped too much for any threshold to work.
Our own handoff docs used to describe it: if retrieval confidence falls below a threshold, hand the visitor to a human instead of generating a weak answer. The default was 0.4. It's the design most people building a RAG chatbot reach for, and I no longer think it works. This post is about why: first a hunch from production, then a small experiment.
What retrieval confidence is
In a typical RAG pipeline, including ours at Asktopus, the visitor's question is embedded (we use OpenAI's text-embedding-3-small). The vector database (Qdrant) returns the five closest chunks by cosine similarity, and those chunks go into the prompt. The best of those five scores is what usually gets called "retrieval confidence".
It's free, already computed and between 0 and 1, so it looks like the system telling you how sure it is.
The tempting design
if confidence < THRESHOLD:
hand_off_to_human() # or reply "I don't know"
else:
answer_from_context()
It promises two things at once: no made-up answers when the knowledge is thin, and a human for the questions the bot can't handle. All you have to do is pick the number.
What first made me doubt it
In production, the scores of 101 real answers on one site ranged from 0.11 to 0.67, with a median of 0.36. At the 0.4 our docs recommended, more than half of them would have gone to a human. I couldn't tell from the numbers which of those answers were good, and 101 answers on one site is far too few to tune a threshold on. So I set up a test where I could check.
The experiment
The knowledge base is asktopus.com itself: 44 crawled pages in English and Danish, split into 219 chunks of up to 1,500 characters. The pipeline is exactly production's: same embedding model, same Qdrant search, top 5, confidence = the best score.
I wrote 65 questions: 40 I expected the site to answer, 20 I expected it not to (a Zapier integration, ISO 27001, an uptime SLA), and 5 plainly off-topic (the capital of France, tomorrow's weather). Then I set my expectations aside and labelled each question by reading the five chunks the model would receive. Does this text contain the answer? Yes, partial (it answers a narrower version: conversations can be exported, but the format isn't stated), or no.
My expectations were off. Of the 40 questions I expected the site to answer, only 24 got retrieved text that answered them; 6 were partial and 10 got nothing useful. Three of those ten were retrieval misses: the answer was in the knowledge base, just not in the top five. All 25 questions I expected it not to answer got a no.
Results
| Label | Questions | Min | Median | Max |
|---|---|---|---|---|
| Yes | 24 | 0.309 | 0.499 | 0.708 |
| Partial | 6 | 0.331 | 0.379 | 0.582 |
| No | 35 | 0.102 | 0.345 | 0.590 |
The score isn't noise: answerable questions score higher on average. But the two groups overlap from 0.31 to 0.59, and that band holds 19 of the 24 answerable questions, 23 of the 35 unanswerable ones and all six partials. Clean answers exist only at the edges. The 12 questions below 0.31 were all unanswerable, four of them off-topic. The five above 0.59 were all answerable.
Here's what each threshold would do if you hand off whenever confidence is below it:
| Threshold | Answerable, handed off (of 24) | Unanswerable, let through (of 35) |
|---|---|---|
| 0.30 | 0 | 25 |
| 0.35 | 1 | 17 |
| 0.40 | 2 | 14 |
| 0.45 | 7 | 9 |
| 0.50 | 12 | 5 |
There's no good row. At 0.40, 14 questions the retrieved text doesn't answer still go through, and 4 of the 6 partial ones are handed off. At 0.50, half of the answerable questions go to a human. The best single cut on this data, found after the fact, is about 0.43, and it still gets 11 of 59 wrong. I found it by looking at labels you won't have in production, and it doesn't travel: at 0.43, more than half of those 101 production answers would have gone to a human.
Two questions show the problem better than the tables.
The highest-scoring "no": "How many employees does Asktopus have?" scored 0.590. The site doesn't say. The top chunk was the features text about inviting colleagues as editors or viewers, which is about the customer's team, not ours. That score beats 19 of the 24 answerable questions.
The lowest-scoring "yes": "How long does it take to set up?" scored 0.309. The best match was a homepage chunk that ends with the FAQ question "How fast can I get started?". The answer to it falls in the next chunk, which wasn't retrieved. The answer came from the installation guide in second place: "up and running on your website in under 5 minutes". 22 of the 30 on-topic unanswerable questions scored higher.
Why this happens
An embedding captures what a text is about. Cosine similarity tells you the question and a chunk are about the same thing, not that the chunk contains the fact the question asks for. "How many employees does Asktopus have?" is about Asktopus and teams, and so is the features page. That's all the score can see.
The clearest pattern in the data: 30 of the unanswerable questions were on-topic. The 14 of them that use the product's own words ("Asktopus", "chatbot", "assistant", "widget", "AI") had a median score of 0.498, the same as the answerable questions. The 16 that don't had a median of 0.307. On a site about AI chatbots, any question about AI chatbots is close to everything.
What's inside the chunks makes it worse in both directions:
- Chunks that mix many topics score moderately for many questions. In this crawl, navigation, page headers and footers ended up in the chunks, and the docs navigation lists the title of every docs page. "Can I remove the Asktopus branding from the chat?" scored 0.507 with five chunks that were all page headers and navigation. The line that answers it, "Remove branding" under the Pro plan, was in the knowledge base but wasn't retrieved.
- The same mixing dilutes real answers. "Can I upload PDF documents?" scored 0.364 although the top chunk answers it almost word for word. That chunk is a homepage FAQ that also covers setup time, wrong answers, languages and GDPR.
- Short questions carry little signal. In a five- or six-word question, one shared word like the product name carries a lot of weight.
Limitations
This is a small test. One small site, one crawl (partly out of date), 65 questions I wrote myself in English, labelled by my own judgement. Different content, languages or chunking would move the numbers. I also measured retrieval, not answers: I didn't test how reliably the model says "I don't know" when the chunks don't contain the answer. But nothing in the mechanism is specific to this site, and the production numbers point the same way.
What we do instead
In Asktopus, the score doesn't make any decision about the conversation.
- The model decides whether the answer is there, and says so in a way the code can read. The prompt tells it to answer only from the retrieved context. When the context doesn't contain the answer, it starts its reply with a fixed marker, which we strip from the stream before the visitor sees it. The model reads the chunks; the score only measures how close they are.
- That signal, not a score, offers a person. Under an answer the assistant marked as missing, the widget shows a "Talk to a person" button. The same button is always in the header, and explicit requests like "talk to a human" open the same contact form.
- Visitors rate answers. A ๐ or ๐ under every answer is a second signal that costs nothing to collect.
- Missing answers go on a list to review, never trigger anything else. The dashboard's "Questions to answer" list shows questions whose newest asking the model marked as missing, a visitor rated ๐, or scored very low (below 0.25). The site owner can write an article that answers one, or hide it, and a question drops off by itself once its next asking gets a real answer. In this experiment, the 0.25 cut alone flagged 6 questions, all of them unanswered, and missed the other 29. That's fine as one input to a review list and useless as a gate.
In a first live check on our own site, the marker flagged all three questions the site doesn't cover (one of them in Danish) and neither a pricing question it does cover nor a plain "Hi there!". Five questions prove nothing statistically, but the signal comes from something that reads the text, which is the point.
We also stopped using confidence to pick the model. We used to send low-confidence questions to the bigger one, but a bigger model can't answer from content it wasn't given.
Takeaways if you're building RAG
- Don't gate handoffs or "I don't know" on a similarity threshold. Answerable and unanswerable questions overlap across exactly the range where you'd put it.
- Let the model judge whether the retrieved text answers the question, and have it say so in a form your code can read, such as a marker you strip from the stream. It reads the text; the score doesn't.
- Give people an explicit way to reach a person. Asking is a better signal than any score.
- Use low scores as a weak signal for a person to review, where a miss costs nothing.
- Label a few dozen of your own questions before you trust any threshold. An afternoon of reading top-5 chunks tells you more than the scores do.
- Keep navigation and footers out of your chunks. In our crawl they matched questions they couldn't answer. We now drop lines that repeat across most of a site's pages before chunking: on our site that removed 27% of the text, and the top-5 results that were mostly menus and footers fell from 69 of 325 to 3.
Asktopus, the chatbot this came out of, is at asktopus.com.
Top comments (2)
Your 65 labelled questions are already the test set for the marker, which is the number I'd most want to see next. Five live questions can't separate "reads the text" from "got lucky", but 24 answerable and 35 unanswerable can: report both error directions (answerable flagged as missing, unanswerable let through) next to the threshold tables.
Even a clean sweep only bounds the error. Zero misses out of 24 answerable still leaves the true rate anywhere up to about 12% (rule of three: 3/24), and zero out of 35 unanswerable up to about 9%. So the claim you can defend is "below roughly 10%", not "solved".
The six partials are where I'd expect the marker to be least stable. Run each question 3 times and count how many flip between answered and missing. If the model flips on "conversations can be exported, format not stated", that's the band where the threshold failed too, and it tells you whether the review list should include partials rather than only hard misses.
One cheap addition: have the model quote the span it answered from, and string-match the quote against the five chunks in code. An answer with no verbatim quote gets treated as missing. That turns "the model says it's there" into something you can check on the three retrieval misses as well.
The mechanism you found is worth naming, because naming it also tells you what cannot fix it: cosine similarity is symmetric โ cos(q, c) = cos(c, q), exactly (I checked, 0.00e+00 across 2000 random pairs) โ while "this chunk answers this question" is asymmetric. The question asks; the chunk answers. A symmetric function of the pair can't represent an asymmetric predicate, and your two medians are its fingerprint: the on-topic-but-unanswerable questions scored 0.498, the answerable ones 0.499. The score is answering "is this question about this content", answerability is a strict subset of on-topic, so the overlap you measure is the mass of on-topic questions that don't answer โ not noise in the score.
The consequence is stronger than "no single threshold works", and it rules out the fix that looks natural. Any threshold on a monotone recalibration of the score is the same threshold: for monotone f, "hand off when f(s) < t" and "hand off when s < fโปยน(t)" flag exactly the same questions. I took your two group medians, generated two overlapping groups from them, and swept every cut under s, sยฒ and logit(s) โ identical set of achievable table rows, identical best cut (16/59 wrong) under all three. So rescaling, calibrating, or fitting a classifier to the score cannot move a single row of your table. Only a new statistic can, and it has to read the pair with a direction โ which is exactly why your marker works and a better-tuned score never would: the model gets the question as the instruction and the chunk as the evidence, an asymmetric read by construction.
One caveat for your takeaway list, from the same place. "Let the model judge" holds only if the prompt keeps the roles fixed. Phrased symmetrically โ "does this chunk relate to this question" โ the check drifts back toward the score with extra steps and a bill, and the difference is invisible on a five-question smoke test. It shows up the way the score did: on the on-topic questions that don't answer.