I built LocalCortex, a chat app for your own documents that runs entirely on a laptop:
- Ollama for the models,
- Qdrant for search,
- Postgres for chat history,
- no cloud APIs.
You upload PDFs, Word files, or Markdown, ask questions, and get answers with sources. If the answer isn't in your documents, it refuses instead of making something up.
But the app itself isn't really the interesting part.
The interesting part was finding out that almost every idea that sounded like "yeah, this should definitely make it better" either did absolutely nothing, made things worse, or required way too much digging through numbers just to figure out whether it helped at all. So I wrote the eval before most of the features, and I kept a log of every experiment I tried.
This post is the short version of that log.
The final result: on a production-style run of 62 questions, reading every single answer by hand, 61 were correct (98%) and all 14 off-topic prompts were properly refused.
Speed-wise, the median answer time was 48 seconds per query. Not exactly fast, but that's pretty much what you sign up for when you're running Llama 3 on a CPU.
Here's what got it there, and what didn't.
The setup
Two eval sets: a hand-written 20-section employee handbook with 58 questions (including near-duplicate "distractor" sections designed to confuse retrieval), and a larger 4-document, initially 208-chunk policy corpus with 55 questions.
That's 113 questions in total.Metrics: Recall@1/3/5 and MRR for retrieval, a hallucination stress test for refusals, and later a hand-graded answer check.
Models:
nomic-embed-textfor embeddings,llama38B for generating the final answer, both through Ollama.
1. Hybrid search didn’t improve overall rank-1 retrieval
Everyone says to add BM25 next to your embeddings. I did.
Every Qdrant point got a dense and a sparse vector, fused with RRF.
The handbook results:
- Hybrid: 54/58.
- Dense: 56/58 correct at rank 1.
- After hardening keyword matching, abbreviation handling, ranking, and relevance filtering: 55/58.
My theory was that the handbook was too small for BM25 to help, so I built the bigger corpus. Hybrid got worse
- Hybrid: 43/55
- Dense: 47/55.
- Adding the missing EAP expansion brought hybrid from 43/55 to 47/55, which closed the entire rank-1 gap.
So why can hybrid lose when it literally contains the dense ranking?
Because RRF fuses rank positions, not scores.
If the dense top result barely beats the second result, RRF sees rank 1 and rank 2. If the top result absolutely destroys the second result, RRF still just sees rank 1 and rank 2.
That means a noisy sparse ranking can come along and flip something that dense search was already very confident about.

Lesson: Hybrid search means extra storage, extra bookkeeping, extra ranking logic, and, naturally, extra ways for retrieval to go wrong.
On these documents, all that extra machinery gave me no overall rank-1 improvement. Dense and hybrid both finished at 105/113.
So LocalCortex uses dense search by default. Hybrid is still there if you want it, but it's optional.
2. The biggest win was a one-line prefix
This one hurt a little.
nomic-embed-text was trained with task prefixes:
search_query: before questions.
search_document: before indexed text.
I was sending it raw text.
I added the prefixes and got this:
| Dense | Dense + prefixes | |
|---|---|---|
| Handbook (58) | 56 | 55 |
| Large corpus (55) | 47 | 50 |
| Total (113) | 103 | 105 |
With prefixes, dense and hybrid tied at 105/113.
But the more interesting part was what happened to the similarity scores.
The prefixes pushed the scores of correct chunks higher. With my old 0.7 relevance threshold, valid chunks that were being incorrectly rejected dropped:
- from 16 down to 7 on the handbook, and
- from 8 down to 3 on the large corpus.
So yes, after trying considerably more exciting things, one of the biggest improvements was basically adding two strings.
3. Why saying "I don't know" needs a measurement, not a vibe
A RAG app should refuse when nothing relevant exists.
Mine had a relevance threshold of 0.7.
Why 0.7?
Because 0.7 looked reasonable.
Excellent science.
So I finally measured it properly using 39 off-topic questions: 10 general-knowledge questions like "What is the capital of France?" and 29 questions asking about information that simply wasn't in the documents.
The top-1 similarity ranges looked like this:
- Answerable questions: 0.640–0.908
- General knowledge: 0.419–0.567
- questions about missing information: 0.514–0.739
At 0.7, my threshold was apparently feeling a bit too picky, locking out 9 of the 113 valid questions. Dropping it to 0.63 let everyone in while successfully keeping random general-knowledge questions outside the door.
Of course, near-miss questions still tried to crash the party by scoring just like the real ones, which made one thing pretty obvious: a threshold can't catch every kind of bad query.
In a separate check, the model correctly refused all 15 such questions that received context.
Lesson: a threshold is a measurement. Plot the distributions before you pick it, and know which failures it can't catch.
4. How the model handled those 15 questions
The model was not told which questions asked for missing information. It compared each question with the retrieved passages.
The prompt told the model to answer using only the supplied passages and say the information was insufficient when they didn’t support an answer.
Say the document tells us that Naruto became the Seventh Hokage, but says absolutely nothing about his salary.
Ask:
"Which Hokage did Naruto become?"
Easy. The answer is there.
Ask:
"How much did Naruto earn as Hokage?"
Now it should refuse. The document might be about Naruto becoming Hokage, but it never tells us how much he earned.
That's basically what happened with those 15 questions.
The retrieved passages were related to the topic, but they didn’t contain the specific answers. Following the prompt, the model correctly said the information was insufficient. This was the result of that test, not a guarantee that it will always decide correctly.
5. My hallucination test was wrong for a whole session
This was a fun one.
The hallucination test asks 10 questions the documents don't answer and checks whether the model refuses them.
Every run came back:
0/10 PASS.
Naturally, I assumed the model was the problem.
It wasn't.
The model was refusing correctly. My regexes were looking for phrases like "insufficient context" and "couldn't." Meanwhile, the model was saying things like "the provided context is insufficient" and "I could not find."
Same refusal. Different words. My grader had opinions.
It also occasionally saw words repeated from the question, like "parking reimbursement," and decided that was evidence the model had fabricated something.
While debugging that, I found another nice surprise: 18 integration tests were always being skipped because Vitest evaluated their skip condition before the service check ran.
Once those tests actually started running, they immediately found four cases using invalid Qdrant point IDs.
Lesson: test your tests. A green or red number from a broken grader is worse than having no number at all.
6. Follow-up questions leak relevance
Conversation memory made follow-up questions work.
"And after five years?"
On its own, that means basically nothing.
The simple fix was to embed the previous question together with the current one.
And it worked.
Follow-ups found the correct section 12/13 times instead of 4/13.
Great.
Then it broke refusals.
Suppose the previous question was about vacation policy and the next question is:
"What is the capital of France?"
On its own, that question scored 0.479.
Combined with the previous vacation question, it scored 0.682.
My relevance threshold was 0.63.
So congratulations to Paris, it now apparently belongs in the employee handbook.
All 6 off-topic follow-ups leaked through.
I added a second gate: the current question by itself also had to clear 0.57.
Why 0.57?
The highest score among the tested general-knowledge questions was 0.567, so 0.57 was simply the smallest rounded threshold above the range I'd actually observed.
Result:
10/13 follow-ups correct.
0/6 leaked.
But I couldn't just send every question through the rewriter either.
In an early test, after a conversation about someone's employment, I asked:
"What is the capital of France?"
Llama3 rewrote it as:
"What was his employment period?"
It basically looked at my completely unrelated new question and said, "Nope, we're still talking about the previous thing."
So now I only rewrite questions that actually appear to refer back to earlier context, and I reject rewrites that throw away too much of the user's original wording.
With those guard rails:
11/13 follow-ups correct.
0/6 leaked.
Lesson: every feature that makes answering easier also seems to find a creative new way to make refusing harder.Measure both.
7. Chunking: one bug fix beat three clever ideas
- A splitter bug: My recursive splitter included "## " as a separator. That sounds harmless until you realize it could split inside the heading marker itself. 61 of 208 chunks were broken. Some were literally just #. Others were headings with no body. I fixed the bug. 208 chunks became 117. Same scores.
Already a win.
- Bigger chunks (800 chars): The handbook's refusals went from 0 to 6, because two short unrelated sections got packed into one chunk whose embedding sat between both topics.
Removed.
- Contextual Retrieval (an LLM-written blurb per chunk): I generated an LLM-written blurb for every chunk. It produced the best rank-1 recall on the resume fixture at that point. It also took 25–31 seconds per chunk on CPU. And one vague generated blurb managed to push the correct chunk out of the top 5 for two questions.
Removed.
- Semantic chunking: It only cuts at sentence ends, so the handbook turned into 48 scraps and bullet lists got glued together.
Not the default.
- What stayed: A "Title › Section" header above each chunk, but automatically for documents with fewer than two Markdown headings like resumes and most other PDFs. On a resume it took rank-1 recall from 5/14 to 8/14. On structured documents it cost a refusal each, so it switches on automatically from the text, not the file type.
Lesson: a surprising number of retrieval problems are just chunking problems wearing a fake moustache.
8. The model needed whole documents, not five fragments
On my resume, I asked:
"How many projects does Sekharendu Dey have?"
The answer came back:
"At least three."
I have two.
The model had counted jobs as projects because I'd handed it five fragments ordered by retrieval relevance instead of giving it the résumé in reading order.
Some project bullets also arrived without their project headings.
So I separated retrieval from context assembly.
Retrieval now answers one question first:
"Is anything here relevant enough to continue, and which documents matched?"
Then context assembly decides what the model should actually read.
If a matched document is 1,000 tokens or less, I include the whole thing when the context budget allows it.
For larger documents, the model gets the strongest matching chunks, the opening chunk of the document, and nearby chunks in reading order, up to a 1,200-token per-document limit and the overall 4,096-token context window.
On a fictional resume fixture, correct answers went from 17/25 to 22/25.
On my actual resume, it now says "two projects" and names both.
There was a trade-off, of course.
Long-document answers became roughly 2× slower.
And one question on a 48-page PDF actually got worse because the answer disappeared inside a prompt that was 2.5× longer.
A later context limit fixed that answer.
The extra latency, unfortunately, did not receive the memo.
9. Read the answers yourself
My keyword grader gave the final run a score of 57/62.
Then I read the answers.
Four of the five "wrong" answers were actually correct.
So the real reviewed score was 61/62.
One answer was correct but apparently too short for the grader's liking.
Another correctly gave the Denver office's opening time, while the grader expected both the opening and closing time.
There was only one actual miss:
"who made the CUGA thing."
The wording was vague enough that its retrieval score fell below the 0.63 threshold, so the app refused it.
When I rephrased it to clearly ask for the authors' organizations, it answered correctly.
Lesson: automated grading tells you where to look. It doesn't get to tell you the final score.
What I'd tell someone building their first RAG
Build an evaluation set early. Do it before adding enough features that you can no longer tell what actually made the system better.
Measure refusals just as seriously as answers. Off-topic questions are half the product.
Keep a log of what you tried and what you removed. A bunch of things I built sounded useful and were removed after I measured them. Without the log, I'd probably forget why.
Try the boring fix first. A prefix. A splitter bug. One missing acronym. Those ended up doing more for me than several much fancier ideas.
The code, the eval sets and every script behind these numbers are on GitHub:
github.com/Sekharendu/LocalCortex (MIT).
It starts with one command: docker compose --profile app up -d.
If you're building something similar, or you know a way to catch slangy questions without letting "the capital of France" through, I'd love to hear it.
Top comments (0)