DEV Community

RagLeap
RagLeap

Posted on

I benchmarked 4 vector DBs with identical vectors — retrieval was identical, scores were not

A comment on my last release: backend portability only matters if retrieval quality stays comparable. Fair.

So I built ragleap-rag/benchmark.

TL;DR:

  • 296 SQuAD paragraphs, 60 questions, same vectors in pgvector, FAISS, Qdrant, Weaviate vs exact float64 scan
  • overlap@10 = 1.000 on all four
  • Found 5 bugs: score scale mismatch + 3 edge cases
  • Fixed in v0.11.0

How I ran it:

cd java/ragleap-rag
./benchmark/run.sh # starts pgvector (docker), FAISS in-mem, Qdrant, Weaviate

generates benchmark/RESULTS.md + results.csv

Core logic:

var groundTruth = exactScan.query(queryVector, 10);
for (var backend : List.of(pgvector, faiss, qdrant, weaviate)) {
var results = backend.query(queryVector, 10);
double overlap = overlapAtK(groundTruth, results, 10);
double scoreDev = meanAbs(results.scores - (cosine+1)/2);
}

The bug that matters:

FAISS IndexFlatIP returns inner product = cosine if normalized. Weaviate returns cosine distance. Both in [-1,1].

But my abstraction promised [0,1]. Threshold-based filtering broke silently.

Fix:

// In FaissStore & WeaviateStore
return (rawScore + 1.0) * 0.5;

Now deviation: 0.17561 -> 0.00003

Other 3 bugs:

  1. insert([0.1, 0.2]) into dim=768 store — should throw IllegalArgumentException

  2. search() on empty Weaviate collection — should return [], not 500

  3. Failed batch write left ragleap_docs row without vectors

Added 19-case matrix in FailureModeTest.java.

What's next?

  • Milvus/Pinecone via Testcontainers if I get CI keys
  • 100K Wikipedia scale test for recall@10 vs efSearch

Repo: https://github.com/antonyrag/ragleap-core
Results: java/ragleap-rag/benchmark/RESULTS.md
Maven: io.github.antonyrag:ragleap-rag:0.11.0

If you use vector DB abstraction, test your score normalization. Overlap is easy, scores lie.

Top comments (0)