DEV Community

TechSimPlus Learnings
TechSimPlus Learnings

Posted on Originally published at techsimplus.com

Pinecone vs Qdrant vs pgvector: The Same RAG Query in All Three (Python Code Comparison)

Comparison tables are useful, but nothing beats seeing the code. Here is the same task in all three: store chunk embeddings and run a similarity search filtered to one tenant.

Assume vec is a 1536-dimension embedding from text-embedding-3-small.

pgvector (Postgres extension)

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE chunks (
  id BIGSERIAL PRIMARY KEY,
  tenant_id TEXT NOT NULL,
  content TEXT,
  embedding VECTOR(1536)
);

CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);

-- query: <=> is cosine distance
SELECT id, content
FROM chunks
WHERE tenant_id = 'acme'
ORDER BY embedding <=> $1
LIMIT 5;
Enter fullscreen mode Exit fullscreen mode

What you get: plain SQL, joins with your other tables, transactions, one database to operate. Since pgvector 0.8, iterative index scans improve results when filters are selective.

Qdrant

from qdrant_client import QdrantClient, models

client = QdrantClient(url="http://localhost:6333")

client.create_collection(
    collection_name="chunks",
    vectors_config=models.VectorParams(size=1536, distance=models.Distance.COSINE),
)

client.upsert(
    collection_name="chunks",
    points=[models.PointStruct(id=1, vector=vec, payload={"tenant_id": "acme", "content": "..."})],
)

hits = client.query_points(
    collection_name="chunks",
    query=vec,
    query_filter=models.Filter(
        must=[models.FieldCondition(key="tenant_id", match=models.MatchValue(value="acme"))]
    ),
    limit=5,
).points
Enter fullscreen mode Exit fullscreen mode

What you get: open source, runs locally with docker run -p 6333:6333 qdrant/qdrant, rich payload filters, quantization to cut memory, sparse vectors for hybrid search. Create a payload index on tenant_id for fast filtering at scale.

Pinecone

from pinecone import Pinecone

pc = Pinecone(api_key="YOUR_API_KEY")
index = pc.Index("chunks")

index.upsert(
    vectors=[{"id": "1", "values": vec, "metadata": {"content": "..."}}],
    namespace="acme",
)

res = index.query(vector=vec, top_k=5, namespace="acme", include_metadata=True)
Enter fullscreen mode Exit fullscreen mode

What you get: fully managed and serverless, no servers to run. Namespaces are a clean way to isolate tenants.

Quick verdict

If you... Pick
Already run Postgres and have up to a few million chunks pgvector
Need open source, self-hosting or advanced filtering at scale Qdrant
Want zero ops and fast time to production Pinecone

Whichever you choose, hide it behind a small retrieve(tenant_id, query) function so switching later is a one-file change.

In Vector 2.0, the live Gen-AI developer cohort by TechSimPlus and Prateek Mishra, you build RegRadar with Pinecone and pgvector, compare against Qdrant, and measure the difference with RAGAS.

Check the Complete Details: https://vector.techsimplus.com

Top comments (0)