DEV Community

Silver_dev
Silver_dev

Posted on

Vector database showroom. Part 3 Qdrant — The Hot Hatchback That Carries More Than It Looks

🏎️ Qdrant — The Hot Hatchback That Carries More Than It Looks

Sharp acceleration, precise handling — and a trunk noticeably bigger than it looks from the outside.

Under the hood: a Rust core shipped as a single binary or container (the Python client even has an embedded/local mode for experiments). Apache 2.0. The signature feature — filterable HNSW: the graph is built with payload indexes in mind, so filters participate in graph traversal instead of being glued on top.

Where it shines:

  • Remember pgvector's three filter tools? Here filtering is part of the index itself: "similar, but only year=2024" is this car's home turf
  • Hybrid out of the box: sparse vectors since v1.7, the Query API with prefetch and fusion (RRF/DBSF) since v1.10 — dense + sparse in a single request
  • Quantization is our "lightweight alloys": int8 shrinks vectors 4×, binary 32×. Originals can live on disk while only compressed copies stay in RAM. Hundreds of millions of vectors on a single node is a real setup, not marketing
  • Multivectors (ColBERT-style):* several vectors per point, late interaction — something very few databases in this showroom offer Operations: Prometheus metrics, snapshots, API keys + JWT RBAC

Where it stalls:

  • Bring your own embeddings — it's a database, not an embedder (contrast with last issue's scooter). Clients for Python/JS/Go/Rust, LangChain/LlamaIndex integrations included
  • Distributed mode (sharding, replication, Raft) is real ops work — simpler than the freight train arriving later in this series, heavier than Postgres
  • Backups are manual snapshots. No WAL, no point-in-time recovery like the one you're used to in Postgres
  • Tuning is part of the contract: m, ef_construct, hnsw_ef, quantization with rescoring. Defaults are decent, but a hot hatch likes to be tuned

What breaks if you skip the manual:

  • The distance function (COSINE/DOT/EUCLID/MANHATTAN) is locked at collection creation. Wrong pick = full reindex — there is no "alter on the fly"
  • Payload indexes: forget to create them and filtered search degrades to scans. The right order: collection → payload indexes → bulk upload. Adding an index afterwards rebuilds the HNSW graph in the background to wire up the filterable edges
  • Benchmark trap #1: while a segment stays below the optimizer's indexing_threshold, no HNSW index is built for it — search falls back to brute force, fast and with perfect recall. A naive benchmark on a small collection measures exact search, not HNSW. Honest numbers require production-like volume — or explicitly lowering indexing_threshold
  • Quantization without rescoring eats recall — and binary quantization is model-sensitive: brilliant on some embeddings, poor on others. Ground truth for comparison: search with exact=true. Showroom thesis: vendor benchmarks are ads — verify on your own data
  • Named vectors: if the collection was created with named vectors, queries must pass the vector name via the using parameter — otherwise a 400

Cost of ownership: free under Apache 2.0; Qdrant Cloud (managed & hybrid) exists. RAM appetite is flexible — quantization + on-disk originals instead of "6 GB per million."

Mechanics & parts: Qdrant GmbH (Berlin) behind it; the docs are among the best in the industry. "Qdrant engineers" are rare on the market, but this is a database a backend developer can drive — not a twenty-pod zoo.

✅ Take it if: millions to hundreds of millions of vectors, filter-heavy RAG, self-hosting, and you want speed without building railroads

❌ Pass if: your logic lives in SQL and you need joins (the station wagon does those better), or you're counting billions with GPUs — the train arrives later in this series

Test drive:

from qdrant_client import QdrantClient, models

# qdrant-client >= 1.10; for local experiments: QdrantClient(path="./qdrant_data")
client = QdrantClient(url="http://localhost:6333")

client.create_collection(
    collection_name="docs",
    vectors_config=models.VectorParams(size=1536, distance=models.Distance.COSINE),
)

# payload index BEFORE upload — this is how filterable HNSW gets built
client.create_payload_index(
    collection_name="docs",
    field_name="year",
    field_schema=models.PayloadSchemaType.INTEGER,
)

# embed() — your embedding function (OpenAI, BGE, anything)
client.upsert(
    collection_name="docs",
    points=[
        models.PointStruct(id=1, vector=embed("Quarterly report 2024"), payload={"year": 2024}),
        models.PointStruct(id=2, vector=embed("Legal contract draft"), payload={"year": 2024}),
    ],
)

# the filter participates in graph traversal, not glued on top
# query_points returns a QueryResponse — the hits live in .points
res = client.query_points(
    collection_name="docs",
    query=embed("annual financial summary"),
    query_filter=models.Filter(
        must=[models.FieldCondition(key="year", match=models.MatchValue(value=2024))]
    ),
    limit=10,
).points

for point in res:
    print(point.id, point.score, point.payload)
Enter fullscreen mode Exit fullscreen mode

A hatchback doesn't pretend to be a station wagon and doesn't build railroads — it just carries fast. And when your data outgrows the trunk, the API doesn't change: you move to a cluster via snapshots, and the application never notices.

Top comments (0)