DEV Community

Cover image for Using the Qdrant Python Client with FastAPI
Ayush Kumar
Ayush Kumar

Posted on Originally published at logiclooptech.dev

Using the Qdrant Python Client with FastAPI

The Qdrant Python client gives you a fully async interface that fits cleanly into FastAPI lifespan handlers, supports dense and sparse vectors for hybrid search, and exposes tuning knobs like quantization and sharding that actually matter at scale. I run this stack in production serving millions of vectors per day. Here is the practical breakdown of what works, what breaks, and the patterns I use daily.

Installation and async client setup with FastAPI lifespan

Start with the async extra. You want qdrant-client[fastapi] or just pip install "qdrant-client[async]". The sync client blocks the event loop. Don't use it in an async app.

pip install "qdrant-client[async]" fastapi uvicorn
Enter fullscreen mode Exit fullscreen mode

Initialize the client inside a lifespan context manager. This handles connection pooling and graceful shutdown. I point the client at a DNS name that resolves to a managed Qdrant cluster or a sidecar container. Never hardcode localhost in config.

# app/qdrant.py
from contextlib import asynccontextmanager
from qdrant_client import AsyncQdrantClient
from qdrant_client.http import models
from fastapi import FastAPI
import os

QDRANT_URL = os.getenv("QDRANT_URL", "http://localhost:6333")
QDRANT_API_KEY = os.getenv("QDRANT_API_KEY")  # None for local dev

_client: AsyncQdrantClient | None = None

def get_client() -> AsyncQdrantClient:
    if _client is None:
        raise RuntimeError("Qdrant client not initialized")
    return _client

@asynccontextmanager
async def lifespan(app: FastAPI):
    global _client
    _client = AsyncQdrantClient(
        url=QDRANT_URL,
        api_key=QDRANT_API_KEY,
        timeout=10.0,          # seconds
        prefer_grpc=True,      # lower latency, higher throughput
    )
    # Warm up connections
    await _client.get_collections()
    yield
    await _client.close()
    _client = None

app = FastAPI(lifespan=lifespan)
Enter fullscreen mode Exit fullscreen mode

Setting prefer_grpc=True uses the gRPC interface. It reduces serialization overhead and supports streaming. The HTTP fallback kicks in automatically if gRPC fails. I set a 10 second timeout because vector search can stall on cold segments. You need a circuit breaker upstream. I cover that pattern in Fix AsyncSession Errors in FastAPI (SQLAlchemy 2.0 Guide) for SQLAlchemy; the same retry and timeout logic applies here.

Collection creation with vector params and payload indexes

Create collections on startup or via a migration script. Do not auto-create on every request. Define your vector config and payload indexes upfront. Payload indexes are critical for filtering performance. Without them, Qdrant scans every point.

# app/schema.py
from qdrant_client.http import models

DENSE_VECTOR_NAME = "text-dense"
SPARSE_VECTOR_NAME = "text-sparse"

def get_collection_config(vector_size: int = 1024) -> models.VectorParams:
    return models.VectorParams(
        size=vector_size,
        distance=models.Distance.COSINE,
        on_disk=True,              # move cold vectors to disk
        hnsw_config=models.HnswConfigDiff(
            m=16,                  # connections per node
            ef_construct=100,      # build time accuracy
            full_scan_threshold=10000,
        ),
        quantization_config=models.ScalarQuantization(
            scalar=models.ScalarQuantizationConfig(
                type=models.ScalarType.INT8,
                quantile=0.99,
                always_ram=True,   # keep quantized vectors in RAM
            )
        ),
    )

def get_sparse_vector_config() -> models.SparseVectorParams:
    return models.SparseVectorParams(
        index=models.SparseIndexParams(
            on_disk=False,
        )
    )

PAYLOAD_INDEXES = [
    models.PayloadSchemaType.KEYWORD,  # for exact match filters
    models.PayloadSchemaType.INTEGER,  # for range filters
    models.PayloadSchemaType.FLOAT,    # for score thresholds
    models.PayloadSchemaType.DATETIME, # for time-based queries
]
Enter fullscreen mode Exit fullscreen mode

The on_disk=True flag on dense vectors saves RAM. It adds ~5-10ms latency on cold reads. Worth it if your dataset exceeds memory. The ScalarQuantization config compresses float32 vectors to int8. You lose ~1-2% recall but cut memory by 4x. I run this on all collections larger than 500k vectors.

Create the collection in a startup task:

# app/startup.py
from qdrant_client.http import models
from app.qdrant import get_client
from app.schema import (
    DENSE_VECTOR_NAME,
    SPARSE_VECTOR_NAME,
    get_collection_config,
    get_sparse_vector_config,
    PAYLOAD_INDEXES,
)

COLLECTION_NAME = "documents"

async def ensure_collection(vector_size: int = 1024):
    client = get_client()
    exists = await client.collection_exists(COLLECTION_NAME)
    if exists:
        return
    await client.create_collection(
        collection_name=COLLECTION_NAME,
        vectors_config={
            DENSE_VECTOR_NAME: get_collection_config(vector_size),
        },
        sparse_vectors_config={
            SPARSE_VECTOR_NAME: get_sparse_vector_config(),
        },
    )
    # Create payload indexes
    for field_name, schema_type in [
        ("doc_id", models.PayloadSchemaType.KEYWORD),
        ("tenant_id", models.PayloadSchemaType.KEYWORD),
        ("created_at", models.PayloadSchemaType.DATETIME),
        ("score", models.PayloadSchemaType.FLOAT),
    ]:
        await client.create_payload_index(
            collection_name=COLLECTION_NAME,
            field_name=field_name,
            field_schema=schema_type,
        )
Enter fullscreen mode Exit fullscreen mode

Upserting points with batching and payload filtering

Batch your upserts. Single-point writes kill throughput. I batch 256 points per request. Use wait=True only when you need immediate consistency. For fire-and-forget ingestion, wait=False returns immediately and Qdrant flushes in the background.

# app/ingest.py
from qdrant_client.http import models
from qdrant_client import AsyncQdrantClient
from app.qdrant import get_client
from app.schema import COLLECTION_NAME, DENSE_VECTOR_NAME, SPARSE_VECTOR_NAME
import uuid

BATCH_SIZE = 256

async def upsert_documents(
    texts: list[str],
    dense_vectors: list[list[float]],
    sparse_vectors: list[dict[int, float]],  # {index: value}
    metadata: list[dict],
):
    client = get_client()
    points = []
    for i, text in enumerate(texts):
        point_id = uuid.uuid4().hex
        points.append(
            models.PointStruct(
                id=point_id,
                vector={
                    DENSE_VECTOR_NAME: dense_vectors[i],
                    SPARSE_VECTOR_NAME: models.SparseVector(
                        indices=list(sparse_vectors[i].keys()),
                        values=list(sparse_vectors[i].values()),
                    ),
                },
                payload={
                    "text": text,
                    **metadata[i],
                },
            )
        )
        if len(points) >= BATCH_SIZE:
            await client.upsert(
                collection_name=COLLECTION_NAME,
                points=points,
                wait=False,
            )
            points.clear()
    if points:
        await client.upsert(
            collection_name=COLLECTION_NAME,
            points=points,
            wait=False,
        )
Enter fullscreen mode Exit fullscreen mode

Filter by payload during search. The filter runs before vector scoring. This is efficient when indexes exist.

# app/search.py
from qdrant_client.http import models
from app.qdrant import get_client
from app.schema import COLLECTION_NAME, DENSE_VECTOR_NAME

async def search_with_filter(
    query_vector: list[float],
    tenant_id: str,
    date_from: str | None = None,
    limit: int = 10,
):
    client = get_client()
    must_conditions = [
        models.FieldCondition(
            key="tenant_id",
            match=models.MatchValue(value=tenant_id),
        ),
    ]
    if date_from:
        must_conditions.append(
            models.FieldCondition(
                key="created_at",
                range=models.DatetimeRange(gte=date_from),
            )
        )
    query_filter = models.Filter(must=must_conditions)
    results = await client.search(
        collection_name=COLLECTION_NAME,
        query_vector=models.NamedVector(
            name=DENSE_VECTOR_NAME,
            vector=query_vector,
        ),
        query_filter=query_filter,
        limit=limit,
        with_payload=True,
        score_threshold=0.7,
    )
    return results
Enter fullscreen mode Exit fullscreen mode

Hybrid search using sparse/dense vectors and fusion

Hybrid search combines dense semantic similarity with sparse keyword matching. Qdrant supports Fusion.RRF (Reciprocal Rank Fusion) and Fusion.DBSF. RRF is parameter-free and works well out of the box. DBSF lets you weight dense vs sparse. I default to RRF.

# app/hybrid.py
from qdrant_client.http import models
from app.qdrant import get_client
from app.schema import COLLECTION_NAME, DENSE_VECTOR_NAME, SPARSE_VECTOR_NAME

async def hybrid_search(
    dense_vector: list[float],
    sparse_vector: dict[int, float],
    tenant_id: str,
    limit: int = 10,
):
    client = get_client()
    prefetch = [
        models.Prefetch(
            query=models.NamedVector(
                name=DENSE_VECTOR_NAME,
                vector=dense_vector,
            ),
            limit=50,
            filter=models.Filter(
                must=[
                    models.FieldCondition(
                        key="tenant_id",
                        match=models.MatchValue(value=tenant_id),
                    ),
                ]
            ),
        ),
        models.Prefetch(
            query=models.NamedSparseVector(
                name=SPARSE_VECTOR_NAME,
                vector=models.SparseVector(
                    indices=list(sparse_vector.keys()),
                    values=list(sparse_vector.values()),
                ),
            ),
            limit=50,
            filter=models.Filter(
                must=[
                    models.FieldCondition(
                        key="tenant_id",
                        match=models.MatchValue(value=tenant_id),
                    ),
                ]
            ),
        ),
    ]
    results = await client.query_points(
        collection_name=COLLECTION_NAME,
        prefetch=prefetch,
        query=models.FusionQuery(
            fusion=models.Fusion.RRF,
        ),
        limit=limit,
        with_payload=True,
    )
    return results.points
Enter fullscreen mode Exit fullscreen mode

The prefetch stage runs each vector search independently with a higher limit. Fusion merges the ranked lists. This is slower than single-vector search because you execute two ANN searches. Budget 50-100ms extra latency. If you need sub-50ms p99, stick to dense only and handle keywords in a separate filter.

Performance tuning: quantization, sharding, and connection pooling

Quantization is the biggest lever. I use ScalarQuantization with INT8 and always_ram=True. For collections over 10M vectors, ProductQuantization (PQ) compresses further but requires a training step. PQ adds ~20ms to search latency. Only use it when RAM is the hard constraint.

Sharding splits a collection across nodes. Configure it at creation time:

await client.create_collection(
    collection_name=COLLECTION_NAME,
    vectors_config={...},
    shard_number=4,        # must divide evenly across nodes
    replication_factor=2,  # HA
)
Enter fullscreen mode Exit fullscreen mode

You cannot change shard count later. Plan for 3-5x growth. Each shard holds its own HNSW graph. More shards means more parallelism but smaller graphs per shard, which can hurt recall.

Connection pooling: the async client maintains a pool internally. The default limit is 100 connections. For high-throughput FastAPI workers, increase it:

_client = AsyncQdrantClient(
    url=QDRANT_URL,
    api_key=QDRANT_API_KEY,
    timeout=10.0,
    prefer_grpc=True,
    # grpc specific options
    grpc_options={
        "grpc.max_receive_message_length": 100 * 1024 * 1024,  # 100MB
    },
)
Enter fullscreen mode Exit fullscreen mode

Monitor grpc_client_msg_received_total and http_client_request_duration_seconds in Prometheus. If you see connection exhaustion, increase the pool or add a read replica.

Production patterns: retries, health checks, and monitoring

Wrap client calls in a retry policy. Transient network blips happen. Use tenacity with exponential backoff. Don't retry on 4xx errors.

# app/resilience.py
from tenacity import (
    retry,
    stop_after_attempt,
    wait_exponential_jitter,
    retry_if_exception_type,
)
from qdrant_client.http.exceptions import UnexpectedResponse
import httpx

def is_retryable(exc: BaseException) -> bool:
    if isinstance(exc, (httpx.TimeoutException, httpx.ConnectError)):
        return True
    if isinstance(exc, UnexpectedResponse):
        return 500 <= exc.status_code < 600
    return False

retry_policy = retry(
    wait=wait_exponential_jitter(initial=0.1, max=2.0),
    stop=stop_after_attempt(3),
    retry=retry_if_exception_type((httpx.TimeoutException, httpx.ConnectError, UnexpectedResponse)),
    retry_error_callback=lambda state: state.outcome.exception(),
)
Enter fullscreen mode Exit fullscreen mode

Apply it to search and upsert:

@retry_policy
async def search_with_retry(...):
    return await client.search(...)
Enter fullscreen mode Exit fullscreen mode

Health checks: hit /readyz on the Qdrant HTTP port. It returns 200 when the node can serve traffic. Do not use /healthz; that only reports process liveness.

# app/health.py
from fastapi import APIRouter, HTTPException
from app.qdrant import get_client

router = APIRouter()

@router.get("/ready")
async def ready_check():
    client = get_client()
    try:
        await client.get_collections()
        return {"status": "ok"}
    except Exception as e:
        raise HTTPException(status_code=503, detail=str(e))
Enter fullscreen mode Exit fullscreen mode

Monitoring: export Qdrant metrics to Prometheus. Key alerts:

  • collection_points_count growing unbounded (ingestion > deletion)
  • search_latency_seconds p99 > 500ms
  • wal_write_latency_seconds spikes (disk IO pressure)
  • replication_lag_seconds > 30s on replicas

I built a small wrapper that logs slow queries and emits custom metrics. It sits in the same pattern as the AI agent observability I wrote about in ai agent python code example for FastAPI and OpenAI SDK.

When not to use Qdrant

If your vector count stays under 100k and you already run Postgres, pgvector is simpler. One less moving part. No separate cluster to operate. Qdrant shines when you need multi-tenancy with strict isolation, hybrid search at scale, or quantization to fit large datasets in memory. It also handles payload filtering better than most pgvector setups.

If you need full-text search with complex linguistics (stemming, synonyms, phrase queries), pair Qdrant with Elasticsearch or Typesense. Qdrant's sparse vectors are BM25-like but not a full search engine.

FAQ

How do I migrate a collection schema in Qdrant?
You cannot alter vector params or payload indexes on an existing collection. Create a new collection with the target config, re-index data via scroll and upsert, then swap the alias. Plan for downtime or run dual-write during migration.

Does the Qdrant Python client support asyncio natively?
Yes. AsyncQdrantClient uses httpx.AsyncClient and grpc.aio under the hood. It integrates with FastAPI lifespan and asyncio.gather for concurrent searches. Do not mix sync and async clients in the same process.

What is the difference between search and query_points?
search is the legacy single-vector endpoint. query_points supports prefetch, fusion, and multi-vector queries. Use query_points for hybrid search and advanced ranking. They share the same underlying engine.

How do I handle multi-tenancy?
Option one: one collection per tenant. Simple isolation, but max collections per cluster is ~10k. Option two: single collection with tenant_id payload filter and a keyword index. This scales to millions of tenants. I use option two with a tenant_id index and filter on every query.

Key Takeaways

  • Use AsyncQdrantClient with prefer_grpc=True inside FastAPI lifespan for connection management
  • Enable ScalarQuantization with INT8 and always_ram=True to cut memory 4x with minimal recall loss
  • Create payload indexes for every field you filter on; unindexed filters scan the entire collection
  • Batch upserts at 256 points with wait=False for throughput; use wait=True only when consistency is required
  • Hybrid search via Fusion.RRF adds latency; benchmark before adopting
  • Shard count is immutable; choose 4-8 shards per node for 3-5x headroom
  • Wrap all client calls in tenacity retries with exponential backoff and 5xx filtering
  • Monitor search latency, WAL write latency, and replication lag; alert on p99 > 500ms

Top comments (0)