DEV Community

Cover image for Qdrant FastAPI RAG API: Build a Grounded Service
Gate of AI
Gate of AI

Posted on Originally published at gateofai.com

Qdrant FastAPI RAG API: Build a Grounded Service

🚀 Technical Briefing: This tutorial is part of our deep-dive series on Agentic Workflows at Gate of AI. For the full technical breakdown, interactive code sandbox, and the native Arabic translation, visit the original article here.

Tutorial
Intermediate

Qdrant FastAPI RAG API: Build a Grounded Service


Build a retrieval-augmented generation API with Qdrant semantic search, BAAI/bge-m3 embeddings, FastAPI, LangChain orchestration, GPT-5 generation, and Pydantic response validation.

What This Qdrant FastAPI RAG API Builds

Retrieval-Augmented Generation, or RAG, connects a language model to an external knowledge collection. Instead of asking a model to answer only from its learned parameters, the application first retrieves relevant passages and then supplies those passages to the generation step. Qdrant provides the vector retrieval layer in this architecture. FastAPI exposes the application endpoint, while LangChain can coordinate the retrieval and generation workflow.

This tutorial follows the component pattern documented in the verified context: Qdrant for semantic search, BAAI/bge-m3 for multilingual embeddings, FastAPI for the backend API, GPT-5 or GPT-5-mini for generation, and Pydantic for structured validation. The documented RAGTIME implementation retrieves 15 top chunks and reconstructs approximately 5–7 documents before generation. Those values are used here as a reproducible baseline, not as universal settings for every corpus.

The design is useful for domain collections such as research material, legal documents, institutional knowledge, and multilingual information sources. A legal research example in the verified context combines Qdrant, FastAPI, and a React frontend. Another documented system uses Qdrant as long-term semantic storage in a privacy-focused assistant. These examples demonstrate an architectural pattern; they do not prove that every Qdrant deployment provides the same accuracy, latency, or privacy properties.

A RAG API has four conceptual stages. First, documents are divided into passages. Second, an embedding model converts those passages into vectors. Third, Qdrant searches for passages close to the question vector. Fourth, the language model generates an answer from the retrieved evidence. The answer should be treated as a synthesis of the selected context, not as an automatic guarantee of truth.

Architecture and Verified Design Choices

The request path is intentionally straightforward:

  • Document preparation: source material is cleaned and divided into retrieval passages.
  • Embedding: BAAI/bge-m3 converts passages and questions into vectors suitable for multilingual semantic search.
  • Vector retrieval: Qdrant returns the passages nearest to the question vector.
  • Context assembly: the service selects the strongest evidence and reconstructs a manageable document context.
  • Generation: GPT-5 or GPT-5-mini receives the question and retrieved evidence.
  • Schema validation: Pydantic validates the API response and can enforce a JSON-shaped contract.

Semantic retrieval differs from ordinary keyword lookup. A vector search system compares numerical representations of meaning, so a question and a passage may match even when they do not share the same wording. This is particularly relevant to multilingual collections and domain language. However, semantic similarity does not establish that a passage answers the question. Retrieval quality must therefore be evaluated with representative queries and expected source documents.

The verified RAGTIME example uses a deliberately compact architecture: Qdrant, BAAI/bge-m3, FastAPI, GPT-5 or GPT-5-mini, and Pydantic JSON-schema enforcement. It reports a pipeline that produces citation-grounded JSON reports. The important lesson is not that this exact stack is always optimal, but that a small, clearly defined interface between retrieval, generation, and validation can be easier to inspect than an unnecessarily complex system.

Step 1: Prepare the Python Project

Create a Python project and install the libraries required for the implementation. The version ranges below are intentionally conservative examples. Test the selected versions together in your own environment before deployment. The verified context establishes the roles of FastAPI, Qdrant, LangChain, BAAI/bge-m3, GPT-5-family generation, and Pydantic; it does not establish a universal package lockfile or hosting configuration.

mkdir qdrant-fastapi-rag
cd qdrant-fastapi-rag
python -m venv .venv

# Linux and macOS
source .venv/bin/activate

python -m pip install --upgrade pip
python -m pip install fastapi uvicorn qdrant-client sentence-transformers openai pydantic
mkdir app

Place your documents in a directory named documents. Use material that you are authorized to process. For a multilingual evaluation, include documents in the languages your users will query. BAAI/bge-m3 is selected here because the verified context identifies it as the embedding model in a multilingual RAG system.

Define configuration through environment variables rather than placing credentials in source files:

export QDRANT_URL="http://localhost:6333"
export QDRANT_COLLECTION="multilingual_rag"
export OPENAI_API_KEY="replace-with-an-authorized-key"
export GENERATION_MODEL="gpt-5-mini"
export EMBEDDING_MODEL="BAAI/bge-m3"

The verified sources do not establish a specific Qdrant hosting method, port, authentication scheme, or deployment topology. Use the connection settings required by your chosen Qdrant environment. Keep the vector store and generation provider configuration separate so either component can be evaluated independently.

Step 2: Create the Ingestion and Retrieval Service

Create app/main.py. This compact service loads BAAI/bge-m3, creates a Qdrant collection using the model’s vector size, indexes plain-text files, retrieves 15 passages for each question, and asks a GPT-5-family model to produce a structured response. The code uses the modern OpenAI client construction for generation. The generation provider configuration must support the selected model in your environment.

import os
from pathlib import Path
from typing import Any

from fastapi import FastAPI, HTTPException
from openai import OpenAI
from pydantic import BaseModel, Field
from qdrant_client import QdrantClient, models
from sentence_transformers import SentenceTransformer

QDRANT_URL = os.environ["QDRANT_URL"]
COLLECTION = os.environ.get("QDRANT_COLLECTION", "multilingual_rag")
EMBEDDING_MODEL = os.environ.get("EMBEDDING_MODEL", "BAAI/bge-m3")
GENERATION_MODEL = os.environ.get("GENERATION_MODEL", "gpt-5-mini")

qdrant = QdrantClient(url=QDRANT_URL)
embedder = SentenceTransformer(EMBEDDING_MODEL)
generator = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
app = FastAPI(title="Qdrant FastAPI RAG API")


class AskRequest(BaseModel):
    question: str = Field(min_length=2, max_length=4000)


class Source(BaseModel):
    source: str
    text: str
    score: float


class AskResponse(BaseModel):
    answer: str
    sources: list[Source]


def passages_from_file(path: Path, size: int = 1200) -> list[str]:
    text = path.read_text(encoding="utf-8", errors="replace")
    text = " ".join(text.split())
    return [text[i:i + size] for i in range(0, len(text), size) if text[i:i + size].strip()]


def ensure_collection(vector_size: int) -> None:
    if qdrant.collection_exists(COLLECTION):
        return
    qdrant.create_collection(
        collection_name=COLLECTION,
        vectors_config=models.VectorParams(
            size=vector_size,
            distance=models.Distance.COSINE,
        ),
    )


def index_directory(directory: str) -> int:
    paths = sorted(Path(directory).glob("*.txt"))
    records: list[tuple[Path, str]] = []
    for path in paths:
        records.extend((path, passage) for passage in passages_from_file(path))
    if not records:
        return 0

    vectors = embedder.encode(
        [passage for _, passage in records],
        normalize_embeddings=True,
    )
    ensure_collection(len(vectors[0]))

    points = []
    for number, ((path, passage), vector) in enumerate(zip(records, vectors)):
        points.append(
            models.PointStruct(
                id=number,
                vector=vector.tolist(),
                payload={"source": str(path), "text": passage},
            )
        )
    qdrant.upsert(collection_name=COLLECTION, points=points, wait=True)
    return len(points)


def retrieve(question: str) -> list[Source]:
    vector = embedder.encode(question, normalize_embeddings=True).tolist()
    results = qdrant.search(
        collection_name=COLLECTION,
        query_vector=vector,
        limit=15,
        with_payload=True,
    )
    return [
        Source(
            source=str(result.payload.get("source", "")),
            text=str(result.payload.get("text", "")),
            score=float(result.score),
        )
        for result in results
        if result.payload and result.payload.get("text")
    ]


def generate(question: str, sources: list[Source]) -> str:
    evidence = "\n\n".join(
        f"[Source {index}] {source.source}\n{source.text}"
        for index, source in enumerate(sources, start=1)
    )
    response = generator.chat.completions.create(
        model=GENERATION_MODEL,
        temperature=0,
        messages=[
            {
                "role": "system",
                "content": (
                    "Answer only from the supplied evidence. If the evidence "
                    "does not answer the question, say that the evidence is "
                    "insufficient. Cite factual statements with [Source N]."
                ),
            },
            {
                "role": "user",
                "content": f"Question:\n{question}\n\nEvidence:\n{evidence}",
            },
        ],
    )
    return response.choices[0].message.content or "Insufficient evidence."


@app.post("/v1/ask", response_model=AskResponse)
def ask(request: AskRequest) -> AskResponse:
    try:
        sources = retrieve(request.question)
        if not sources:
            return AskResponse(answer="Insufficient evidence.", sources=[])
        return AskResponse(
            answer=generate(request.question, sources),
            sources=sources,
        )
    except Exception as error:
        raise HTTPException(status_code=503, detail=str(error)) from error


@app.post("/v1/index")
def index() -> dict[str, Any]:
    try:
        return {"indexed_passages": index_directory("documents")}
    except Exception as error:
        raise HTTPException(status_code=503, detail=str(error)) from error

The example intentionally retrieves 15 chunks, matching the verified RAGTIME configuration. A larger set is not automatically better: irrelevant passages can dilute the evidence supplied to the generator. The verified implementation also reconstructs approximately 5–7 documents, which is a useful refinement when multiple chunks belong to the same source. To reproduce that behavior, group retrieved payloads by the source field, order chunks by their original position, and cap the number of reconstructed documents before generation.

Step 3: Index Documents and Run the API

Start the service and call the indexing endpoint:

uvicorn app.main:app --host 127.0.0.1 --port 8000

curl -X POST http://127.0.0.1:8000/v1/index

curl -X POST http://127.0.0.1:8000/v1/ask \
  -H "Content-Type: application/json" \
  -d '{"question":"What does the indexed material say about the topic?"}'

The response contains an answer and the retrieved sources. Inspect both fields. If the expected document is absent from the sources, the failure is in document preparation, embeddings, or retrieval. If the correct source is present but the answer is wrong, investigate context assembly, generation instructions, and schema handling. Separating retrieval evaluation from generation evaluation is one of the most important debugging practices in a RAG project.

For structured reporting, extend AskResponse with fields such as a short answer, a list of citations, and an evidence-status value. Pydantic can validate that the returned object has the expected shape before it reaches a client. The verified RAGTIME system specifically uses Pydantic JSON Schema to enforce valid structured output, so schema validation should be treated as part of the pipeline rather than as a cosmetic API feature.

Step 4: Evaluate Retrieval and Grounding

Create a small evaluation set containing questions, expected source documents, and answer requirements. Measure whether the expected source appears among the top 15 results, whether the reconstructed context contains the required passage, and whether the generated answer cites the appropriate source. Include questions that should not be answerable from the collection. A grounded system must be able to indicate insufficient evidence rather than filling a gap with unsupported content.

Evaluate multilingual questions separately. BAAI/bge-m3 is documented in the verified context for multilingual semantic search, but a multilingual capability should still be tested on the languages, terminology, and document formats relevant to your users. Check whether a question in one language retrieves a relevant passage written in another language, and record failures rather than assuming cross-language retrieval is equally strong in every domain.

Do not convert a result from one experiment into a universal accuracy promise. The verified research reports specific outcomes in specific corpora and model settings, including a study that evaluated compact Gemma models and a TREC system designed for citation-grounded reports. Those findings support the architecture’s plausibility, not a guaranteed score for your own documents.

Limitations and Production Considerations

RAG can reduce unsupported generation by supplying external evidence, but it does not eliminate hallucinations. Qdrant may retrieve a passage that is semantically close but not sufficient. A language model may misread conflicting passages or cite the wrong evidence. The source material may also be incomplete, synthetic, outdated, or limited to a particular specialist domain. These limitations are explicitly relevant to the verified research, which notes constraints involving synthetic personal data, corpus coverage, and hallucination behavior in some settings.

Access control requires design beyond the example. The code demonstrates retrieval and generation, but it does not establish tenant isolation, identity management, legal compliance, or a complete privacy program. Before indexing sensitive material, determine what data may be sent to each component, how long it is retained, who may retrieve it, and how source-level permissions will be enforced. Treat these as deployment requirements, not as capabilities implied by Qdrant or FastAPI alone.

For larger collections, add document identifiers, chunk positions, language metadata, and version information to each Qdrant payload. Then evaluate filtered retrieval and lifecycle operations with the same discipline used for the initial semantic search. If the application serves multilingual or legal users, preserve the original source text and citation metadata so users can inspect the evidence behind a response.

Key Takeaways

  • Qdrant is the vector retrieval component in the verified RAG architecture.
  • FastAPI provides the API boundary, while LangChain can orchestrate retrieval and generation workflows.
  • BAAI/bge-m3 is the verified multilingual embedding choice in the documented RAGTIME system.
  • GPT-5 and GPT-5-mini are the verified generation options described in that system.
  • Retrieving 15 chunks and reconstructing 5–7 documents provides a documented baseline for evaluation.
  • Pydantic JSON Schema helps enforce a predictable, citation-oriented response contract.
  • Retrieval quality, evidence coverage, and abstention behavior must be measured on the target corpus.

Sources and Verification Notes

This tutorial was updated against the verified context available to the Gate of AI Editorial & Engineering Teams. The primary references describe Qdrant, FastAPI, LangChain, BAAI/bge-m3, GPT-5-family generation, Pydantic schema enforcement, multilingual RAG, and domain-specific applications. Architecture and research findings should be reproduced and evaluated in the reader’s own environment before production adoption.

Top comments (0)