Qdrant FastAPI RAG API: Build a Grounded Service

Share:
Tutorial Intermediate

Qdrant FastAPI RAG API: Build a Grounded Service

Build a retrieval-augmented generation API with Qdrant semantic search, BAAI/bge-m3 embeddings, FastAPI, LangChain orchestration, GPT-5 generation, and Pydantic response validation.

What This Qdrant FastAPI RAG API Builds

Retrieval-Augmented Generation, or RAG, connects a language model to an external knowledge collection. Instead of asking a model to answer only from its learned parameters, the application first retrieves relevant passages and then supplies those passages to the generation step. Qdrant provides the vector retrieval layer in this architecture. FastAPI exposes the application endpoint, while LangChain can coordinate the retrieval and generation workflow.

This tutorial follows the component pattern documented in the verified context: Qdrant for semantic search, BAAI/bge-m3 for multilingual embeddings, FastAPI for the backend API, GPT-5 or GPT-5-mini for generation, and Pydantic for structured validation. The documented RAGTIME implementation retrieves 15 top chunks and reconstructs approximately 5–7 documents before generation. Those values are used here as a reproducible baseline, not as universal settings for every corpus.

The design is useful for domain collections such as research material, legal documents, institutional knowledge, and multilingual information sources. A legal research example in the verified context combines Qdrant, FastAPI, and a React frontend. Another documented system uses Qdrant as long-term semantic storage in a privacy-focused assistant. These examples demonstrate an architectural pattern; they do not prove that every Qdrant deployment provides the same accuracy, latency, or privacy properties.

A RAG API has four conceptual stages. First, documents are divided into passages. Second, an embedding model converts those passages into vectors. Third, Qdrant searches for passages close to the question vector. Fourth, the language model generates an answer from the retrieved evidence. The answer should be treated as a synthesis of the selected context, not as an automatic guarantee of truth.

Architecture and Verified Design Choices

The request path is intentionally straightforward:

  • Document preparation: source material is cleaned and divided into retrieval passages.
  • Embedding: BAAI/bge-m3 converts passages and questions into vectors suitable for multilingual semantic search.
  • Vector retrieval: Qdrant returns the passages nearest to the question vector.
  • Context assembly: the service selects the strongest evidence and reconstructs a manageable document context.
  • Generation: GPT-5 or GPT-5-mini receives the question and retrieved evidence.
  • Schema validation: Pydantic validates the API response and can enforce a JSON-shaped contract.

Semantic retrieval differs from ordinary keyword lookup. A vector search system compares numerical representations of meaning, so a question and a passage may match even when they do not share the same wording. This is particularly relevant to multilingual collections and domain language. However, semantic similarity does not establish that a passage answers the question. Retrieval quality must therefore be evaluated with representative queries and expected source documents.

The verified RAGTIME example uses a deliberately compact architecture: Qdrant, BAAI/bge-m3, FastAPI, GPT-5 or GPT-5-mini, and Pydantic JSON-schema enforcement. It reports a pipeline that produces citation-grounded JSON reports. The important lesson is not that this exact stack is always optimal, but that a small, clearly defined interface between retrieval, generation, and validation can be easier to inspect than an unnecessarily complex system.

Step 1: Prepare the Python Project

Create a Python project and install the libraries required for the implementation. The version ranges below are intentionally conservative examples. Test the selected versions together in your own environment before deployment. The verified context establishes the roles of FastAPI, Qdrant, LangChain, BAAI/bge-m3, GPT-5-family generation, and Pydantic; it does not establish a universal package lockfile or hosting configuration.

mkdir qdrant-fastapi-rag
cd qdrant-fastapi-rag
python -m venv .venv

# Linux and macOS
source .venv/bin/activate

python -m pip install --upgrade pip
python -m pip install fastapi uvicorn qdrant-client sentence-transformers openai pydantic
mkdir app

Place your documents in a directory named documents. Use material that you are authorized to process. For a multilingual evaluation, include documents in the languages your users will query. BAAI/bge-m3 is selected here because the verified context identifies it as the embedding model in a multilingual RAG system.

Define configuration through environment variables rather than placing credentials in source files:

export QDRANT_URL="http://localhost:6333"
export QDRANT_COLLECTION="multilingual_rag"
export OPENAI_API_KEY="replace-with-an-authorized-key"
export GENERATION_MODEL="gpt-5-mini"
export EMBEDDING_MODEL="BAAI/bge-m3"

The verified sources do not establish a specific Qdrant hosting method, port, authentication scheme, or deployment topology. Use the connection settings required by your chosen Qdrant environment. Keep the vector store and generation provider configuration separate so either component can be evaluated independently.

Step 2: Create the Ingestion and Retrieval Service

Create app/main.py. This compact service loads BAAI/bge-m3, creates a Qdrant collection using the model’s vector size, indexes plain-text files, retrieves 15 passages for each question, and asks a GPT-5-family model to produce a structured response. The code uses the modern OpenAI client construction for generation. The generation provider configuration must support the selected model in your environment.

import os
from pathlib import Path
from typing import Anyfrom fastapi import FastAPI, HTTPException
from openai import OpenAI
from pydantic import BaseModel, Field
from qdrant_client import QdrantClient, models
from sentence_transformers import SentenceTransformerQDRANT_URL = os.environ["QDRANT_URL"]
COLLECTION = os.environ.get("QDRANT_COLLECTION", "multilingual_rag")
EMBEDDING_MODEL = os.environ.get("EMBEDDING_MODEL", "BAAI/bge-m3")
GENERATION_MODEL = os.environ.get("GENERATION_MODEL", "gpt-5-mini")qdrant = QdrantClient(url=QDRANT_URL)
embedder = SentenceTransformer(EMBEDDING_MODEL)
generator = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
app = FastAPI(title="Qdrant FastAPI RAG API")class AskRequest(BaseModel):
question: str = Field(min_length=2, max_length=4000)class Source(BaseModel):
source: str
text: str
score: floatclass AskResponse(BaseModel):
answer: str
sources: list[Source]def passages_from_file(path: Path, size: int = 1200) -> list[str]:
text = path.read_text(encoding="utf-8", errors="replace")
text = " ".join(text.split())
return [text[i:i + size] for i in range(0, len(text), size) if text[i:i + size].strip()]def ensure_collection(vector_size: int) -> None:
if qdrant.collection_exists(COLLECTION):
return
qdrant.create_collection(
collection_name=COLLECTION,
vectors_config=models.VectorParams(
...

Continue Reading

Log in for free to read the rest of this article and access exclusive AI tools.

Log in / Register

Was this tutorial helpful?

GateOfAI AI Guide
Online
Hello! Welcome to GateOfAI. I am your guide copilot. I can answer questions about our SaaS tools, pricing, vetted developers, and escrow safety. How can I help you today?