DEV Community

HiDevs
HiDevs

Posted on

Turbo4’s Real Cost: What You Give Up When You Drop the Full-Precision Copy

A reproducible benchmark measuring recall, storage, and throughput across five embedding models - comparing Qdrant 1.19's storage-only Turbo4 datatype against the two-copy TurboQuant model with rescoring.
A reproducible benchmark measuring recall, storage, and throughput across five embedding models - comparing Qdrant 1.19's storage-only Turbo4 datatype against the two-copy TurboQuant model with rescoring.

When Qdrant 1.19 introduced the Turbo4 datatype, it presented vector search engineers with a compelling value proposition: drop full-precision vectors entirely, retain only a tiny 4-bit compressed version, and slash collection storage footprints ninefold. But this massive compression saving comes with an immediate operational trade-off.

Because original vectors are completely discarded from disk, there is no underlying full-precision baseline left to double-check candidate points against during search. As a result, distance metrics derived from 4-bit compressed representations become the final search score - eliminating the traditional safety net used to correct quantization errors.

Prior to this update, standard production setups running on Qdrant 1.18 relied on TurboQuant, Qdrant's standard 4-bit quantization layer. To guarantee high retrieval accuracy, TurboQuant maintains a dual-copy architecture:

  • In-Memory Copy: A 4-bit compressed vector held in fast RAM for rapid candidate scanning across HNSW index graphs.
  • On-Disk Copy: The original float32 vector saved to disk, used specifically during the final rescoring pass to re-evaluate exact distances for top candidates.

While this rescoring approach provided an impressive 98.2% recall on a 1-million-document OpenAI collection, maintaining dual-copy storage scales steeply. At 1,536 dimensions per vector, the raw coordinate data alone requires 6.9 GB of space - meaning a planned expansion to 1*0 million documents* would demand over 65 GB just to preserve full-precision rescoring safety.

Figure 1: TurboQuant keeps both copies and double-checks results. Turbo4 keeps only the compressed copy - no double-checking possible.
Figure 1: TurboQuant keeps both copies and double-checks results. Turbo4 keeps only the compressed copy - no double-checking possible.

Mechanics of Hadamard Rotations and Rescoring Overhead

Our starting point was TurboQuant from Qdrant 1.18. It uses a trick from Google Research: before compressing vectors to 4 bits, it scrambles them with a mathematical rotation (Hadamard transform) that spreads the important information evenly across all coordinates.

Qdrant stores both the original and the compressed copy. During search, it uses the compressed copy for speed, then pulls the originals from disk for the top candidates and re-checks distances. This rescoring step catches most compression mistakes.

Here is what that rescoring search looks like in code:

#Python — TurboQuant Search with Rescoring

results = client.query_points(
    collection_name="wiki_tq_rescore",
    query=query_vector, limit=10,
    search_params=models.SearchParams(
        hnsw_ef=128,
        quantization=models.QuantizationSearchParams(
            rescore=True,         # Re-check against full-precision originals
            oversampling=2.0,     # Pull 20 candidates, return best 10
        ),
    ),
)
Enter fullscreen mode Exit fullscreen mode

The rescore and oversampling parameters are part of Qdrant’s quantization search configuration. Qdrant documents oversampling as a way to pre-select additional candidates from the quantized index before rescoring them with the original vectors.

Architecture Choices: Dual-Copy Quantization vs. Pure 4-Bit Storage

Qdrant handles this via two explicit setup types:

  • TurboQuant (FLOAT32 with Quantization): Holds both quantized representations and float32 original vectors. It optimizes for maximum recall by accepting a higher disk footprint.
  • Turbo4 Datatype (TURBO4): Stores vectors exclusively as 4-bit quantized values. It cuts storage radically by eliminating raw vector copies.

Side-by-Side Accuracy, Memory, and Throughput Measurements

We wanted to test one thing: what happens to search accuracy and speed when you remove the safety net? To make the comparison fair, we kept everything else identical - same hardware, same index settings, same queries. We computed the "correct" answers using a slow but perfect brute-force search, then measured how close each setup got.

Benchmark Setup & Environment
Table 1

Evaluated Embedding Models
We picked five embedding models that cover a wide range:
Table 2

For the data, we used 500,000 Wikipedia abstracts. Wikipedia is messy and varied, which is the point - synthetic test data would have made the results look better than they really are. We wanted numbers we could trust in production.

Collection Setup Scripts :

Python - Turbo4 Collection (compressed only, no originals kept)

client.create_collection(
    collection_name="wiki_turbo4",
    vectors_config=models.VectorParams(
        size=1536,
        distance=models.Distance.COSINE,
        datatype=models.Datatype.TURBO4,   # 4-bit only, no originals
    ),
    hnsw_config=models.HnswConfigDiff(m=16, ef_construct=128),
)
Enter fullscreen mode Exit fullscreen mode

The TURBO4 datatype used above is documented in Qdrant’s vector datatype documentation.

Python - TurboQuant Collection (keeps both compressed and original)

client.create_collection(
    collection_name="wiki_tq_rescore",
    vectors_config=models.VectorParams(
        size=1536,
        distance=models.Distance.COSINE,
        datatype=models.Datatype.FLOAT32,
    ),
    quantization_config=models.TurboQuantConfig(
        encoding=models.TurboQuantEncoding.BITS4,
    ),
)
Enter fullscreen mode Exit fullscreen mode

Accuracy Measurement Function :

Python - Recall Measurement

def compute_recall(predicted_ids, ground_truth_ids):
    """What fraction of the correct answers did we find?"""
    return len(set(predicted_ids) & set(ground_truth_ids)) / len(ground_truth_ids)

# Average over 1,000 queries
recalls = [compute_recall(approx[i], exact[i]) for i in range(len(queries))]
mean_recall = sum(recalls) / len(recalls)
Enter fullscreen mode Exit fullscreen mode

Recall Performance Across Vector Lengths

When we ask for the top 10 results (the most common setting for RAG), Turbo4 found between 96.1% and 98.9% of the correct answers, depending on the model. The gap compared to TurboQuant-with-rescoring ranged from about half a percentage point to nearly three percentage points.

Table 3

Why do longer vectors handle compression better?
Think of it this way: when you round each number in a vector, you introduce a tiny error. If the vector only has 384 numbers, each error matters more. With 3072 numbers, the errors tend to cancel each other out - some round up, some round down, and the overall distance barely changes. The math behind this is the law of large numbers, and it shows up clearly in our results.

Figure 2: Recall@10 across all configurations. The gap shrinks as vectors get longer.
Figure 2: Recall@10 across all configurations. The gap shrinks as vectors get longer.

In practical terms, the 384-dimension gap means about 1 in 37 queries would lose a relevant document from the results. At 1536 dimensions and above, it drops to about 1 in 100 - and if you have a reranker sitting after the retrieval step, even that small difference tends to disappear.

Figure 3: At deeper retrieval (k=100), the gap widens slightly, especially for shorter vectors.
Figure 3: At deeper retrieval (k=100), the gap widens slightly, especially for shorter vectors.

Storage Metrics & Throughput Trade-Offs

The storage numbers are direct: TurboQuant adds a 4-bit compressed copy to 32-bit float originals (36 bits total), while Turbo4 keeps only the 4-bit version. That provides a deterministic 9x reduction. Speed also improved because Turbo4 never has to read original vectors from disk to rescore results.

Table 4

Figure 4: Storage comparison. Turbo4 is consistently 9× smaller.
Figure 4: Storage comparison. Turbo4 is consistently 9× smaller.

To put this in perspective: a 10M-vector collection with OpenAI's largest model takes about 130 GB with TurboQuant. With Turbo4, it is about 14.4 GB. That is the difference between needing a special high-memory server and running comfortably on a standard cloud instance.

Speed also improved, for a simple reason. TurboQuant has to read the original vectors from disk every time it rescores a result. With Turbo4, there are no originals to read - everything stays in memory. The speed gain was 25–33% depending on the model, and it got bigger with longer vectors because those take more time to read from disk.

Table 5

Figure 5: Turbo4 is faster across the board, and the gap grows with vector size.
Figure 5: Turbo4 is faster across the board, and the gap grows with vector size.

Deployment Recommendations & Architectural Trade-offs

Four clear rules emerge from these benchmark results__:

  1. Vector Dimension Sensitivity: Short vectors (384d) lose nearly 3 percentage points of recall without rescoring, whereas high-dimensional vectors (1536d+) lose only ~1 point.
  2. Deterministic Footprint Reduction: Storage reduction remains exactly 9x, making large-scale deployments dramatically cheaper.
  3. I/O Efficiency Gains: Eliminating disk rescoring yields a 25–33% throughput boost depending on coordinate length.

Production Setup Choices

  • Turbo4 Strategy: Recommended for OpenAI embeddings (1536d+) where a downstream cross-encoder reranker is present, making the 1*% recall difference* negligible while saving substantial infrastructure costs.
  • TurboQuant Strategy: Recommended for shorter vectors (384d) in low-latency search or autocomplete tasks where maximum first-pass recall is required without downstream re-ranking.

Top comments (0)