DEV Community

Cover image for Native Quantization: Let OpenSearch Service Compress Your Vectors
Jon Handler
Jon Handler

Posted on Originally published at Medium

Native Quantization: Let OpenSearch Service Compress Your Vectors

Send FP32, get 2x to 32x compression, and leave your ingestion pipeline untouched. The engine does the work.

The previous article in this series put the quantization work on you. You convert vectors to a reduced precision before indexing, and Amazon OpenSearch Service stores exactly what you send. That gives you full control, and it gives you a standing job: pick a method, run the conversion in your pipeline, and re-check accuracy every time your embedding model changes.

There is another way. OpenSearch Service can compress at index time. You send the full-precision FP32 vectors your embedding model already produces, and the engine converts each one to a lower-bit representation during ingestion and stores the compressed vectors in the index. Both the Faiss and Lucene engines offer native quantization, each with its own methods. Your ingestion pipeline does not change. The control over the accuracy tradeoff moves out of your code and into index configuration.

The baseline is the same one from the previous article. At 100 million vectors of 1,536 dimensions, full-precision FP32 needs about 643 GB for one copy, roughly 1.3 TB with a replica. That figure is the best-latency deployment, holding every vector in memory; with memory-optimized loading you can serve the same index in less RAM and page from disk at some latency cost. Every option here reduces the footprint while leaving ingestion identical. This is the second of three articles on the quantization spectrum in OpenSearch Service. The first covered pre-quantized vectors. The third covers on-disk mode and cost stacking.

Scalar quantization: the engine compresses each dimension

Faiss scalar quantization is the native path you reach for first. You set the encoder on the knn_vector field to sq and give it a bits value, and OpenSearch Service converts each FP32 dimension to the target precision during ingestion, storing the quantized vectors in the index. Valid bit widths are 16, 4, 2, and 1, spanning the range from near-lossless to aggressive compression.

At 16 bits, scalar quantization halves memory, the same reduction as pre-quantized FP16, dropping a replicated 100-million-vector index to about 656 GB, without a conversion step in your pipeline. OpenSearch stores the 16-bit vectors and builds the HNSW graph over them, and at search time it computes distances directly on the 16-bit values, accelerated by SIMD instructions on recent processors, so there is no round trip back to full precision. The type parameter selects the 16-bit format: fp16, the default, for the highest precision within a plus or minus 65,504 range, or bf16 (introduced in OpenSearch 3.9) for the full 32-bit value range at slightly lower precision. A companion clip parameter governs fp16 only: left at its default it rejects any vector with a value outside the fp16 range, and set to true it rounds out-of-range values to the limits instead. Both type and clip apply to 16-bit quantization only.

PUT /my-sq-index
{
  "settings": { "index.knn": true },
  "mappings": {
    "properties": {
      "embedding": {
        "type": "knn_vector",
        "dimension": 1536,
        "space_type": "l2",
        "method": {
          "name": "hnsw",
          "engine": "faiss",
          "parameters": {
            "m": 16,
            "ef_construction": 256,
            "encoder": { "name": "sq", "parameters": { "bits": 16 } }
          }
        }
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

The low-bit end: 1, 2, and 4-bit quantization

Below 16 bits, scalar quantization reaches the aggressive end of the range. Multi-bit scalar quantization added two-bit (16x) and four-bit (8x) options for Faiss in OpenSearch 3.9, and one-bit (32x) has been available since OpenSearch 3.6. These bit widths are supported for the HNSW method. You configure the multi-bit options by setting the sq encoder bits to 2 or 4, or compression_level on the field to 8x or 16x. For the 32x end, set compression_level to 32x and let OpenSearch choose the algorithm, which since 3.7 is the optimized scalar quantizer described below. At 32x, a replicated 100-million-vector index drops to about 66 GB, from the 1.3 TB FP32 baseline. Collapsing each dimension toward a single bit discards magnitude, so benchmark against your own data; at 32x the optimized scalar quantizer recovers much of that recall for you, as the next section describes.

PUT /my-sq-2bit-index
{
  "settings": { "index.knn": true },
  "mappings": {
    "properties": {
      "embedding": {
        "type": "knn_vector",
        "dimension": 1536,
        "space_type": "l2",
        "method": {
          "name": "hnsw",
          "engine": "faiss",
          "parameters": {
            "encoder": { "name": "sq", "parameters": { "bits": 2 } }
          }
        }
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

The 32x path: optimized scalar quantization on both engines

At the 32x end of the dial, OpenSearch Service uses optimized scalar quantization (OSQ), a RaBitQ-derived technique that compresses float32 vectors down toward a bit per dimension while holding recall far better than naive binary rounding. Where naive binary quantization centers each vector on the global centroid and rounds, OSQ adds a calibration step: it computes statistics, optimizes the quantization intervals, and keeps small corrective factors with each vector that sharpen the distance estimates. The payoff is a better compression-to-quality ratio at the same 32x reduction. You do not pick the algorithm by hand. Starting in OpenSearch 3.7, asking for 32x compression selects OSQ automatically, and OpenSearch applies it across both the Faiss and Lucene engines, handling the underlying mechanism for you. Set the compression level and let the engine choose.

The reason OSQ holds recall at roughly a bit per dimension is a set of techniques OpenSearch applies for you, with no knobs to turn. It keeps small corrective factors with each quantized vector so the first-pass distances stay close to the true ones. It uses asymmetric distance computation, keeping the query vector at full precision and rescaling it to compare meaningfully against the compressed document vectors, so the query side keeps information the documents gave up. And it applies a random rotation that spreads variance across dimensions, so more signal survives the compression. Each of these was once a separate option to wire up by hand; at 32x, OpenSearch does all of it under the covers.

PUT /my-osq-index
{
  "settings": { "index.knn": true },
  "mappings": {
    "properties": {
      "embedding": {
        "type": "knn_vector",
        "dimension": 1536,
        "space_type": "l2",
        "compression_level": "32x"
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Product quantization, briefly

Product quantization (PQ) is the one native method that works differently enough to call out. Instead of reducing the precision of each dimension, PQ splits each vector into m subvectors and encodes each one against a learned codebook, reaching compression ratios beyond what scalar quantization offers. The cost is a training step: PQ runs k-means over a representative sample to build the codebooks before you can index, which scalar quantization does not require. That training requirement, plus a larger accuracy gap at high compression, makes PQ a fit for very large corpora, hundreds of millions to billions of vectors, where the savings justify the added operational step. For the parameters and the training workflow, see the OpenSearch product quantization documentation.OpenSearch product quantization documentation

Oversampling to recover accuracy

The low-bit methods search over reduced-precision vectors, so the first-pass scores are approximate. Across these methods, you raise recall at query time with oversampling: set an oversample_factor and OpenSearch retrieves that multiple of k candidates before ranking, trading latency for accuracy, with 2x to 4x a reasonable starting point. The same rescore.oversample_factor parameter applies whether you are running scalar quantization or the optimized scalar quantizer at 32x. Measure recall and latency together on your own evaluation set, because the right oversampling depends on your data and your bit width.

GET /my-sq-index/_search
{
  "size": 10,
  "query": {
    "knn": {
      "embedding": {
        "vector": [0.1, 0.2],
        "k": 10,
        "method_parameters": { "ef_search": 100 },
        "rescore": { "oversample_factor": 3.0 }
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

When exact k-NN is the better call

In-engine quantization still builds and searches an HNSW graph, and that graph is what holds memory. When your queries pre-filter to a small candidate set, you can skip the graph entirely. Exact k-NN scores the query against every vector that passes the filter, with no graph to build or hold in RAM, and when a structured filter narrows the field to roughly 100,000 vectors or fewer it returns perfect recall at latency competitive with the approximate path. It stacks with quantization too: a quantized exact scan over a filtered subset is both small and precise. Reach for exact k-NN when the access pattern is search within this tenant, this category, this date range rather than search the whole corpus. The graph earns its memory only when the candidate set stays large.

Pick the method, keep the pipeline

The native methods trade a little configuration for a large cut in memory while your ingestion stays exactly as it was. Scalar quantization at 16 bits halves memory at recall close to full precision, and the 2- and 4-bit settings take you to 16x and 8x. At the 32x end, you set the compression level and the optimized scalar quantizer takes over on either engine, recovering recall with corrective factors, asymmetric distance computation, and random rotation that you never configure. Product quantization serves the largest corpora, and oversampling tunes recall across all of them. When your queries filter down to a small set, exact k-NN sidesteps the graph cost altogether.

The next article moves from compressing the vectors to moving them off memory. On-disk mode keeps a compressed graph in RAM and the full-precision vectors on disk, with a two-phase search that rescores against those full-precision vectors, and cost-stacking techniques layer on top of any method here for more savings.

Top comments (0)