FP16, INT8, and binary. Pick your spot on the accuracy-cost dial before Amazon OpenSearch Service ever sees the vector.
If you use vectors for search, you are always trading off cost, latency, and accuracy. OpenSearch Service gives you a set of knobs to adjust that balance, depending on where you want the emphasis. The one that dominates the bill is memory, because the approximate-nearest-neighbor graph that makes vector search fast lives in RAM, and RAM is the resource you provision the most of. How much you need comes down to one choice you make before indexing a single document: how much precision each vector carries.
At 100 million vectors of 1,536 dimensions, you need about 643 GB for a single copy. Run it with one replica, which most production deployments do, and the resident footprint is about 1.3 TB. That full footprint is what you deploy to hold every vector in memory for the best latency; with memory-optimized loading you can run in less RAM and page from disk on demand, trading some latency for a smaller bill. That number is not fixed either way. You decide how much of it to pay for, by choosing how much precision each vector carries before OpenSearch Service ever stores it, and you make that choice knowing what accuracy it costs.
Quantization is a dial with a continuous range of settings, each trading a measured amount of accuracy and latency for a lower memory bill. This is the first of three articles walking the full spectrum available in Amazon OpenSearch Service. Here you shrink the vectors yourself before they reach the engine. The next two cover OpenSearch-native quantization, where the engine compresses for you, and on-disk mode, where a compressed copy stays in memory and the full-precision vectors move to disk for a second pass that restores accuracy.
The baseline: full-precision FP32
Every compression option is measured against the full-precision baseline, so start there. A single FP32 vector at 1,536 dimensions holds 1,536 values of 4 bytes each, which is 6,144 bytes. The OpenSearch estimate for Faiss HNSW adds the graph that connects those vectors and multiplies the whole by a storage factor: 1.1 times the quantity 4 times dimension plus 8 times m. The 4 times dimension term is the raw vector, 4 bytes per FP32 dimension. The 8 times m term is the graph: m is how many neighbor links HNSW keeps per vector (m defaults to 16), and each link is stored as an 8-byte node identifier, so the graph adds 8 times m bytes per vector. The 1.1 is the source-to-index ratio, the stored index running about ten percent larger than the raw input.
memory_per_vector = 1.1 x (4 x dimension + 8 x m)
= 1.1 x (4 x 1,536 + 8 x 16)
= 1.1 x (6,144 + 128)
= ~6,899 bytes/vector
Per copy: 100,000,000 x 6,899 bytes = ~643 GB
With 1 replica: 643 GB x 2 = ~1,285 GB (~1.3 TB)
The raw vectors dominate that number. A replica holds a full second copy of the graph in memory for availability and query throughput, so every footprint figure in this article is per copy, and you double it for one replica. 1.3 TB is the resident anchor for a replicated 100-million-vector index. Every option below reduces it.
Turn 1.3 TB into instances to see the bill. Take the r8gd.2xlarge, with 64 GB of RAM. By default OpenSearch Service gives half of that to the JVM heap, leaving 32 GB off-heap, and the k-NN memory circuit breaker (knn.memory.circuit_breaker.limit) caps native vector memory at 50% of what remains, about 16 GB per instance. At 16 GB usable, 1.3 TB needs roughly 80 instances. If the node is dedicated to vector search and not running heavy text queries that lean on the heap, raise the circuit breaker to 75% of off-heap memory, which lifts usable vector memory to about 24 GB per instance and brings the count down to roughly 54 instances. On demand, 54 to 80 r8gd.2xlarge instances runs on the order of $31,000 to $46,000 per month, before reserved-instance or savings-plan discounts. Treat the dollars as order-of-magnitude. The point is that full precision at this scale is a five-figure monthly line item, and the dial below moves it.
Pre-quantized vectors: you shrink, OpenSearch stores
The first category of compression happens before your vectors reach OpenSearch Service. You, your embedding model, or a step in your ingestion pipeline produce vectors in a reduced-precision format, and OpenSearch Service stores and searches exactly what you send. The engine does no conversion here. The appeal is control: you pick the method, you measure recall against your own evaluation set, and nothing about the accuracy tradeoff is hidden. Three formats are available, and they sit at different points on the dial: FP16, INT8, and binary. The footprints below are per copy; double each for a replica, the same as the baseline.
FP16: half the bytes, nearly identical recall
Half-precision floating point stores each dimension in 16 bits instead of 32. Per-vector size drops by exactly half, from 6,144 to 3,072 bytes, and the 100-million-vector footprint falls from about 643 GB to roughly 328 GB per copy, about 655 GB with a replica. Many embedding models carry little useful signal in the low-order bits of each FP32 dimension, so truncating to FP16 costs almost nothing in quality. OpenSearch benchmarks put FP16 recall at effectively indistinguishable from full precision, because most datasets do not use the full 32-bit range. Some models emit FP16 natively, in which case there is no conversion step to add.
In OpenSearch Service, the data type for native FP16 is half_float, introduced in OpenSearch 3.9 and supported for the HNSW method on both the Faiss and Lucene engines. You ingest and query half_float vectors the same way as float vectors, and each dimension must fall within the representable range of roughly plus or minus 65,504. The mapping:
PUT /my-vector-index
{
"settings": { "index.knn": true },
"mappings": {
"properties": {
"embedding": {
"type": "knn_vector",
"dimension": 1536,
"data_type": "half_float",
"space_type": "l2",
"method": {
"name": "hnsw",
"engine": "faiss"
}
}
}
}
}
INT8 byte vectors: a quarter of the footprint
Byte vectors store each dimension as a signed 8-bit integer in the range -128 to 127. At one byte per dimension, a 1,536-dimension vector occupies 1,536 bytes, a quarter of the FP32 baseline, and the footprint drops to roughly 170 GB per copy, about 341 GB with a replica. The cost is a conversion step: you map your FP32 embeddings into the byte range before indexing. A linear scaling from the observed min and max across a representative sample works for many models, and calibrated or per-channel approaches narrow the accuracy gap further. OpenSearch benchmarks on byte vectors report recall@100 of 0.99 on Euclidean datasets, matching full precision, with angular datasets measured by cosine similarity giving up more and benefiting most from careful calibration. Measure it against your own queries before you commit.
PUT /my-byte-index
{
"settings": { "index.knn": true },
"mappings": {
"properties": {
"embedding": {
"type": "knn_vector",
"dimension": 1536,
"data_type": "byte",
"space_type": "l2",
"method": {
"name": "hnsw",
"engine": "lucene"
}
}
}
}
}
Binary vectors: 32x smaller, a different bargain
Binary is the aggressive end of the pre-quantization dial. Each dimension collapses to a single bit, one if the value is non-negative and zero otherwise. A 1,536-dimension vector becomes 1,536 bits, which is 192 bytes, 32 times smaller than FP32, and the 100-million-vector footprint lands near 33 GB per copy, about 66 GB with a replica, small enough to fit on a single large instance either way. The tradeoff is that binary discards magnitude entirely. For high-dimensional embeddings whose vectors are well separated by angle, which is common for modern text embeddings, binary retains a surprising share of the original recall. OpenSearch large-scale tests with rescoring hold recall@100 in the 0.94 to 0.98 range across 8x, 16x, and 32x compression, close to the in-memory baseline. For embeddings that lean on subtle magnitude differences, the drop is steeper. Always benchmark against your own data.
When you convert to binary yourself and store the result, you pack the bits into bytes: eight bits become one int8 value, so the dimension you declare must be a multiple of 8, and the engine compares vectors with Hamming distance. This pre-converted path is Faiss-only. If you would rather hand FP32 vectors to the engine and let it compress them, OpenSearch does that for you at the aggressive end of the dial through the optimized scalar quantizer, which is the subject of the next article.
A production pattern pairs a fast binary first pass with a full-precision second pass that restores ordering. OpenSearch Service implements that pattern for you in on-disk mode rather than leaving you to wire up a separate full-precision field, so the mechanics are covered in the third article of this series.
When exact k-NN beats the graph entirely
Quantization assumes you are running approximate search over an HNSW graph, and that graph is the thing eating your RAM. If your queries pre-filter to a small candidate set, you may not need the graph at all. Exact k-NN scores the query against every vector that passes the filter, with no graph to build, hold in memory, or rebuild on merge. When a structured filter narrows the field to roughly 100,000 vectors or fewer, exact search returns perfect recall at latency that is competitive with the approximate path, and it stacks cleanly with quantization: a byte or binary exact scan over a filtered subset is both inexpensive and precise. Reach for exact k-NN when your access pattern is 'search within this tenant, this category, this date range' rather than 'search the whole corpus.' The graph earns its memory only when the candidate set stays large.
Dimensionality reduction stacks with all of it
Precision is one axis. Dimension count is another, and the two multiply. Reducing 1,536-dimension embeddings to 768, whether through a model that emits shorter vectors or a reduction step such as principal component analysis, halves memory on its own and composes with any format above. FP16 at 768 dimensions is half of FP16 at 1,536. Binary at 768 is half of binary at 1,536. The dial is really a surface, precision on one axis and dimensionality on the other, and you can move along both.
Pick your spot on the dial
Pre-quantized vectors are the easiest segment of the spectrum to reason about, because you own the conversion and you measure the result. For many workloads FP16 alone halves the memory bill at recall indistinguishable from full precision, and that is often the end of the story. Byte vectors take you to a quarter of the baseline when a few points of recall are acceptable. Binary takes you to a thirty-second when your embeddings tolerate it and a rescoring pass covers the rest. And when your queries filter down to a small set, exact k-NN sidesteps the graph cost altogether.
When you want more compression without managing the conversion yourself, OpenSearch Service will do the work. The next article covers OpenSearch-native scalar and binary quantization, where you send full-precision vectors and the engine compresses them at index time, with oversampling to recover accuracy. After that, on-disk mode keeps a compressed copy in memory and moves the full-precision vectors to disk for a second-pass rescore, the lowest-cost end of the dial.
Top comments (0)