DEV Community

Cover image for I Tested Amazon S3 Vectors' New Pre-Filtering Against Exact Ground Truth

I Tested Amazon S3 Vectors' New Pre-Filtering Against Exact Ground Truth

A filtered vector query that asks for 10 results can come back with 2, and nothing errors. AWS's own launch post shows exactly that on a CLASSIC index: a query scoped to one tenant "returns two of the ten results requested". The rest of that tenant's matching documents just aren't in the answer.

On Sep 30, 2026, AWS launched metadata pre-filtering for Amazon S3 Vectors (the ENHANCED index mode), saying it fixes exactly this failure on selective filters. I wanted to know what really changes, so I measured it. I loaded 50,000 vectors into a real ENHANCED index, ran 6,900 logical query requests (6,000 of them filtered, across 10 filters, one of which matches nothing) at 4 values of K, and scored every result against exact nearest neighbors computed locally. The scope is narrow: one synthetic index in one region, 50 query vectors, and no real CLASSIC measurement (AWS doesn't allow CLASSIC on new buckets).

The short version:

  • ENHANCED returned the exact answer on every filter with 1 to 5 matches, where a post-filtering approach collapsed.
  • Between 50 and 135 matches, its recall dipped measurably below the unfiltered search.
  • On broad filters, ENHANCED and a large-budget post-filter agreed.
  • I could not get CLASSIC at all on a new bucket.

The concept, briefly

In CLASSIC mode, S3 Vectors applies the metadata filter during the vector search, so a selective filter "may return fewer than top K results". In ENHANCED mode, it finds the vectors that match the filter first and then searches only those. Buckets created on or after Sep 30, 2026 create ENHANCED indexes, and older buckets keep CLASSIC until you change them.

AWS reports that on highly selective filters, pre-filtering "returns up to 5x more of the matching vectors than the same query returned before on CLASSIC indexes" (AWS News Blog). That's a ratio of result counts. It isn't a Recall@K multiplier. This article doesn't re-explain the feature; it measures it.

Three query paths: filter first (ENHANCED, as AWS documents it), filter during search (CLASSIC; the candidate bound is an assumption), and the client-side post-filter baseline used in this experiment
Figure 1. Three query paths. Only the first and third were run on AWS. The CLASSIC row shows AWS's documented behavior plus an assumed candidate bound. The bottom row is my baseline, and it is not CLASSIC.

Setup and methodology

The dataset:

  • It is deterministic and synthetic: N = 50,000 vectors, 384 dimensions, cosine distance, seed 20260930. The 50 query vectors come from the same 64 clusters but a separate random stream, so they are not vectors from the index.
  • It was loaded into one disposable ENHANCED index in us-east-1 on Oct 5, 2026 (run 20261005t0953-a9ba).
  • The filters use the metadata keys tenant, category and year. Tenant labels are independent of vector position, and no keys were configured as non-filterable.
Filter Definition Matching vectors Selectivity
F50 tenant = tenant-s50 25,000 50%
F10 tenant = tenant-s10 5,000 10%
FAND3 $and: tenant-s10, category c2, year ≥ 2020 656 1.312%
F1 tenant = tenant-s1 500 1%
FAND2 $and: tenant-s1, category c1 135 0.27%
F01 tenant = tenant-s01 50 0.1%
F001 tenant = tenant-s001 5 0.01%
FFEW tenant = tenant-few 3 0.006%
FONE tenant = tenant-one 1 0.002%
FZERO tenant = tenant-absent 0 0%

Table 1. Filters. Counts and percentages were measured on the canonical local data.

Three kinds of query ran against the real index, all with 50 query vectors and K ∈ {5, 10, 20, 50}.

Query class (REAL AWS) What it is Requests
ENHANCED filtered query QueryVectors with the filter 6,000 (3 repeats)
Unfiltered ANN reference QueryVectors, no filter 600
Client-side post-filter baseline Unfiltered query at topK = B ∈ {100; 1,000; 10,000}, then the filter applied locally, keeping the first K matches 300 logical requests, about 11,100 pages (2 repeats)

Table 2. Query classes. All rows are REAL AWS measurements.

How results were scored:

  • Ground truth is brute force, computed locally over all 50,000 vectors. I apply the filter to the canonical data, compute exact cosine distances to every matching vector, and take the top K.
  • Recall@K = |returned ∩ exact top-K| / min(K, matching).
  • Completeness = returned / min(K, matching).
  • For the zero-match filter, recall is undefined. I report whether the query correctly returned nothing.
  • 95% confidence intervals are bootstrapped over the 50 query vectors.

The post-filter baseline is not CLASSIC, and it does not reproduce CLASSIC. It's what you could build yourself if you didn't trust the service filter. B = 10,000 is the documented topK maximum. Here that's 20% of N, which makes this baseline far more generous than it would be on a production-size index.

Console: index details showing Enhanced index mode enabled, 384 dimensions, cosine
Figure 2. The benchmark index in the S3 Console: Enhanced index mode enabled, dimension 384, cosine. ENHANCED is the documented default for a bucket created on or after Sep 30, 2026. The account segment of the ARN is blanked.

Result 1: on the most selective filters, ENHANCED is exact and post-filtering collapses

Filter Matches ENHANCED Recall@10 Post-filter B=10,000 B=1,000 B=100
FONE 1 1.000 0.420 0.000 0.000
FFEW 3 1.000 0.280 0.053 0.013
F001 5 1.000 0.240 0.008 0.004
F01 50 0.898 [0.874, 0.922] 0.836 0.131 0.008
FAND2 135 0.932 0.932 0.280 0.034

Table 3. REAL AWS, K = 10. Mean Recall@10 over 50 query vectors. The post-filter columns are my client-side baseline, not AWS CLASSIC.

Results on the most selective filters:

  • For the filters with 5 or fewer matches, ENHANCED recall was 1.000 at every K, for every query vector.
  • ENHANCED completeness was 1.000 at K = 10 on every filter in Table 3.
  • The post-filter baseline found the single FONE match for only 42% of query vectors, even after scanning 10,000 results (20% of the index).

REAL AWS: Recall@10 vs measured selectivity for ENHANCED and the post-filter baseline at three budgets
Figure 3. REAL AWS. Recall@10 vs measured selectivity (log scale), with 95% CIs. The zero-match filter is omitted. All four series were measured on the same ENHANCED index.

REAL AWS: completeness vs measured selectivity for ENHANCED and the post-filter baseline
Figure 4. REAL AWS. Completeness (returned / min(K, matching)) at K = 10. ENHANCED fills the result at every selectivity. The baseline's fill rate depends on B × selectivity.

The baseline's returned count depends almost entirely on B × selectivity. For example, F01 at B = 10,000 and K = 50 returned 10.86 matches on average, which is about 20% of the 50 that exist. With only 1 to 5 matching vectors the share is noisier (FONE 42%, FFEW 28%, F001 24%), because it depends on where those few vectors sit relative to just 50 query vectors. ENHANCED has no B, so its count doesn't depend on one.

Result 2 (negative): ENHANCED recall dips at 0.1–0.3% selectivity

This was the result I didn't expect. Inside the filtered set, ENHANCED recall falls as K grows, while completeness stays at 1.000 for K ≤ 20.

Filter Matches Recall@5 Recall@10 Recall@20 Recall@50
F01 50 0.948 0.898 0.811 0.610
FAND2 135 0.952 0.932 0.883 0.724

Table 4. REAL AWS, ENHANCED filtered query.

For comparison:

  • The unfiltered ANN reference on the same index had Recall@10 = 0.994.
  • Filters with 5 or fewer matches had recall 1.000 at every K. Broad filters were approximate too, but close to the reference (F50 at K = 10: 0.984).

So in this experiment ENHANCED recall was not monotonic in selectivity. At K = 10 it dipped around 50–135 matches and recovered on broader filters (F1: 0.974; F10: 0.992). At K = 50 the drop reached further: F1 scored 0.911 and FAND3 0.914, against 0.992 for the unfiltered ANN reference at K = 50, while F10 (0.996) and F50 (0.991) did not drop.

There was also one underfill. With F01 and K = 50, where matches = K = 50, ENHANCED returned 30.52 of the 50 matching vectors on average (completeness 0.610, recall 0.610). It returned the identical result set on all 3 repeats. No other ENHANCED row had completeness below 1.000. AWS doesn't document ENHANCED's search internals, so I don't know the cause.

One possible explanation is that the search inside the filtered set is still approximate and has a working-set bound that a 50-match set at K = 50 runs into. I have not tested this. Even here, ENHANCED was well ahead of the baseline: the post-filter baseline at B = 10,000 returned 10.86 and scored recall 0.217 on the same row.

Where the differences disappear

Broad filters. When B ≥ 1,000, ENHANCED and the post-filter baseline agree to three decimals at every K.

Row ENHANCED Post-filter B=1,000 Post-filter B=10,000
F50 / K=10 0.984 0.984 0.984
F10 / K=10 0.992 0.992 0.992
F10 / K=50 0.996 0.996 0.996

Table 5. REAL AWS, mean Recall@K.

More rows where the two methods agree:

  • For F50, even B = 100 matched ENHANCED at K ≤ 20.
  • At 1% selectivity and on the 3-condition $and, B = 10,000 matched ENHANCED within 0.001 (F1/K=10: 0.974 for both; F1/K=50: 0.911; FAND3/K=10: 0.984).
  • FAND2 at K ≤ 20 was identical between ENHANCED and B = 10,000 (0.952, 0.932, 0.883).

The takeaway: on filters that match a large share of the index, this experiment shows no recall difference between ENHANCED and a client-side post-filter with a large enough budget. It says nothing direct about CLASSIC, which I couldn't run. AWS documents CLASSIC's shortfall only for selective filters: on CLASSIC "a query with a selective filter can return fewer than top K results".

Latency (client round trip). Warm ENHANCED medians ranged from 329.2 to 330.2 ms across every filter, from 50% down to a single match. The unfiltered ANN reference was 329.9 ms. These are client-observed round trips that include the network, so at this index size any server-side effect of selectivity is below what the round trip can resolve. They are not AWS server latency. AWS documents that ENHANCED query work grows with index size, with the share of vectors the filter matches, and with the number of filter constraints. At 50,000 vectors I couldn't see that effect from the client.

REAL AWS: client-observed round trip vs selectivity, first and warm phases
Figure 5. REAL AWS. Client-observed round trip including the network, not AWS server latency. "first" means repeat 1 and is not claimed to be cold. Post-filter latency is the sum of its sequential pages. 1 of 6,900 requests is excluded (see Limitations).

The post-filter baseline is slow because it paginates: its warm median was 3,350.5 ms at B = 1,000 and 33,553.8 ms at B = 10,000. That's a cost of the method, not of the service.

Correctness checks:

  • The zero-match filter returned 0 results in all 600 ENHANCED requests.
  • Every REAL AWS row had 0 precision violations (no returned vector failed the filter) and 0 duplicate keys.
  • Repeat consistency was 1.000 in most rows. Four ENHANCED rows had 0.98 (F1/K=20, F50/K=50, FAND2/K=50, FAND3/K=50), with mean Jaccard ≥ 0.9987. Six post-filter baseline rows at B = 1,000 had 0.980, 48 of 49 cells (F01 at every K, F1 at K = 20 and 50).
  • Ties at the K boundary are reported, not excluded: 3 in the unfiltered ANN reference at K = 20, 3 in ENHANCED F10 at K = 5, and 2 per budget in the post-filter baseline F10 at K = 5.

Migration implications: CLASSIC isn't available on new buckets

I tried to get CLASSIC behavior on a bucket created on Oct 5, 2026 every documented way. AWS rejected each attempt with HTTP 400 ValidationException.

Step Call AWS error message
3 PutVectorBucketDefaultIndexMode CLASSIC "defaultIndexMode cannot be set to CLASSIC because this vector bucket does not support CLASSIC mode"
8 UpdateIndexMode CLASSIC "indexMode cannot be set to CLASSIC because the vector bucket does not support CLASSIC mode"
10 QueryVectors queryMode=CLASSIC (probe index) "queryMode cannot be CLASSIC when indexMode is ENHANCED"
13 QueryVectors queryMode=CLASSIC (main index) "queryMode cannot be CLASSIC when indexMode is ENHANCED"

Table 6. REAL AWS CLASSIC probe, recorded in results/aws/probe_classic.json.

This matches the documentation: CLASSIC can only be set for an index in a bucket created before Sep 30, 2026, and queryMode CLASSIC can't be used on an ENHANCED index. GetIndex reported indexMode: ENHANCED for the probe index after the UpdateIndexMode CLASSIC attempt, and for the main index, whose only CLASSIC attempt was the rejected queryMode.

Console: probe bucket Properties tab still showing Enhanced index mode enabled after the CLASSIC attempts
Figure 6. The probe bucket after the CLASSIC attempts. AWS rejected PutVectorBucketDefaultIndexMode CLASSIC, and the bucket still shows Enhanced index mode enabled. The account segment of the ARN is blanked.

What this means for migration:

  • On new buckets there is no CLASSIC at all, so there's nothing to fall back to and nothing to A/B against on the same index.
  • If you have a pre-Sep-30 bucket, it's your only place to compare the modes. Test there before you flip, because AWS documents that switching back to CLASSIC works only through the CLI, SDKs, or REST API, not the console.

The 100-constraint limit is real. AWS documents that ENHANCED allows at most 100 filter constraints per query, with each evaluated value counting as one. An $in with 100 values (filter C100, matching 19,441 vectors) was accepted and returned 10 results. With 101 values (C101), the query was rejected with ValidationException "Filter must have at most 100 constraints". So a query that worked on CLASSIC can fail after migration if it uses long value lists. AWS's suggested fixes are to consolidate those values into a grouping key, or to split the filter into parallel queries and merge the results by distance.

How my baseline ratios compare with AWS's "up to 5x"

I can't test AWS's 5x claim directly, because I have no real CLASSIC measurement. AWS doesn't publish the methodology behind the figure. The What's New headline says "5x higher recall", while the blog and the What's New body say "5x more of the matching vectors".

The closest real ratio I have is ENHANCED returned ÷ post-filter returned. At K = 10, against the most generous baseline (B = 10,000):

Filter Ratio Returned (ENHANCED / baseline)
F001 4.2x 5.00 / 1.20
FFEW 3.6x 3.00 / 0.84
FONE 2.4x 1.00 / 0.42
F01 1.09x 10.00 / 9.20

Table 7. REAL AWS. These are ratios against my post-filter baseline, not against CLASSIC.

Against B = 100 (a single page), the ratios for Table 7's filters are 75x for FFEW (3 / 0.04), 125x for F01 (10 / 0.08) and 250x for F001 (5 / 0.02). FONE's single match never came back at B = 100, so it has no ratio. So "N times more matches" depends almost entirely on what you compare against, and none of my ratios measures AWS's 5x.

Possible reasons for any gap. None of these is a finding:

  • AWS's figure may come from a much larger index, where any fixed candidate budget covers a smaller fraction of N. The blog's own worked example is an 8-million-ticket knowledge base.
  • AWS's CLASSIC may evaluate the filter during search with a candidate bound unlike my B or the simulator's C.
  • "Up to" is a maximum over AWS's own test filters.
  • A matching-count ratio says nothing about ranking errors inside the filtered set. Result 2 shows that a full result and imperfect Recall@K can coexist.

SIMULATED: a mechanism model of CLASSIC (not AWS)

To show the mechanism, I built a numpy IVF model of both modes. CLASSIC is modeled as filtering during the search with a bounded candidate budget C, and ENHANCED as filter first, then search. AWS doesn't document its ANN internals, and nothing in this model is calibrated to AWS's 5x.

At K = 10 on F01, the SIMULATED CLASSIC model returned 0.70, 2.26 and 7.76 results at C = 500, 2,000 and 8,000. The SIMULATED ENHANCED model returned 10 at every C and scored recall 1.000 everywhere.

The real ENHANCED index didn't score 1.000 at 0.1–0.3% selectivity, so this model is more ideal than what I measured on AWS.

SIMULATED (numpy IVF model, not AWS): Recall@10 vs selectivity for CLASSIC and ENHANCED models
Figure 7. SIMULATED, not AWS. Recall@10 of a mechanism model at three assumed CLASSIC candidate budgets. These are not AWS numbers, and the model contains no timing.

Cost

The run's tally was $0.0731, against the run's spend cap of $0.374573 (the project's overall budget cap was $5). The tally is the runner's conservative model, built from the official S3 pricing page. It isn't an AWS bill.

AWS documents that pre-filtering has no additional charge and that standard S3 Vectors pricing applies: storage, PUT, and queries, where queries are billed as request fee + data processed + data returned. AWS defines data processed by index size, not by filter matches. I found no source saying ENHANCED changes it, so that remains unknown. I make no savings claim.

Limitations

  • There is no real CLASSIC data. Every CLASSIC statement here is either AWS documented or SIMULATED.
  • One failed request, caused by host sleep. Exactly 1 of 6,900 logical query requests failed.
    • It was a post-filter baseline request at B = 1,000.
    • Pages 1–5 ran between 10:49:27 and 10:49:29 UTC. Then my PC went to sleep, and page 6 was sent at 11:05:25 UTC.
    • The first attempt hit ConnectionClosedError. The retry got ValidationException "Invalid page token".
    • The request is recorded and excluded, which is why the B = 1,000 rows have 99 requests instead of 100. The retry warning is recorded in the repo's run log (results/aws/run_log.txt).
  • Synthetic data, one index, one region. 50,000 vectors at 384 dimensions, with metadata independent of vector position. Real embeddings, correlated metadata, and larger indexes may behave differently.
  • 50 query vectors. The confidence intervals are bootstrapped over them.
  • Latency is a client-observed round trip. The round trip was about 330 ms and includes network and client time, so it can't isolate server-side ENHANCED latency.
  • Unknowns. AWS's ANN internals, the cause of the F01/K=50 underfill and the mid-selectivity dip, and whether ENHANCED changes data-processed billing.

Recommendation

Based on this experiment:

  • When ENHANCED matters: filters that match a tiny slice of the index. A filter with 1 to 5 matches was exact with ENHANCED, and lost most of its matches with post-filtering, even at a 10,000-result budget.
  • When it doesn't: filters that match about 1% or more of a 50,000-vector index, compared against a 10,000-result post-filter (20% of this index). Recall was the same either way. At a 1,000-result budget, only the 10% and 50% filters matched ENHANCED at every K. F1 and FAND3 fell well behind at larger K (F1 at K = 50: 0.206 vs ENHANCED 0.911; FAND3 at K = 50: 0.256 vs 0.914).
  • What to watch: sets of roughly 50–135 matches. ENHANCED usually fills the result (the one exception was F01 at K = 50, 30.52 of 50) but misses some of the true nearest neighbors, more so at larger K. At K = 50 the shortfall also showed up at 500–656 matches (about 0.91 recall). If exact top-K inside a small tenant matters to you, measure it on your own data.
  • Before migrating: check every filter for the 100-constraint limit. Remember that CLASSIC isn't available on new buckets, so there is nothing to fall back to there.

A hypothetical example: a multi-tenant RAG service where a tenant owns 3 documents is the FFEW case. With ENHANCED, all 3 came back for every query vector. With a client-side post-filter over 10,000 results, 0.84 of the 3 came back on average (28%).

Reproduce it

The code, the raw evidence (every request, response, and timing), and the generators for the tables and figures are in the repo: https://github.com/Extraordinarytechy/s3vectors-prefilter-bench. The offline steps need no AWS credentials.

python3 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python -u -m pytest -q tests
.venv/bin/python -u -m src.runner offline-all     # generate -> ground-truth -> simulate -> report
.venv/bin/python -u -m src.runner estimate        # writes results/aws/cost_estimate.md
Enter fullscreen mode Exit fullscreen mode

To run the real AWS part (this creates billable resources; read the cost estimate first):

.venv/bin/python -u -m src.runner aws-probe   --aws --profile default --region us-east-1 --confirm-cost
.venv/bin/python -u -m src.runner aws-ingest  --aws --profile default --region us-east-1 --confirm-cost --run-id <id>
.venv/bin/python -u -m src.runner aws-query   --aws --profile default --region us-east-1 --confirm-cost --run-id <id>
.venv/bin/python -u -m src.runner aws-cleanup --aws --profile default --region us-east-1 --run-id <id>
Enter fullscreen mode Exit fullscreen mode

Cleanup deletes only the resources recorded for that run and verifies the deletion with list calls. I ran it after capturing the screenshots, and both buckets from this run are gone.

What's still open

In one line: pre-filtering fixed the missing-matches problem on tiny filters, but it doesn't make the search exact. When a filter matched only a handful of vectors, ENHANCED found all of them while a large-budget post-filter missed most. When a filter matched about 1% or more of the index, the two agreed. Between 50 and 135 matches, ENHANCED still missed some true nearest neighbors.

Three questions this run can't answer:

  • How ENHANCED compares to real CLASSIC on the same index. That needs a vector bucket created before Sep 30, 2026. The repo has a borrowed-bucket mode for this: it creates one disposable index in an existing bucket and never modifies anything else in it. It's covered by the test suite, but I haven't run it against a real pre-Sep-30 bucket.
  • Whether the dip at 50–135 matches holds on real embeddings and larger indexes.
  • What causes the F01 / K = 50 underfill.

If you have a pre-Sep-30 bucket and try it, I'd like to see your numbers in the comments.

Top comments (1)

Collapse
 
ahmetozel profile image
Ahmet Özel •

Would the mid-selectivity dip persist if tenant labels were correlated with the vector clusters? Your independent metadata assignment is a clean control, but a real tenant corpus may occupy a much narrower region of the embedding space, which could change the difficulty at the same matching count.

I would keep N, K and eligible-set size fixed while comparing random labels with cluster-aligned labels. The F01/K=50 row is especially diagnostic because every eligible vector belongs in the result, so underfill can be studied separately from nearest-neighbour ordering. That would test the distribution effect without assuming an undocumented internal search budget.