Over the past two years a familiar story has played out across the industry: a team reads a benchmark blog post, concludes Postgres cannot possibly serve its embeddings, and spends days migrating to a dedicated vector database. The migration adds a second database, a sync job, and a new failure mode. Queries get measurably faster. Nobody using the product notices.
That pattern came to mind when I read turbopuffer's September 30 post, RIP, vector database. This is not a hot take from a startup nobody uses. turbopuffer is the search layer behind Cursor and Notion. Cursor manages billions of vectors across millions of codebases on it. Notion runs more than 10 billion vectors over millions of namespaces. If anyone has earned the right to say what a vector database should be, it is the company serving those workloads. And their conclusion is blunt: they are removing the vector index as the primary data structure in their next major version.
If the most successful vector database of this cycle is demoting vectors, the question worth asking is the one your architecture should have answered in the first place: do you actually need a dedicated vector database, or was Postgres fine all along?
I have not used turbopuffer myself. Everything about their architecture below comes from their own blog posts and the Hacker News thread, linked inline. This is a sourced breakdown of both sides of the decision, not a migration story.
What turbopuffer actually announced
Read the title and you could easily assume vector search is dying. It is not. The announcement is narrower and more interesting.
- The vector index loses its throne, not its job. In turbopuffer v3, documents get keyed by a stable document ID, and the approximate nearest neighbor index becomes "just another" secondary index. Vector search keeps working. The claim they make elsewhere: single indexes of 100B+ vectors, 200 ms p99 reads, 1k+ QPS.
- The reason is write amplification. Their current layout stores a document's full contents keyed by its vector's cluster address. When the clustering index rebalances, which happens on routine inserts and updates, the whole document and every inverted index referencing it moves too. Their words: updating one vector can move hundreds of attributes.
- The economics stay object-storage-first. Their original pitch was roughly $70 per TB per month with an SSD cache versus around $1,600 for cache-plus-SSD incumbents. Object storage as source of truth is what got them Cursor and Notion. Nothing in v3 changes that.
- They admit v3 is currently slower. As of early October, 100% of CI passes, but they describe the new engine as a significant performance regression that they are only starting to tune. This is a direction, not a shipping product.
So the real headline: even a company whose entire original identity was vector-first storage hit a wall where treating vectors as special made everything else worse. Storage amplification for multi-vector documents, write amplification on every rebalance, and query plans locked to cluster-sized blocks of 100 to 200 documents, when engines like DuckDB work in 2,048-row batches and ClickHouse up to around 65k. That is the engineering argument. Now the decision it should trigger in your stack.
Postgres vs the dedicated vector database: what each side actually claims
- The Postgres-first claim: pgvector is included with your existing database, supports HNSW and IVFFlat indexes, handles the hybrid queries vector databases struggle with, and costs nothing extra. Supabase benchmarked pgvector on a 16-core, 256 GB server and got about 1,800 QPS at 0.91 accuracy@10 on a 1M-vector dataset.
- The dedicated-DB claim: beyond roughly 100 million vectors, purpose-built engines win on latency and memory. HNSW graphs are memory-hungry, Postgres has no built-in sharding for vector workloads, and managed engines handle the tiering that keeps RAM costs survivable.
- The turbopuffer data point both camps miss: their whole architecture exists because single-digit-millisecond disk-based search at massive scale is mostly a storage economics problem. When 1M 768-dim vectors fit in 3 GB, cold queries still land around 444 ms p90 from object storage, and warm ones hit 10 ms p90. Tiering beats brute force RAM at a fraction of the cost.
Now the part that matters more than any benchmark, and the question nobody's blog post asks: what does the rest of your query look like?
The SQL that decides it: when Postgres was always the answer
Here is the situation that keeps repeating across teams. You need "find me similar documents, but only in this workspace, only in this category, that the user has not already dismissed, joined back to the documents table for display."
In a dedicated vector database, that query shape is awkward by design. Vector engines are built around one primary query: nearest neighbors, optionally with attribute filters. The join back to your relational data happens in application code, often as a second round trip. In Postgres it is one query:
SELECT d.id, d.title, d.snippet,
e.embedding <=> $1 AS distance
FROM embeddings e
JOIN documents d ON d.id = e.document_id
WHERE d.workspace_id = $2
AND d.category = ANY($3)
AND NOT EXISTS (
SELECT 1 FROM dismissed_items x
WHERE x.user_id = $4
AND x.document_id = d.id
)
ORDER BY e.embedding <=> $1
LIMIT 20;
Filters, joins, permissions, and ranking in one statement, inside one transaction, backed by one backup and restore story. This is the workload where Postgres does not just compete with dedicated vector databases, it is the better tool by design.
The common counter-argument is "but the ANN index gets slower when filters are selective." True, and there is a standard fix:
CREATE INDEX ON embeddings
USING hnsw (embedding vector_cosine_ops)
WHERE workspace_id = 42;
A partial HNSW index per high-traffic workspace keeps the graph small and the recall high. pgvector's partial index support is the underrated feature in this entire debate. For the full pattern, pgvector's README documents iterative index scans and exact fallback when post-filtering would otherwise starve your LIMIT, so filtered queries stay correct, not just fast.
One more thing the benchmarks never model: the embeddings table is not the only thing touching those rows. When the document changes, you want to re-embed it and delete stale vectors in the same transaction as the content update. In two-database architectures, that consistency is a distributed systems problem. In Postgres, it is a WHERE clause.
When you genuinely need the dedicated engine
Honesty requires the other side of the ledger. Three situations where Postgres is the wrong choice, and I say that as someone who argues for Postgres by default.
- Multi-tenant at Notion scale. When you have millions of namespaces where only a fraction are hot at any moment, object-storage-first tiering is not an optimization, it is the business model. That is exactly turbopuffer's wedge, and Notion's workload is the proof.
- Very large single indexes. The 100M+ vector mark is where the HNSW graph in Postgres starts demanding more RAM than the dataset itself. If your workload is a few giant indexes at sustained high QPS, dedicated engines and their memory-tiering win.
- Late interaction and multi-vector documents. ColBERT-style retrieval stores dozens of vectors per document. This is precisely the storage amplification problem turbopuffer v3 exists to fix, and it is a terrible fit for row-based Postgres storage too.
Notice the pattern in that list. These are scale problems. If you are arguing about them before launch, you are arguing about the wrong thing. Here is the checklist I now apply before any migration discussion.
The decision checklist
- Under 1M vectors: Postgres, full stop. pgvector exact search needs no index at all and gives perfect recall. This covers the majority of RAG apps.
-
1M to 10M vectors: Postgres with HNSW, tune
hnsw.ef_search, use partial indexes for hot tenants. Benchmark with your own data, not blog numbers. - 10M to 100M vectors: Postgres still viable with partitioning and careful memory budgeting, but this is where you should run a real bake-off. pgvectorscale extends the ceiling before you add a second database.
- 100M+ vectors or millions of tenants: dedicated engine territory, and tiered object-storage architectures deserve a hard look.
- Heavy relational query shape (joins, filters, permissions per row): Postgres advantage at any scale, because the alternative is shipping join logic into your app.
- Multi-vector per document or ColBERT-style retrieval: dedicated engine, and not pgvector.
The threshold that matters most is not vector count. It is the question from earlier: does your retrieval query look like SQL, or does it look like nearest-neighbors-plus-a-prayer? Benchmarks measure the second shape. Most applications need the first.
What turbopuffer's retreat tells us about the next two years
The vector database gold rush of 2023 sold "your data is different now, you need new infrastructure." Three years later, the leading vector-first database is re-architecting around exactly the design Postgres people always argued for: stable document IDs as the primary key, vectors as one index among several. Turbopuffer arrived at the relational model from the vector side. Postgres arrived at vector search from the relational side. They are meeting in the middle, and that convergence tells you where the industry actually landed.
Meanwhile the durable trend is consolidation. Supabase, Neon, and the managed Postgres providers all ship pgvector by default now, and the default answer to "where do my embeddings live" keeps collapsing into "next to the rest of my data." The teams still running separate vector infrastructure are increasingly the ones with the three scale signatures above, not the ones who read an exciting benchmark in 2024.
The pattern behind most of these migrations is the same: the team benchmarked raw nearest-neighbor speed when their actual workload was the filtered query shape above. Measure your query, not the marketing.
I write about databases, AI infrastructure, and the engineering decisions behind the hype every week. Subscribe, it is free.
Over to you: are you running embeddings in Postgres, or did you end up on a dedicated vector database? If you migrated in either direction, what triggered it? I read every comment.
Sources
- RIP, vector database (turbopuffer, Sept 30, 2026)
- turbopuffer: fast search on object storage (launch post with the $70/TB cost breakdown)
- turbopuffer v3 progress page (open benchmark tracking)
- pgvector README (index types, limits, partial indexes)
- pgvector performance benchmarks (Supabase methodology and results)
- Notion's turbopuffer workload (10B+ vectors, millions of namespaces)
Top comments (0)