For two years the default answer to “we need similarity search” was a new database. A whole product category grew up around the idea that vectors were special enough to be the primary key of their own system. I never bought one, mostly out of laziness: my data already lived in Postgres, pgvector was good enough, and running a second database plus a sync job felt like paying twice for the same rows.
On September 30, turbopuffer announced it agrees with the lazy version of me. The company launched as a serverless vector database, and counts Cursor and Notion among its earliest customers by its own account. Its new v3 architecture demotes the ANN vector index from the primary index to just another secondary index. The announcement post is literally titled “RIP, vector database,” from a vector database vendor. That took some nerve, and the engineering reasoning underneath is the most interesting storage post I have read this year.
Keyed by where the vector lives
To see what is being buried, you have to see what turbopuffer v1 was. Documents were nothing but an ID and a vector. Vectors got clustered into groups, centroids clustered in turn, up to a single root: a tree, built for object storage with a clustering index rather than the graph indexes everyone else used.
The load-bearing detail is what everything was keyed by. Each cluster got an ID, each vector a dense local ID inside it, and the pair, something like C0L1, was the address of the document. turbopuffer calls this the ANN address, and it was the primary key. The document, its ID, and later all of its attributes lived under the address that described where its vector sat in the cluster tree.
Then came the query shapes every customer actually wanted. Attribute filtering arrived as an inverted index mapping attribute values to ANN addresses. Full-text search arrived the same way, with postings pointing at ANN addresses plus term counts for BM25 scoring. Aggregations, regex, fuzzy matching, ordering by attributes: all of it bolted onto the same vector-primary layout.
This worked. turbopuffer reports pushing single indexes past 100 billion vectors while serving 200 millisecond p99 reads at over a thousand queries per second. Those are the vendor’s own numbers, and even discounted they describe an engine that did exactly what it was built to do.
Everything moves when one vector moves
The problem is what the layout does to every query that is not a vector search.
Start with writes. The clustering has to stay balanced or recall decays, so on insert, update, or delete the engine rebalances vectors between clusters. But the full document contents are stored under the ANN address, so rebalancing a vector moves the whole document with it, plus every inverted index entry that referenced the old address. Updating one vector can relocate hundreds of attributes and their index entries. turbopuffer says its indexing-throughput tuning has started hitting diminishing returns against this amplification, which is the kind of sentence vendors only write when the wall is real.
Storage has the same shape of problem. One vector per document means the payload is stored once. But multi-vector representations, document nesting or late interaction, duplicate the whole document contents per vector. That is not a corner case; it is where retrieval quality has been heading for two years.
The third one is my favorite because it is the least obvious. Modern query engines scan in blocks: DuckDB works in batches of 2,048 rows, ClickHouse up to around 65,000, Lucene posting blocks hold 256 documents, and turbopuffer’s own ANN clusters hold 100 to 200. Block size is a per-engine tuning choice until your primary key takes it away. With the ANN tree as the primary index, every scan reads documents one cluster at a time, capped at cluster size, even when the query plan wants blocks of thousands to keep the CPU pipeline full. turbopuffer already proved the point against itself once: its first full-text search partitioned postings along cluster boundaries, the median block held about 1.5 postings, and reworking postings into fixed blocks of 256 made the index 10 times smaller and queries up to 20 times faster. Those are turbopuffer’s numbers again. Postings could escape because they are stored separately. Document scans could not, because the documents lived inside the clusters.
The fix is a stable row id
So v3 stops keying on the ANN address. Documents get a stable identity, and the ANN tree becomes one secondary index among others, pointing at rows instead of containing them. Rebalancing moves pointers, not payloads. Block sizes become each engine’s own choice again. Aggregations and group-bys stop paying a tax levied by a clustering they never asked for.
If this sounds familiar, it is because it is a row store with indexes. Postgres was right all along, not about vectors specifically but about the general principle: the primary key should be the thing that does not move, and the fancy access path hangs off the side.
Would I still run a vector store
Here is where I would take a position. The dedicated vector store earns its keep in exactly one situation: ANN dominates your query mix and your scale is past what a general engine handles well. Hundred-billion-vector indexes at a thousand queries per second is a real place, and turbopuffer got there by specializing. Most of us do not live there.
If I were mid-migration off pgvector today, the question I would actually check is what share of my queries are pure similarity. The moment the workload is filters plus full-text plus aggregations with similarity as one predicate among several, the general-primary engine wins, because every bolted-on query shape in a vector-primary system pays storage or write amplification for the privilege. That was my lazy instinct two years ago. Now the vendor that built the other side has published the same conclusion with benchmarks attached.
The category is not dead, whatever the title says. But “vector database” as the default answer to retrieval is over, and the company that just told you so sells one. That is the part worth sitting with. turbopuffer is betting its v3 that the future is a general engine where vectors are one index type among many, which is to say it is betting on becoming the thing Postgres already is, with better object-storage economics. I would take that bet from the Postgres side.
