hn.today

RIP, vector database

turbopuffer.com374 points105 comments
Screenshot of RIP, vector database

turbopuffer is redesigning its storage engine in v3 to stop treating the approximate nearest neighbor (ANN) index as the primary key and instead make ANN a secondary index. The system began as a serverless vector database that stored each document keyed by an ANN address built from hierarchical clustering (SPANN/SPFresh), which made vector search on object storage cheap and fast. As features were added - attribute filtering via inverted indexes and BM25 full-text search - those indexes were implemented as postings that point to ANN addresses, and non-vector data was stored alongside vectors. That single ANN-centric layout supported massive single-index ANN workloads but constrained other query shapes and data layouts.

Three concrete technical problems drove the redesign: storage amplification from duplicating full document contents for multi-vector representations, write amplification because SPFresh rebalances move complete documents and all their inverted-index references, and limited vectorization since all query plans are forced into small ANN cluster block sizes (~100-200) rather than their optimal batch sizes. The v3 change is to stop keying on ANN addresses, enabling separate storage layouts for postings and documents, larger vectorized blocks, and lower write/storage amplification. v3 reached 100% CI passing but initially regressed on performance; tuning and benchmark updates are in progress to recover and surpass current results.

Read on turbopuffer.com105 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in Security

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.