turbopuffer began as a serverless vector database optimized for extremely cheap, fast approximate nearest neighbor (ANN) search by storing each document as an ID and a vector and organizing vectors into a hierarchical clustering tree (SPANN → SPFresh). As attribute filtering and BM25 full-text search were added, inverted indexes were implemented that map attribute values and terms to ANN addresses, and document attributes were stored alongside vectors. Over time many query plans - aggregations, regex, fuzzy matching, sparse vectors - were built around the ANN address as the primary key. That layout works very well for large-scale ANN workloads but creates three concrete limits: storage amplification from duplicating document contents for multi-vector documents, write amplification when SPFresh rebalances clusters and forces moving full documents and all inverted-index pointers, and constrained vectorization because all query plans are locked to cluster-sized blocks (~100-200 documents) rather than their optimal block sizes.
turbopuffer v3 replaces the ANN-address primary key: ANN becomes a secondary index and documents are keyed independently, eliminating much duplication, reducing cascading writes, and allowing query engines to use block sizes suited to each plan. The redesign is substantial; CI is green across the codebase, but initial performance regressed versus production as correctness and foundational design were prioritized. Work is now focused on tuning, benchmark-driven optimizations, and publishing results so the new storage layout can restore or exceed prior ANN performance while unlocking better efficiency and scale for non-vector queries.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.