PlanetScale released TIN, a Postgres full-text search extension, and reported large speedups versus ParadeDB on BM25 top-K and COUNT queries, attributing gains to using Postgres ctid values as document identifiers. ParadeDB accepted the challenge, noting TIN was indeed much faster in initial benchmarks (at least 8x vs ParadeDB 0.25), but argued that ctid-based IDs alone couldn’t explain the BM25 top-K gap because ParadeDB defers ctid lookups until after collecting the final Top-K. Rather than rearchitecting around ctids, ParadeDB focused on targeted optimizations and fairer benchmark settings to close the gap.
Two concrete changes closed most of the difference. First, fieldnorms (per-field length data used by BM25) were moved from a global DocId-indexed array into a small, per-term array aligned with each postings list; this turned thousands of random page reads into mostly sequential access (fieldnorm page reads fell from ~1,500 to ~30) at a modest storage cost (~9% larger index). Second, ParadeDB added a MAXSCORE Blockmax pruning path alongside WAND and chose MAXSCORE for dense multi-term disjunctions (≥3 terms), eliminating WAND’s overhead and cutting p50 by ~6x and p95 by ~8x for those queries. On a 28.7M Hacker News top-K benchmark these changes produced 344.6 QPS versus TIN’s 145.7 QPS. COUNT optimizations and other benchmark-details are deferred to a follow-up.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.