Kapa built a Company Knowledge Bench: a 1,000-case retrieval benchmark drawn from real, messy company data (docs, tickets, chat, code) to measure how different retrieval strategies perform on practical agent use cases. The benchmark defines retrieval success tightly: the retriever must return a minimal, complete set of chunks that fully answer the query and prefer higher-authority and more current sources. It encodes rules for ambiguity, multi-solution problems, and no-answer situations. Each eval case pairs a real query with a corpus snapshot and a formal retrieval criterion (a boolean expression naming acceptable chunks). Because hand-labeling at scale is impractical, Kapa first produced 170 human-labelled cases to train and validate labeling agents, then used a two-stage agent pipeline - candidate agents for recall and criteria agents to formalize selections - to generate the full 1,000 cases after reaching high agreement with humans.
Seven retrievers were compared on score versus time per query: Hybrid search (0.41), Hybrid+rerank (0.50), Query decomposition+rerank (0.56), Kapa Default (0.61), Kapa Deep (0.65), and two agentic grep variants (Luna 0.54, Sol 0.61). Key findings: a frontier model using only grep can match a tuned modern pipeline score (~0.61) but is about five times slower, while an optimized agentic retriever (Kapa Deep) reaches 0.65 in ~5s. Results note a conflict of interest: all retrievers and corpora were built and ingested by Kapa.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.