This describes building a production retrieval-augmented generation (RAG) pipeline, Air Context, to enable semantic code search so LLM-driven coding agents can find relevant code by meaning rather than brittle keyword matches. The motivation is that agents working on large codebases waste time and tokens locating the right symbols; a semantic index lets free-text queries return specific functions, classes, or snippets with evidence. The piece frames the project as moving from a simple prototype to a robust, production-grade system and announces a multipart series; this installment focuses on the preprocessing and embedding stages of the pipeline.
The technical heart covers parsing-aware chunking and vectorization. Fixed-size or line-based splits fail semantically, so the system leverages mature parsers for nine languages to produce syntax-node streams, then groups nodes into properly scoped chunks (keeping doc-comments, decorators, prefixes/suffixes together and removing semantically meaningless annotations). Chunks are normalized, paired with metadata, and judged - using an LLM - to detect bad boundaries. Chunk vectors are produced by embedding models so semantically related code sits nearby in high-dimensional space. A major operational finding is storage and cost pressure: millions of chunks yield large indexes (tens of GBs in 32-bit floats), so optimizing bytes-per-vector and retrieval infrastructure is critical. The approach resembles cAST but encodes richer language-specific semantics.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.