hn.today

Let's See Paul Allen's SIMD CSV Parser

chunkofcoal.com17 points0 comments
Screenshot of Let's See Paul Allen's SIMD CSV Parser

Explains a SIMD-based approach to CSV parsing that processes many bytes at once to find structural characters (commas, newlines, quotes) in a branchless, high-throughput way. It positions simdjson as the motivating prior work and notes this implementation is research-focused (hand-waving validation) using ARM NEON intrinsics in Rust, while x86 equivalents exist. The core idea is to treat parsing as three steps: classify bytes, filter out delimiters that occur inside quoted fields, and record offsets of the remaining structural characters so fields can be sliced from the original byte stream.

The classification uses a clever nibble-based vectorized lookup: split each byte into high and low nibbles, use two 16-entry lookup tables to produce candidate classes, AND the results to yield exact classifications (implemented with NEON ops like vld1q_u8, vqtbl1q_u8, vshrq_n_u8, vandq_u8, vst1q_u8). Classified bytes are compressed into per-class bitmasks (u64 per 64 bytes). Filtering determines quoted regions by maintaining parity over quote positions - odd means inside a quote - so structural characters inside quotes are masked out. The writeup covers practical details like remainder padding, the lookup-table grid false-positive check (unique highs × lows equals class size), and how the resulting bitmasks give fast offsets for slicing rows and fields.

Read on chunkofcoal.com0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in Programming

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.