hn.today

The Query Transformation Pipeline

readyset.io14 points1 comments
Screenshot of The Query Transformation Pipeline

Readyset explains how it converts arbitrary SQL into a form its streaming dataflow engine can maintain incrementally, shifting work from repeated query execution to continuous view maintenance. The core argument is that traditional pull-based engines re-run entire plans on each read, while a compiled dataflow graph propagates upstream changes through operators (joins, filters, aggregates) so reads become lookups into materialized results. That shift imposes structural constraints - binary equality joins, no per-row correlated execution, explicit GROUP BY keys, and only INNER/LEFT/CROSS joins are natively supported - so many valid SQL queries must be rewritten to be compilable and efficient.

The pipeline is a multi-pass, semantics-preserving transformer organized into normalization, deep rewrites, and cleanup. Normalization resolves schemas, expands stars, qualifies columns, and desugars USING. Deep rewrites perform heavy lifting: rewriting PostgreSQL ARRAY(SELECT...) into LATERAL + array_agg + COALESCE, eliminating redundant self-joins, hoisting top-k derived tables so LIMIT stays parameterizable for a native Top-K operator, and decorrelating subqueries (turning correlated subqueries into joins or grouped derived tables) to allow incremental maintenance. Cleanup removes redundancies and parameterizes literals. These targeted passes reduce intermediate materialization, preserve semantics for all data, and produce canonical queries that the dataflow compiler can turn into efficient, incrementally maintained graphs.

Read on readyset.io1 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in Other

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.