hn.today

>1B tokens/minute/GPU by combining query planner and inference engine

modal.com7 points2 comments
Screenshot of >1B tokens/minute/GPU by combining query planner and inference engine

Describes Quail (QUery-Aware Inference Layer), an inference engine that fuses SQL-style query planning with transformer inference to accelerate AI-SQL workloads. AI-SQL produces many small LLM requests constructed from database rows (e.g., AI.IF and PROMPT-driven joins) that are intelligible with small models but inefficient for general-purpose inference stacks. Quail exploits the structured query plan to reorder and batch requests, control cache lifetime, prefetch and evict KV entries safely, and avoid the expensive decode phase by mapping boolean and small-class classification onto single-token prefill predictions. That alignment makes transformers a near-perfect fit: high arithmetic intensity on GPUs, linearized KV use, and much higher throughput for many short sequences.

Empirically, Quail delivers dramatic gains: over 1 billion tokens per minute on a single H100 for a multi-join query, more than 10× faster than a vLLM baseline on identical hardware, with a Modal cost of under $0.06 per billion tokens. On a new AI-SQL benchmark Quail is 1.84× faster (geometric mean) than vLLM across tasks. The implementation reduces host overhead and GPU underutilization for many small requests, enables multi-tier KV caching and prefill-only inference patterns, and is presented as open-source, the result of a Modal and CMU Full Stack Data Lab collaboration aimed at making inference+database performance engineering more expert-parallel and cost-efficient.

Read on modal.com2 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.