hn.today

Rai: CPU-only LLM inference engine in pure Rust

github.com4 points0 comments
Screenshot of Rai: CPU-only LLM inference engine in pure Rust

RAI is a CPU-only LLM inference engine implemented in Rust that runs 4-bit quantized decoder models with hand-written AVX2 kernels - no GPU, CUDA, Python runtime, PyTorch, GGML, or BLAS. It loads a single .raimodel binary and performs on-the-fly dequantization inside GEMM inner loops so no fp32 weight copy exists in RAM. Measured on a 4-core laptop, RAI decodes TinyLlama-1.1B at ~22 tok/s (versus 4.3 tok/s for HuggingFace transformers fp32) while using roughly 1/7 the memory; conversion and streaming conversion tools keep peak RAM low and can convert 7B models on a 16 GB machine. Supported families include Llama-2/3, Mistral-7B, Qwen, Gemma, Phi-3, OLMO2, various MoE models, TinyLlama, SmolLM and their fine-tunes; unsupported architectures are explicitly refused at conversion with actionable diagnostics.

Design choices emphasize auditability, minimal runtime dependencies, and performance portability: AVX2+FMA+F16C optimized paths with scalar fallbacks, one flat validated model file, a lean inference crate (rai-infer) containing loaders, kernels, transformer layers, KV cache, sampling and speculative decoding, and a separate rai-compress toolkit for quantization research. Local serving includes an HTTP chat UI and REST/MCP adapters for agent tooling. The workspace separates inference from auxiliary memory/reasoning crates that are not part of releases. Requirements: Rust (pinned to 1.95.0), x86-64 CPU with AVX2/FMA/F16C recommended, Linux/Windows/macOS; GPU and Python are optional for calibrated export only.

Read on github.com0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.