RAI is a CPU-only LLM inference engine implemented in Rust that runs 4-bit quantized decoder models with hand-written AVX2 kernels - no GPU, CUDA, Python runtime, PyTorch, GGML, or BLAS. It loads a single .raimodel binary and performs on-the-fly dequantization inside GEMM inner loops so no fp32 weight copy exists in RAM. Measured on a 4-core laptop, RAI decodes TinyLlama-1.1B at ~22 tok/s (versus 4.3 tok/s for HuggingFace transformers fp32) while using roughly 1/7 the memory; conversion and streaming conversion tools keep peak RAM low and can convert 7B models on a 16 GB machine. Supported families include Llama-2/3, Mistral-7B, Qwen, Gemma, Phi-3, OLMO2, various MoE models, TinyLlama, SmolLM and their fine-tunes; unsupported architectures are explicitly refused at conversion with actionable diagnostics.
Design choices emphasize auditability, minimal runtime dependencies, and performance portability: AVX2+FMA+F16C optimized paths with scalar fallbacks, one flat validated model file, a lean inference crate (rai-infer) containing loaders, kernels, transformer layers, KV cache, sampling and speculative decoding, and a separate rai-compress toolkit for quantization research. Local serving includes an HTTP chat UI and REST/MCP adapters for agent tooling. The workspace separates inference from auxiliary memory/reasoning crates that are not part of releases. Requirements: Rust (pinned to 1.95.0), x86-64 CPU with AVX2/FMA/F16C recommended, Linux/Windows/macOS; GPU and Python are optional for calibrated export only.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.