hn.today

LLaMA Now Goes Faster on CPUs

justine.lol3 points3 comments
Screenshot of LLaMA Now Goes Faster on CPUs

A developer implemented 84 new matrix-multiplication kernels in llamafile, a Cosmopolitan-packaged fork of llama.cpp, to accelerate CPU inference. Benchmarks claim prompt-evaluation speedups of roughly 30%-500% over llama.cpp when using f16 and q8_0 weights, with the biggest wins on ARMv8.2+ (Raspberry Pi 5), Intel Alder Lake, and AVX512-capable Zen 4 CPUs. The kernels reportedly outperform Intel MKL by about 2x for matrices that fit in L2 cache, making the approach especially effective for short prompts (<1,000 tokens). Optimizations currently target q8_0, f16, q4_1, q4_0, and f32 weight formats and change memory-bandwidth behavior enough that quantization can become the new bottleneck rather than raw memory throughput.

Practical examples include TinyLlama-1.1B spam-filtering runs: about 3.2 seconds on an RPI5 and 420 ms on an Alder Lake gaming PC for a single verdict, versus significantly slower times on earlier llamafile and llama.cpp builds. The work also emphasizes system-level choices - avoiding efficiency cores on Alder Lake, using mmap weight loading, and tradeoffs between fp16 speed and potential rounding issues versus q8_0’s dot-product robustness - arguing that faster CPU kernels make local LLMs more practical across hobbyist, gaming, enterprise, and Apple hardware.

Read on justine.lol3 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.