Magnitude is an open-source inference engine that runs large language models locally and optimizes itself for the exact hardware it’s installed on. It compiles and tunes compute kernels on-device rather than shipping generic precompiled kernels, and provides hand-optimized kernels for popular open-weight model families. That device-specific tuning yields substantial performance gains versus llama.cpp in their benchmarks (up to 2x overall: for example, 92% faster decode on Apple Metal and 19% faster decode on CUDA), and the engine claims reduced memory usage and faster prefill/decode times. The distribution is a desktop app (macOS, Windows, Linux) that includes a CLI, supports CPU-only setups as well as Apple Silicon, NVIDIA, and AMD GPUs, and exposes a list of supported models with optimized kernels.
Functionally, it connects with existing agent frameworks (one-click integrations for Pi, OpenCode, Hermes, Codex and others or via an OpenAI-compatible API), shares prefix caches for concurrent sessions to avoid slowdown, and frees memory when agents stop. It is privacy-focused and offline-capable once models are downloaded, released under Apache 2.0 with no token costs. The project on GitHub is active and popular - about 5.5k stars, ~394 forks and over 1,000 commits - while documentation, installer downloads and model discovery are integrated into the desktop app.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.