hn.today

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

github.com164 points84 comments
Screenshot of Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

Magnitude is an open-source inference engine that runs large language models locally and optimizes itself for the exact hardware it’s installed on. It compiles and tunes compute kernels on-device rather than shipping generic precompiled kernels, and provides hand-optimized kernels for popular open-weight model families. That device-specific tuning yields substantial performance gains versus llama.cpp in their benchmarks (up to 2x overall: for example, 92% faster decode on Apple Metal and 19% faster decode on CUDA), and the engine claims reduced memory usage and faster prefill/decode times. The distribution is a desktop app (macOS, Windows, Linux) that includes a CLI, supports CPU-only setups as well as Apple Silicon, NVIDIA, and AMD GPUs, and exposes a list of supported models with optimized kernels.

Functionally, it connects with existing agent frameworks (one-click integrations for Pi, OpenCode, Hermes, Codex and others or via an OpenAI-compatible API), shares prefix caches for concurrent sessions to avoid slowdown, and frees memory when agents stop. It is privacy-focused and offline-capable once models are downloaded, released under Apache 2.0 with no token costs. The project on GitHub is active and popular - about 5.5k stars, ~394 forks and over 1,000 commits - while documentation, installer downloads and model discovery are integrated into the desktop app.

Read on github.com84 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.