hn.today

Qwen3.8-27B at ~200 tok/s peak on an Apple M5 Max

twitter.com4 points0 comments
Screenshot of Qwen3.8-27B at ~200 tok/s peak on an Apple M5 Max

Zhihao Jia announces an open-source runtime called lithos-metal that combines megakernels and a DSpark speculative decoding system to run Qwen3.8-27B at over 200 tokens per second peak per user on a single Apple M5 Max. The claim emphasizes ultra-fast, laptop-scale inference and one-command integration with coding agents; code is released on GitHub and a technical writeup is linked. The core specifics are the use of megakernel kernels plus speculative decoding to parallelize and prefetch generation work, the target model (Qwen 3.8, 27B parameters), and the benchmarked peak throughput number (200+ tok/s/user) on Apple’s M5 Max hardware.

Responses highlight practical limits and skepticism about sustained performance: commenters point out that Apple silicon hits a memory-bandwidth bottleneck as the KV cache grows, so the 200+ tok/s figure may be a short-lived peak rather than steady throughput in realistic multi-turn agent loops with tool calls and retrieval. The announcement generated notable attention and excitement, but the substantive takeaway is a promising engineering step that delivers remarkable peak latency and throughput on consumer silicon while leaving open questions about sustained multi-agent or long-context performance under real-world workloads.

Read on twitter.com0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.