hn.today

From the creator of Redis; run LLM locally with ds4

dwarfstar.sh69 points6 comments
Screenshot of From the creator of Redis; run LLM locally with ds4

DwarfStar 4 (ds4) is a compact C inference engine by Salvatore Sanfilippo (antirez) for running frontier open weights locally on high-memory Macs and Linux machines with CUDA or ROCm. It targets a small set of routed-MoE and dense families - DeepSeek V4/V4.1 Flash, GLM 5.x and Qwen 3.8 Flash Next - exposing a unified model state through a CLI, an OpenAI/Anthropic-compatible HTTP server and a native agent. Key engineering choices make large MoE models practical locally: asymmetric 2-bit quantization compresses routed experts while keeping shared paths precise, a disk-backed KV cache persists long prompt prefixes keyed by SHA1 to survive restarts, and the stack validates specific GGUF layouts end-to-end rather than acting as a generic runner.

Operationally, running ds4 involves downloading project GGUF weights, building for the chosen backend (Metal, CUDA, ROCm) and launching ./ds4, ./ds4-server or ./ds4-agent. Hardware guidance spans Apple Silicon (64-512 GB), NVIDIA DGX Spark and ROCm desktops; V4 Flash Q2 is the baseline with larger models streaming from SSD at higher memory. Benchmarks show high prefill throughput (e.g., M5 Max 128 GB q2: ~790 T/s prefill, 39.4 T/s gen; DGX Spark 128 GB q2: ~825 T/s prefill, 18.1-13.8 T/s gen for long contexts). The project is MIT-licensed, available on GitHub, and designed for local, persistent, low-latency inference.

Read on dwarfstar.sh6 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.