hn.today

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

github.com551 points268 comments
Screenshot of Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

Strata is a free, open-source system that runs large models - specifically QWEN3.8-FLASH-NEXT (125B) - locally on consumer gaming PCs so chats, coding, image understanding and agent workflows stay on your machine. It automates hardware detection and installation (Windows/Linux), streams a ~70 GB model download, and exposes an OpenAI/Anthropic-compatible HTTP API and a browser UI at localhost for chat, monitoring and settings. Minimum practical requirements are a GPU with 12 GB+ VRAM, 32+ GB RAM (64 GB recommended for larger sizes), and ~80 GB disk (SSD recommended). Multi-GPU setups are supported; AMD and older GPUs and CPUs work in community-tested, experimental modes. The installer chooses model size and context length for your RAM and offers image support, adjustable “thinking” effort, and parallel request options.

The repository includes performance measurements and model-sizing specifics: smaller compressed builds trade speed for quality while larger sizes increase capability but need more RAM or SSD paging. Examples show consumer cards achieving tens to low hundreds of tokens/sec (the project highlights ~100 tokens/s on high-end cards; an RTX 3090 is cited at ~100-140 tokens/s). Variants include a CODER build that fits 32 GB of RAM, IQ2/IQ3 sizes for 48-96+ GB, and Unsloth ~4-bit releases that reduce memory at some speed cost because parts may be streamed from SSD. First startup can pause the system for 1-3 minutes while 35-55 GB of RAM is reserved; troubleshooting and detailed setup docs are included.

Read on github.com268 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.