ByteShape released a full ShapeLearn quantization suite for Qwen 3.8 27B that improves on their earlier ShapeLearn-Lite run and pushes the measured quality-speed frontier across six GPU classes. Five dense GGUF models (BPW from ~2.56 to 3.84) were benchmarked against competing quants (AtomicChat, Bartowski, ISTA‑DASLab, Unsloth Dynamic v3). Across all GPUs tested each full ShapeLearn model sits on the frontier - meaning no plotted model is both faster and more accurate - while larger ShapeLearn variants give higher aggregate accuracy and smaller variants yield higher throughput. GPU‑5 (IQ4_XS, 3.84 bpw, ~13.1 GB) is the default recommendation, reaching about 99.63% of the BF16 baseline and high token/s on cards like RTX Pro 6000 and RTX 5090; GPU‑4 (≈11.0 GB) is a strong alternative at ~98.72% BF16 with better speed/size tradeoffs. ISTA‑DASLab also provides a competitive point on the frontier.
Benchmarks show ShapeLearn‑Lite held up well relative to its KLD ranking, with three of six Lite models still frontier‑competitive. Speculative decoding improves throughput for every ShapeLearn model: DFlash2 typically gives the highest text‑only throughput but needs more memory and lacks llama.cpp image support, while MTP supports multimodal inputs with lower memory requirements. Ready‑to‑run commands and recommended sampling settings are provided for each GGUF.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.