hn.today

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

byteshape.com77 points15 comments
Screenshot of Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

ByteShape released a full ShapeLearn quantization suite for Qwen 3.8 27B that improves on their earlier ShapeLearn-Lite run and pushes the measured quality-speed frontier across six GPU classes. Five dense GGUF models (BPW from ~2.56 to 3.84) were benchmarked against competing quants (AtomicChat, Bartowski, ISTA‑DASLab, Unsloth Dynamic v3). Across all GPUs tested each full ShapeLearn model sits on the frontier - meaning no plotted model is both faster and more accurate - while larger ShapeLearn variants give higher aggregate accuracy and smaller variants yield higher throughput. GPU‑5 (IQ4_XS, 3.84 bpw, ~13.1 GB) is the default recommendation, reaching about 99.63% of the BF16 baseline and high token/s on cards like RTX Pro 6000 and RTX 5090; GPU‑4 (≈11.0 GB) is a strong alternative at ~98.72% BF16 with better speed/size tradeoffs. ISTA‑DASLab also provides a competitive point on the frontier.

Benchmarks show ShapeLearn‑Lite held up well relative to its KLD ranking, with three of six Lite models still frontier‑competitive. Speculative decoding improves throughput for every ShapeLearn model: DFlash2 typically gives the highest text‑only throughput but needs more memory and lacks llama.cpp image support, while MTP supports multimodal inputs with lower memory requirements. Ready‑to‑run commands and recommended sampling settings are provided for each GGUF.

Read on byteshape.com15 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

OpenJev

OpenJev

OpenJev allows users to run decision models directly in their browser using local models like MiniCPM5 2B or Qwen3 0.6B, without needing a backend. It compares two methods: reading option probabilities directly or generating them token by token, with all data staying on the user's device. (openjev.com)

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.