hn.today

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not

arxiv.org13 points1 comments
Screenshot of Quantized Reasoning Models Think They Need to Think Longer, but They Do Not

This work investigates how post-training quantization (PTQ) affects large language models tuned for step-by-step reasoning across math, coding, and science QA. Aggressive PTQ is found to lower final-answer accuracy while paradoxically lengthening chain-of-thought (CoT) outputs. A striking pattern emerges: up to 52% of quantized-model failures contain the correct answer in intermediate reasoning steps but fail to surface it as the final response, a phenomenon the authors call overthinking. Token-level KL divergence between quantized and full-precision output distributions pinpoints where deviations occur; those high-KL positions align with high next-token entropy and an increased tendency for quantized models to produce hesitation or redirection markers (e.g., wait, but, alternatively).

To mitigate this, a simple, training-free intervention applies a logit penalty to a curated set of overthinking markers. This penalty shortens CoT length by 12-23% while preserving or improving accuracy across five model sizes (1.5B-32B), three quantization methods, and five benchmarks, and it cuts overthinking errors by up to 58%. The method produces a favorable accuracy-versus-reasoning-cost Pareto frontier compared with penalizing other tokens. The practical takeaway is that quantized reasoning models do not truly need longer deliberations; targeted logit adjustments reclaim efficiency and reliability without retraining.

Read on arxiv.org1 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.