This work investigates how post-training quantization (PTQ) affects large language models tuned for step-by-step reasoning across math, coding, and science QA. Aggressive PTQ is found to lower final-answer accuracy while paradoxically lengthening chain-of-thought (CoT) outputs. A striking pattern emerges: up to 52% of quantized-model failures contain the correct answer in intermediate reasoning steps but fail to surface it as the final response, a phenomenon the authors call overthinking. Token-level KL divergence between quantized and full-precision output distributions pinpoints where deviations occur; those high-KL positions align with high next-token entropy and an increased tendency for quantized models to produce hesitation or redirection markers (e.g., wait, but, alternatively).
To mitigate this, a simple, training-free intervention applies a logit penalty to a curated set of overthinking markers. This penalty shortens CoT length by 12-23% while preserving or improving accuracy across five model sizes (1.5B-32B), three quantization methods, and five benchmarks, and it cuts overthinking errors by up to 58%. The method produces a favorable accuracy-versus-reasoning-cost Pareto frontier compared with penalizing other tokens. The practical takeaway is that quantized reasoning models do not truly need longer deliberations; targeted logit adjustments reclaim efficiency and reliability without retraining.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.