Lon Lundgren reports a clear drop in Fable 5’s reasoning performance after Anthropic made the model permanently available in subscription plans. Over a six‑week capture and analysis period he measured performance five different ways and found August delivered dramatically fewer “thinking tokens” than July. Reasoning quality declined across the whole window and bounced in multi‑day episodes, with some fluctuations coinciding with product announcements and releases, so the model could feel excellent one day and poor the next.
Despite consistently selecting xhigh or max effort, most invocations received little or no thinking tokens, and the occasional longer reasoning runs rarely reached published benchmark levels. He frames this as an “inference gap”: having access to a frontier model no longer guarantees access to the inference regime that reproduces frontier capabilities. The practical takeaway is to investigate which inference regime is being served (token budgets, routing, or throttling) before assuming a model has been intentionally nerfed; others suggested increased user load and resource dilution as a likely mechanism.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.