hn.today

Fable 5 – Median thinking declined in August

twitter.com271 points183 comments
Screenshot of Fable 5 – Median thinking declined in August

Lon Lundgren reports a clear drop in Fable 5’s reasoning performance after Anthropic made the model permanently available in subscription plans. Over a six‑week capture and analysis period he measured performance five different ways and found August delivered dramatically fewer “thinking tokens” than July. Reasoning quality declined across the whole window and bounced in multi‑day episodes, with some fluctuations coinciding with product announcements and releases, so the model could feel excellent one day and poor the next.

Despite consistently selecting xhigh or max effort, most invocations received little or no thinking tokens, and the occasional longer reasoning runs rarely reached published benchmark levels. He frames this as an “inference gap”: having access to a frontier model no longer guarantees access to the inference regime that reproduces frontier capabilities. The practical takeaway is to investigate which inference regime is being served (token budgets, routing, or throttling) before assuming a model has been intentionally nerfed; others suggested increased user load and resource dilution as a likely mechanism.

Read on twitter.com183 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

Xiaomi MiMo v2.6

Xiaomi MiMo v2.6

Xiaomi MiMo v2.6 is an open-source language model that demonstrates diverse capabilities, including use in scientific environments and multimedia tasks. The release is seen as a significant step forward in lightweight models at an affordable cost. (mimo.xiaomi.com)

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.