hn.today

Tokens Too Cheap to Meter

jyn.dev343 points223 comments
Screenshot of Tokens Too Cheap to Meter

This argues that the cost of using machine-learning intelligence is collapsing by orders of magnitude and will turn LLMs from standalone products into ubiquitous computing infrastructure within a year or two, with frontier-quality models running locally on commodity hardware within 3-6 years. Evidence spans hardware, models, and software: GPUs are doubling efficiency roughly every two years; model cost-per-task has fallen dramatically (a two-orders-of-magnitude drop in 2025-2026 on the pareto frontier); inference engines like vLLM and vendor stacks show 40-50% efficiency gains in months; and advances in architectures and serving tech are shifting limits from token counts to model quality and access.

Concrete architectural and product developments make this plausible and consequential. Mixture-of-experts yields much higher quality-per-joule and smaller effective models; Mamba-style transformer hybrids cut VRAM needs by ~5x so a Nemotron-quality model can hold a million tokens in 32 GB versus ~120 GB for older models; and specialized classifiers like Jev and Laya deliver yes/no or probability outputs at trivial per-token cost (example pricing cited at about $42 per billion input tokens), enabling very cheap streaming tools such as jgrep. The net result is tokens becoming cheaper than tool calls, creating supply- and demand-side Jevons effects and forcing new business and infrastructure models around essentially free token access.

Read on jyn.dev223 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.