hn.today

What if AI worked at 1.000.000 tokens per seconds?

echohive.ai7 points13 comments
Screenshot of What if AI worked at 1.000.000 tokens per seconds?

A thought experiment imagines a single capable agent generating one million output tokens per second and examines what that speed would change. It teases apart four meanings of “a million tokens” - context size, input processing, aggregate throughput across many streams, and single-agent output - and emphasizes that memory size is not the same as per-request write speed. Concrete worked examples show writing that would take hours at normal speeds collapsing to seconds or fractions of a second: a thousand 1,000-token story endings in ~1s, forty 20,000-token app drafts in ~0.8s, ten thousand 300-token critiques in ~3s. Engineering tricks and parallel streams raise aggregate throughput, but sequential token dependency keeps single-agent generation inherently chained. Faster generation does not equate to better reasoning.

The central finding is that when drafting becomes extremely cheap, judgment becomes the scarce resource. Using Amdahl-style tables, a fixed check (like a 60-second test) turns writing into a tiny fraction of total time, so end-to-end speedups fall far short of raw generation gains; long checks (hours or days) almost erase them. What gains value are taste, specs and tests, physical evidence, fidelity, and cost control. Practical prescriptions: define “done” before generation, set selection criteria up front, identify and shorten the slow validation step where safe, and request diverse assumptions rather than volume alone. This is a thought experiment, not a measurement of current models.

Read on echohive.ai13 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.