A thought experiment imagines a single capable agent generating one million output tokens per second and examines what that speed would change. It teases apart four meanings of “a million tokens” - context size, input processing, aggregate throughput across many streams, and single-agent output - and emphasizes that memory size is not the same as per-request write speed. Concrete worked examples show writing that would take hours at normal speeds collapsing to seconds or fractions of a second: a thousand 1,000-token story endings in ~1s, forty 20,000-token app drafts in ~0.8s, ten thousand 300-token critiques in ~3s. Engineering tricks and parallel streams raise aggregate throughput, but sequential token dependency keeps single-agent generation inherently chained. Faster generation does not equate to better reasoning.
The central finding is that when drafting becomes extremely cheap, judgment becomes the scarce resource. Using Amdahl-style tables, a fixed check (like a 60-second test) turns writing into a tiny fraction of total time, so end-to-end speedups fall far short of raw generation gains; long checks (hours or days) almost erase them. What gains value are taste, specs and tests, physical evidence, fidelity, and cost control. Practical prescriptions: define “done” before generation, set selection criteria up front, identify and shorten the slow validation step where safe, and request diverse assumptions rather than volume alone. This is a thought experiment, not a measurement of current models.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.