hn.today

Why didn't we get GPT-2 in 2005?

dynomight.net5 points2 comments
Screenshot of Why didn't we get GPT-2 in 2005?

An inquiry into why GPT-2-level language models didn't exist in 2005 shows that the core barrier was not pure hardware but the combination of algorithmic knowledge, incentives, and investment. Training GPT-2 likely required on the order of 10^21 FLOPs: various back-of-envelope estimates include ~1.9×10^20 FLOPs from the Chinchilla formula for a single pass, ~8×10^20 from two-day TPUv3 reports, and papers reporting up to 2.5×10^21; rounding yields 10^21 as a convenient figure. At BlueGene/L's peak achieved throughput in 2005 (~2.8×10^14 FLOPs/sec) that computation would take roughly 41 days, so compute availability alone could have sufficed. However, transformers, Adam, layer normalization, modern tokenizers and other algorithmic innovations did not exist then.

More fundamentally, the absence resulted from uncertainty and lack of a feedback loop: researchers lacked a clear roadmap, didn't know the costs or potential performance, and so did not pour resources into large-scale experimentation. As compute became cheaper, experiments produced promising results, which attracted more researchers and money, creating an accelerating cycle that produced rapid progress after about 2018. Demonstrations of capability, not secret low-level details, collapse uncertainty and trigger broad investment; while a monopolistic strategy might in principle delay diffusion, openness and demos have in practice accelerated adoption.

Read on dynomight.net2 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.