hn.today

Show HN: TurboGPT: train 22KiB transformer in 13s

github.com54 points10 comments
Screenshot of Show HN: TurboGPT: train 22KiB transformer in 13s

TurboGPT is a compact CUDA C++ implementation that trains a byte-level GPT model extremely quickly: the project advertises training a 22 KiB transformer in about 13 seconds (or "under a minute" for tiny models) using NVIDIA GPUs. The codebase (MIT licensed) targets Windows with Visual Studio 2022 and CUDA 13.4, and is built with a provided PowerShell script that accepts a CudaArch argument. Core components include C++ source files for model, trainer, sampling, dataset handling, checkpointing, optimization, and logging; a run example shows launching the executable on a text dataset and saving checkpoints to runs/ctx4/ctx4.pt.

The implementation emphasizes practical training workflow details: a sample run reports 2.52435 bits-per-byte after 1.5 billion training tokens, logs TensorBoard-compatible reports per batch (capped at 8Mi reports), and stores full trainer/optimizer/scheduler state to allow resuming. Tests are available via python tests\verify.py. The repository is focused on minimal, high-performance GPU-only training for tiny transformers, offering reproducible commands, checkpoints, and metrics for rapid experimentation and verification.

Read on github.com10 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.