hn.today

How many GPUs is 1M/B/T tokens?

cedana.com16 points6 comments
Screenshot of How many GPUs is 1M/B/T tokens?

A tokens-to-GPUs calculator turns a token-volume target into a hardware plan by combining a chosen model, GPU type, serving time window, token mix (input vs output), average utilization and per-GPU throughput estimates. It uses presets (measured when available, estimated otherwise) for tokens-per-second per GPU, allows swapping GPU reference types, and exposes advanced assumptions - input-token cost, multi‑GPU efficiency and memory fit - so planners can replace defaults with their own measurements. The tool shows a headline GPU count, a plausible low/high range, per‑GPU detail, a sensitivity chart that ranks which assumptions move the result most, and clear limits (rough 2× planning error, latency-dependence, no prompt caching or speculative decoding, and theoretical values for some older GPUs).

Using the calculator: converting 1 trillion tokens in 30 days for Llama 3.3 70B with a 47.5% output token share and 50% utilization yields 367 H100 GPUs (range 211-853). Equivalent base estimates for other hardware are ~815 A100, ~2,547 V100 (theoretical), and ~92 B200 single GPUs; in 8‑GPU nodes those are ~46 8×H100 nodes or ~12 8×B200 nodes. The sensitivity analysis shows per‑GPU speed, average utilization and input-token cost drive the result most, so measuring model throughput and choosing utilization targets are the priority actions for accurate capacity planning.

Read on cedana.com6 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

AI Model Groupthink

AI Model Groupthink

Asking multiple AI models for their top opinions reveals that Chinese models tend to align more with the mainstream consensus, while others show more contrarian responses. The study scores models based on their agreement with group consensus, highlighting regional differences in AI responses. (magicnumbers.io)

Claude Haiku 5.5

Claude Haiku 5.5

Claude Haiku 5.5 is the most capable small model from Anthropic, offering a 75% reduction in running costs compared to its predecessor. It is optimized for high-volume, cost-sensitive tasks like summaries, classification, and customer support, with improvements in alignment and efficiency. (twitter.com)

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.