Google AI Edge Foresight – offline, private meeting transcripts
Google AI Edge Foresight enables offline, private transcription of meeting content. It aims to enhance data privacy and accessibility for users. (developers.google.com)
A tokens-to-GPUs calculator turns a token-volume target into a hardware plan by combining a chosen model, GPU type, serving time window, token mix (input vs output), average utilization and per-GPU throughput estimates. It uses presets (measured when available, estimated otherwise) for tokens-per-second per GPU, allows swapping GPU reference types, and exposes advanced assumptions - input-token cost, multi‑GPU efficiency and memory fit - so planners can replace defaults with their own measurements. The tool shows a headline GPU count, a plausible low/high range, per‑GPU detail, a sensitivity chart that ranks which assumptions move the result most, and clear limits (rough 2× planning error, latency-dependence, no prompt caching or speculative decoding, and theoretical values for some older GPUs).
Using the calculator: converting 1 trillion tokens in 30 days for Llama 3.3 70B with a 47.5% output token share and 50% utilization yields 367 H100 GPUs (range 211-853). Equivalent base estimates for other hardware are ~815 A100, ~2,547 V100 (theoretical), and ~92 B200 single GPUs; in 8‑GPU nodes those are ~46 8×H100 nodes or ~12 8×B200 nodes. The sensitivity analysis shows per‑GPU speed, average utilization and input-token cost drive the result most, so measuring model throughput and choosing utilization targets are the priority actions for accurate capacity planning.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.
Google AI Edge Foresight enables offline, private transcription of meeting content. It aims to enhance data privacy and accessibility for users. (developers.google.com)
Meta and Microsoft are reducing employee use of Anthropic’s Claude AI, shifting focus to their own AI tools. Microsoft is cutting internal AI spending, while Meta is replacing Claude with its proprietary models for internal use. (rswebsols.com)
Qwen 3.8-27B, a free local language model, replaces paid cloud AI services like ChatGPT and Claude for a user’s daily tasks. It provides similar performance at no cost, removing the need for subscriptions. (xda-developers.com)
Asking multiple AI models for their top opinions reveals that Chinese models tend to align more with the mainstream consensus, while others show more contrarian responses. The study scores models based on their agreement with group consensus, highlighting regional differences in AI responses. (magicnumbers.io)
Claude Haiku 5.5 is the most capable small model from Anthropic, offering a 75% reduction in running costs compared to its predecessor. It is optimized for high-volume, cost-sensitive tasks like summaries, classification, and customer support, with improvements in alignment and efficiency. (twitter.com)
Claude Haiku 5.5 is the most affordable and fastest small model released by Anthropic, optimized for high-volume, cost-sensitive tasks. It offers improved performance and new adjustable effort settings for users to balance cost and intelligence. (anthropic.com)
Today's best Hacker News stories, summarized and screenshotted, one email a day.